agentsclimarketplace

Geo tracking

Skill Thibaultbm/claude-seo-geo/skills/geo-tracking

SEO & GEO skills for Claude Code, built with Claude Mythos 5. Rank in Google AND in LLMs (ChatGPT, Perplexity, Gemini): technical audits, backlinks, AI-optimized content, local SEO, and social amplification.

Install
npx -y skills add Thibaultbm/claude-seo-geo --skill geo-tracking

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 12 stars12 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Measure AI visibility without paid tools or API keys. Input: your site (GA4 and server logs) and a buyer prompt panel. Output: GA4 AI-traffic reporting (custom channel group plus referrer regex above Referral), monthly brand mention rate, citation rate, and share of voice versus competitors across ChatGPT, Perplexity, AI Overviews, Claude, and Gemini, plus verified AI-crawler hits from logs. Repo reference for measurement. Use to track AI traffic, monitor brand mentions, run prompt tracking, compute share of voice, or build a monthly GEO report.

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

21.1 KB, as published. Nobody here has run it

GEO Tracking: Measuring AI Visibility

AI visibility is measurable today, for free, at three layers of the funnel. Each layer answers a different question:

LayerData sourceQuestion it answersLag
TrafficGA4How many visits and conversions do AI assistants send?Trailing (after the click)
AnswersPrompt panelWhat do AI engines say about the brand when buyers ask?Current state
RetrievalServer logsAre AI systems reading the pages right now?Leading (before citations appear)

Run all three. GA4 alone undercounts structurally, panels alone miss revenue proof, logs alone say nothing about what answers contain. The combination is the measurement system; paid platforms are an optional layer on top, never the starting point.

This skill is the canonical measurement reference in this repo. Build the prompt panel with seo-keyword-research; act on the gaps with geo-visibility.

Company knowledge first (Obsidian)

If the working environment contains an Obsidian vault or any local knowledge base (a folder of .md notes, often with a .obsidian directory), read the relevant notes before acting: brand and product facts, target keywords, competitors, and the SEO action log of what was already tried. Ground every recommendation in that context instead of asking the user for facts the vault already holds. At the end of the session, append the actions taken to the vault's SEO action log so the next session starts informed. Vault structure, read-first and write-back protocols: the obsidian-brain skill.

When to use this skill

Use this skill when the user:

  • Asks how much traffic comes from ChatGPT, Perplexity, Gemini, Claude, or Copilot.
  • Wants AI traffic visible in GA4, Looker Studio, or a client report.
  • Wants to track brand mentions, citations, or share of voice in AI answers (prompt tracking, AI rank tracking).
  • Asks whether AI bots crawl the site, or what ChatGPT-User hits in the logs mean.
  • Needs a monthly GEO report, a baseline before a GEO project, or proof of GEO results for a client.
  • Asks whether to buy an AI visibility tool (Profound, Otterly, Brand Radar, Semrush AI tools).

Workflow

Day zero: capture the baseline

Do this before any GEO work starts, because every later claim of progress is judged against it; without a baseline, improvements are invisible and unprovable.

  1. Create the custom GA4 channel group (Layer 1 below) and annotate the property with the date.
  2. Run the full prompt panel once: that run is month zero.
  3. Pull 30 days of server logs and count verified AI hits per agent (Layer 3 below).
  4. Screenshot the current AI answers for the 10 most commercial prompts, brand present or not.

Layer 1: AI traffic in GA4

Step 1: Know what the default gives you

Since May 13, 2026, GA4 includes a default "AI Assistant" channel group. It captures referrers from chatgpt, gemini, and claude, but not perplexity and not copilot, and it is not retroactive: it only classifies data from its activation forward (https://www.searchenginejournal.com/google-analytics-adds-ai-assistant-as-default-channel-group/574974/). Useful as a zero-effort baseline, insufficient as the measurement.

Step 2: Build a custom channel group (the real setup)

Why: full assistant coverage, and a definition you control and can extend when new assistants appear.

  1. In GA4: Admin, Data display, Channel groups, create a new channel group (or copy the default).
  2. Add a channel named "AI Assistants" with the condition Session source matches the regex:
chatgpt\.com|chat\.openai\.com|perplexity\.ai|claude\.ai|gemini\.google\.com|copilot\.microsoft\.com|bing\.com/chat|deepseek\.com|grok\.com|meta\.ai|mistral\.ai|you\.com
  1. Drag the "AI Assistants" channel above Referral in the channel order. Why: channel groups evaluate top-down, and these sources all qualify as referrals; placed below Referral, every AI session is swallowed by Referral first and the channel stays empty (setup walkthrough: https://www.analyticsmania.com/post/ai-traffic-in-google-analytics-4/).
  2. Custom channel groups are not retroactive either. Create the group on day one of any GEO project and annotate the date; the baseline starts there.
  3. Review the regex quarterly and append new assistant domains.

Step 3: Read the numbers honestly

Structural limit: native mobile and desktop apps frequently send no referrer, and those sessions land in Direct. Measured AI traffic is a floor, not the total. State this in every report; never present the AI channel as exhaustive, and never read a Direct uptick as noise without considering assistant apps.

Step 4: Report value, not just volume

AI assistant visits convert about 4.4x better than organic search visits (Semrush study, reported at https://authoritytech.io/blog/ai-traffic-attribution-how-to-track-chatgpt-perplexity-gemini). The pattern: the assistant did the comparison work, so the visitor arrives pre-sold. Always report conversion rate and conversions next to session counts; AI traffic that looks small by sessions is often material by revenue.

Also segment landing pages: in Explore, build a free-form report with Landing page as dimension and a filter on Session source matching the AI regex. The URLs receiving AI sessions are the pages engines already like; their structure is the template to replicate via geo-visibility.

Step 5: Know what Google Search Console cannot tell you

Google includes AI Overviews impressions and clicks inside the regular Search performance data with no separate breakdown (https://developers.google.com/search/docs/appearance/ai-features). There is no "AI Overviews" filter to pull. Consequence: the prompt panel (Layer 2) is the only direct observation of AI Overviews presence available without paid tools. Treat unexplained GSC click drops on stable rankings as a possible AI Overview substitution effect, then verify with the panel.

Layer 2: The prompt panel protocol (share of voice)

This is the heart of the skill: a reproducible, free protocol measuring what buyers actually see in AI answers.

Why a panel at all: AI engines publish no impression count and no search volume for a brand mention, so visibility inside them is a black box. The panel is how you mime ranking measurement: send the questions buyers would actually ask, record whether the brand appears, and track that rate over time. It is a sampled proxy, not a traffic metric, and that is the best signal available without paid tools (field heuristic from 115+ agency audits).

Step 1: Freeze the panel

Take the 50-100 buyer prompt panel built with seo-keyword-research (patterns: "best [category] for [use case]", "[brand] vs [competitor]", "is [brand] worth it", "how to [job to be done]"). Freeze it. Comparable months require identical prompts; edit the panel quarterly at most, append rather than replace, and version it (v1, v2) so trend charts never silently mix panels.

Sizing and cadence: 50-150 prompts is the realistic working range, scaled to how broad the topic is, not bigger for its own sake. Replay monthly by default; move to weekly only for fast-moving or hotly contested topics where a month is too coarse. Daily replay rarely earns its cost in money and noise (field heuristic from 115+ agency audits).

Step 2: Replay monthly under controlled conditions

  • Same week each month (for example the first business week).
  • Fresh sessions: logged out or temporary chat, memory and personalization off, no custom instructions. Why: personalized answers measure your history, not the market's visibility.
  • Note the country and language and keep them constant; answers differ materially by locale.
  • For a multilingual market, build a separate panel per language AND culture, not a translation of one panel. The reference sources an engine draws on differ between cultural groups even inside one country (Flemish versus French-speaking Belgium, for example), so run a distinct setup and a distinct report per locale, and set the engine region by country and city where the interface allows it (field heuristic from 115+ agency audits).
  • Engines: ChatGPT (search enabled), Perplexity, Google AI Overviews (run the prompt as a Google query and record the AI Overview if one appears), Claude, Gemini. Keep the engine set constant.

Step 3: Record one row per prompt and engine

ColumnValues
dateYYYY-MM
promptverbatim text
enginechatgpt, perplexity, google-aio, claude, gemini
brand_mentionedyes or no
cited_with_linkyes or no
position1st, 2nd-3rd, listed, absent
sentimentpositive, neutral, negative
competitors_mentionedcomma-separated names
notesfree text, screenshot reference

Keep screenshots for wins and losses: clients believe screenshots, and they settle disputes about what an answer said.

Step 4: Compute the metrics

MetricFormulaRead it as
Mention rateprompts with brand_mentioned yes / total prompts, per engine and overallPresence in answers
Citation rateprompts with cited_with_link yes / total promptsPresence as a clickable source
Share of voicebrand mentions / (brand mentions + tracked competitor mentions)Competitive position
Trendcurrent month minus previous month, per metric per engineDirection

Alert threshold: investigate any relative drop greater than 20% on mention rate or share of voice, but confirm with a re-run before acting (next step explains why).

Worked example, panel of 50 prompts, ChatGPT, one month:

  • brand_mentioned yes on 9 prompts: mention rate 9/50 = 18%
  • cited_with_link yes on 4 prompts: citation rate 4/50 = 8%
  • Mentions across the answers: brand 9, competitor A 21, competitor B 13: share of voice 9 / (9 + 21 + 13) = 20.9%
  • Next month the mention rate reads 14% (7/50): that is a 22% relative drop, above the alert threshold. Re-run the panel; if confirmed, list the lost prompts and hand them to geo-visibility as rewrite targets.

Step 5: Respect non-determinism

AI answers are non-deterministic: the same prompt can name different brands across two consecutive runs. One run of one prompt is noise. Practical discipline:

  • Run each prompt 2-3 times per engine when feasible and score the majority outcome, or accept that single-run numbers only become meaningful as multi-month trends across the whole panel.
  • Never report a single-prompt, single-run change as a result. Aggregate first (panel-level rates), trend second (2-3 months), conclude third.
  • Budget honestly: 50 prompts across 3 engines at one run each is roughly half a day of manual work per month (field practice). Start with 50 prompts and 3 engines rather than skipping months on an oversized panel.

Layer 3: Server logs (the leading indicator)

Why logs: referrer analytics only see clicks, but assistants read pages while composing answers. A user-triggered fetcher hit means a human asked something and the assistant pulled your page into the conversation, whether or not anyone clicked. Log activity rises before citations and traffic do, which makes it the earliest signal a GEO effort is working.

Step 1: Identify the agents

User agentOperatorA hit means
GPTBotOpenAITraining data crawl
OAI-SearchBotOpenAICrawl for the ChatGPT search index (eligibility to appear)
ChatGPT-UserOpenAIA live user conversation fetched the page right now
ClaudeBotAnthropicTraining data crawl
Claude-UserAnthropicA live user conversation fetched the page right now
PerplexityBotPerplexityIndex crawl
Perplexity-UserPerplexityA live user conversation fetched the page right now

Filter the access logs for these strings (a plain text search is enough; no special tooling required). The full crawler table, robots.txt implications, and rendering notes live in the seo-technical skill, references/ai-crawlers.md: that file is the repo's canonical crawler reference.

Step 2: Verify before trusting (anti-spoofing)

User agent strings are trivially forged by scrapers. Verify that hit IPs fall inside the official published ranges before counting them:

Count only verified hits in reports; unverified hits are noise at best and scraper traffic at worst.

Step 3: Read the signals

  • Rising OAI-SearchBot coverage: the site is entering or refreshing in the ChatGPT search index.
  • ChatGPT-User, Claude-User, Perplexity-User hits: pages are being used in live answers. Track which URLs receive them; those are the AI-favored pages, and their structure (answer-first blocks, tables, definitional sentences) is what to replicate via geo-visibility.
  • Zero verified AI hits over a month, while competitors are visible in answers: check robots.txt and server-side blocking with seo-technical before blaming content.

Layer 4 (optional): paid platforms

Adopt only after the DIY system runs, and keep the DIY panel as the control.

ToolWhat it adds over DIY
Semrush AI Visibility IndexLarge shared prompt corpus, competitive benchmarks
ProfoundEnterprise answer monitoring, volume across regions
Ahrefs Brand RadarAI mentions alongside classic SEO data
OtterlyLightweight prompt tracking automation

Why DIY first: it is free, the methodology is transparent, the prompt set matches your actual buyers (vendor indexes use their own corpora), and the numbers stay comparable over time even if you change vendors later.

Measurement cadence

FrequencyTask
MonthlyReplay the prompt panel, compute rates, produce the report
MonthlyReview GA4 AI channel sessions, conversions, landing pages
MonthlyCount verified AI bot hits per agent from logs
QuarterlyReview the GA4 regex for new assistant domains
QuarterlyRevise the prompt panel (versioned, append-first)
On alertRe-run the panel to confirm any over 20% relative drop

Reading the signals together

The three layers diagnose each other. Use this table before drawing conclusions from any single number.

SymptomLikely meaningAction
Panel shows mentions, GA4 shows near zero AI sessionsAssistant apps send no referrer; sessions hide in DirectKeep reporting both; never conclude from GA4 alone
High mention rate, low citation rateBrand known to models, pages not retrieved or not quotablePassage citability work in geo-visibility
Verified fetcher hits but no mentions in answersRetrieval without selection: fetched, not quotedAnswer-first rewrite of the fetched URLs
Zero verified AI bot hits over a monthBlocking, rendering, or indexation problemrobots.txt and server HTML checks in seo-technical
Mention rate swings wildly month to monthPanel too small or sessions not clean50+ prompts, fresh sessions, 2-3 runs per prompt
GSC clicks drop on stable rankingsPossible AI Overview substitutionVerify with the panel on the affected queries

Rules and thresholds

RuleValueWhy
GA4 AI channel positionAbove ReferralTop-down evaluation; below Referral the channel captures nothing
RetroactivityNone, default or customGA4 channel groups classify forward only; baseline starts at creation
GA4 readingFloor, not totalAssistant apps send no referrer; the rest lands in Direct
Panel size50-100 prompts, frozenBelow 50 the rates are noise; unfrozen panels break trends
Panel revisionQuarterly at most, versioned, append-firstComparability is the entire value of the panel
Session hygieneMemory off, logged out, fixed locale, same weekPersonalization and locale shifts contaminate the measurement
Runs2-3 per prompt when feasible, majority outcomeSingle runs of a non-deterministic system prove nothing
Alert thresholdDrop over 20% on mention rate or share of voiceBelow that, normal variance; above, investigate and re-run
Bot hitsCount only IP-verified hitsUser agents are trivially spoofed
ReportingConversions next to sessionsAI visits convert about 4.4x better; volume alone undersells

Output format

Produce the monthly GEO report in this structure:

# GEO Report, [Month YYYY], [site]

## 1. AI traffic (GA4, floor values)
- AI assistant sessions: N (previous month: N, delta %)
- AI conversions: N, conversion rate vs organic search
- Top 5 landing pages from AI assistants
- Note: app-based assistant visits land in Direct; these figures are a floor.

## 2. Answer visibility (prompt panel vN, N prompts, locale)
- Mention rate per engine: table with previous month and delta
- Citation rate per engine: same table format
- Share of voice vs [competitor list]: overall and per engine
- Wins: prompts where the brand appeared this month and not last month
- Losses: the reverse, each with a screenshot reference

## 3. Retrieval signals (server logs, IP-verified only)
- Verified hits per agent (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Perplexity-User), with trend
- Top 10 URLs fetched by user-triggered fetchers

## 4. Actions for next month
- Pages to rewrite or create, from losses and gaps (hand to geo-visibility)
- Entity or third-party surface fixes (hand to geo-visibility)
- Technical blockers found (hand to seo-technical)
- Panel notes: prompts to add at the next quarterly revision (do not edit mid-quarter)

Keep every monthly report in the same format: the comparability of the document is part of the measurement.

Common mistakes

MistakeConsequenceFix
Comparing periods before and after channel group creationFake growth storyChannel groups are not retroactive; baseline starts at creation, annotate it
Leaving the AI channel below ReferralEmpty channel, false "no AI traffic" conclusionOrder the channel above Referral
Trusting the default AI Assistant group as completeMisses Perplexity and Copilot entirelyCustom group with the full regex
Presenting GA4 AI numbers as the totalSystematic undercount presented as truthState the floor caveat in every report
Concluding from one run of one promptActing on noiseAggregate rates, 2-3 runs or multi-month trends
Editing prompts every monthTrend lines compare different thingsFreeze the panel, version it, revise quarterly
Testing logged in with memory onMeasures personalization, not visibilityFresh sessions, memory off, fixed locale
Counting bot hits by user agent stringScrapers counted as AI visibilityVerify IPs against the official JSON ranges
Reporting sessions without conversionsAI channel dismissed as too smallReport the conversion multiple next to volume
Buying a platform before any DIY baselineNo control group, vendor-locked numbersDIY first, platform later as a layer

Cross-references

  • seo-keyword-research: builds the 50-100 buyer prompt panel this skill measures (do this first).
  • geo-visibility: turns measured gaps into citability fixes, entity work, and third-party surface plans (canonical citability rules).
  • seo-technical: canonical AI crawler reference (references/ai-crawlers.md), robots.txt and rendering checks when retrieval signals are absent.
  • seo-geo-audit: the bundled audit script (scripts/seo_audit.py) scores pages flagged by this skill's reports.
  • seo-backlinks: brand mentions workflow when share of voice lags despite good content.

Sources

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.