AI Search Tracking: How to Measure AI Visibility (2026)

TL;DR
AI search tracking means running a fixed set of buyer prompts through ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews on a schedule and recording how often your brand shows up. A single check tells you very little, because AI answers change on almost every run: SparkToro found the same brand list comes back less than 1% of the time, and Ahrefs found AI Overviews have a 70% chance of changing between observations. Use the Prompt Panel Method instead: pick 40 buying prompts, run each one 5 times per engine every month, and score every answer on four yes-or-no checks (mentioned, cited, recommended, accurate). Report those as rates per engine, never as a rank position, and only act when a rate moves 10 points or more two months in a row. Add AI referral traffic from GA4 and AI crawler hits from your server logs. Before you pay for a tool, ask whether it scrapes the real interface or calls the API, because Surfer found the brands in API answers overlapped with the real interface only 15.5% to 23.8% of the time.
Free Playbook
The exact 90-day playbook we run for clients
AI search tracking measures how often ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews name or cite your brand for the prompts your buyers ask. Because AI answers change on almost every run, you track rates across repeated runs of a fixed prompt set rather than a single position.
One check proves very little. SAASY LINKS ran 40 buying prompts five times each in ChatGPT for 12 SaaS clients in 2026, and when a client appeared at all, it showed up in all five runs only 31% of the time. A screenshot from Monday can be wrong by Tuesday.
If you're still working on getting cited in the first place, start with our AI SEO guide. This post is for the next question: is that work showing up in the answers, and how would you know?
What Is AI Search Tracking?
AI search tracking is the practice of running a fixed set of buyer prompts through AI assistants on a schedule and logging whether your brand is named, linked or recommended. It does the job rank tracking does for Google. The difference is the output: a rate across many runs instead of a position from 1 to 100.
Rank tracking vs AI search tracking
| Google rank tracking | AI search tracking | |
|---|---|---|
| Unit you track | A keyword | A prompt, often 8 to 20 words |
| Output | A position from 1 to 100 | Mention, citation and recommendation rates |
| How stable it is | Usually the same for days | AI Overviews change 70% of the time between checks |
| What skews it | Location and device | Location, chat memory, reasoning mode and the model version |
| How data is collected | Scraping the results page | Calling the model's API or scraping the chat interface |
Mentions and citations are two separate things, so measure them separately. Kevin Indig's analysis of more than 600,000 ChatGPT citations across 1,094 US categories found the correlation between citations and brand mentions was -0.229. The most-cited domain was also the most-mentioned brand in only 20.8% of cases.
In practice, ChatGPT can link to your pricing page while recommending a competitor, or name you without linking anywhere. Both happen every day, and a tool that blends them into one visibility score hides which one you have.
Why Your Rank Tracker Misses AI Answers
Rank trackers assume a result sits still long enough to measure. AI answers don't. Ahrefs found AI Overviews have a 70% chance of changing from one observation to the next, and SparkToro found ChatGPT, Claude and Google's AI returned the same brand list less than 1% of the time for the same prompt.
The SparkToro study, run by Rand Fishkin with Patrick O'Donnell of Gumshoe, tested 2,961 prompts with 60 to 100 runs per platform. The same brands in the same order came back less than 0.1% of the time. Fishkin's verdict on tools that report an AI rank was that any such tool "is full of baloney."
Louise Linehan's Ahrefs data adds a second problem. Across 43,000 keywords watched for a month, AI Overview content lasted 2.15 days on average, and only 54.5% of cited URLs carried over between consecutive answers.
Engines also disagree with each other. Despina Gavoyannis at Ahrefs compared 540,000 query pairs and found AI Mode and AI Overviews cited the same URLs only 13.7% of the time, even though both are Google. Kevin Indig's H1 2026 report found 91% of citations appear in only one of ChatGPT, Perplexity or AI Overviews.
Keep your rank tracker, though. The OppAlerts study of about 167,000 domains found your best Google ranking correlates at +0.148 with AI visibility, so Google positions still work as a leading indicator. They just can't stand in for the answers themselves.
The Prompt Panel Method
The Prompt Panel Method is how we run AI search tracking for SaaS clients. Pick 40 buying prompts and run each one 5 times per engine every month. Score every answer on four yes-or-no checks: mentioned, cited, recommended and accurate. Report the results as rates, and only act on moves of 10 points or more.

How to score each answer in the Prompt Panel Method
| Check | Counts as yes when | Rate you report |
|---|---|---|
| Mentioned | Your brand name appears anywhere in the answer | Mention rate |
| Cited | A URL on your domain appears in the answer's sources | Citation rate |
| Recommended | You're named as the top pick or in the first two options | Recommendation rate |
| Accurate | Pricing, features and category match your site today | Accuracy rate |
The numbers are deliberate. Forty prompts times five runs times five engines gives 1,000 answers a month, or 200 per engine. At that size, a 10-point swing in one engine's mention rate is roughly twice the random wobble you'd expect from sampling alone, so it usually means something changed.
Hold the panel steady. Swap no more than 10% of prompts a quarter, or you'll be comparing different questions month to month and calling it progress. And log the date, engine, country and model for every run, because a model update can move every rate at once.
What a monthly Prompt Panel report looks like
Keep the report to one table per month: one row per engine, one column per rate, plus the change since last month. The example below is illustrative, but it shows the pattern we look for, where one engine moves and the others hold.
Example Prompt Panel report (illustrative): 40 prompts, 5 runs each, 200 answers per engine
| Engine | Mention rate | Citation rate | Recommendation rate | Accuracy rate |
|---|---|---|---|---|
| ChatGPT | 34% → 46% | 12% → 15% | 9% → 17% | 88% |
| Perplexity | 41% → 43% | 29% → 31% | 14% → 15% | 92% |
| Gemini | 22% → 24% | 8% → 7% | 5% → 6% | 81% |
| Claude | 18% → 19% | 3% → 3% | 4% → 4% | 76% |
| Google AI Overviews | 27% → 25% | 19% → 18% | not scored | 90% |
Read it row by row. ChatGPT's mention rate jumped 12 points, which clears the bar once. If it holds next month, check which new sources it cited. Every other move is inside the noise. Claude's 76% accuracy rate is the other flag: a quarter of its answers get something wrong about the product.
Which AI Search Metrics Matter in 2026?
Four rates from your prompt panel carry most of the signal: mention rate, citation rate, recommendation rate and accuracy rate. Add share of voice against named competitors, then two numbers from your own site: AI referral sessions in analytics and AI crawler hits in your server logs.
AI search metrics, what each one shows and where to get it
| Metric | What it shows | Where it comes from |
|---|---|---|
| Mention rate | How often AI names you for buying prompts | Your prompt panel |
| Citation rate | How often AI links to your pages as a source | Your prompt panel |
| Recommendation rate | How often you're the pick, not just on the list | Your prompt panel |
| Accuracy rate | Whether AI describes your pricing and features correctly | Your prompt panel, checked by hand |
| Share of voice | Your mentions ÷ all tracked brands' mentions | Your prompt panel |
| AI referral sessions | Visits and signups arriving from AI assistants | GA4 or your analytics tool |
| AI crawler hits | Whether OAI-SearchBot and PerplexityBot can reach your pages | Server logs or CDN logs |
Referral traffic is small but worth watching closely. Patrick Stox reported that AI search was 0.5% of Ahrefs' traffic but 12.1% of its signups, which is why the conversion number matters more than the session count.
Accuracy gets skipped most often. A 60% mention rate isn't much use if half of those answers quote a price you retired last year, so read a sample of 20 answers by hand each month and log what's wrong.
How to Set Up AI Search Tracking
Setting up AI search tracking takes about a day. Write the prompt panel, pick the engines your buyers use, fix the conditions every run happens under, then connect the two first-party sources: AI referrals in GA4 and AI crawler hits in your logs. After that it's a monthly routine.
Build your prompt panel
Write prompts the way buyers type them into ChatGPT, not the way they type keywords into Google. "Best CRM for a 10-person accounting firm" beats "best CRM". Split the 40 prompts across category prompts, comparison prompts like "HubSpot vs Pipedrive for agencies", problem prompts and brand prompts such as "is [your brand] good for startups".
Weight the panel toward money. We put about 60% of prompts in category and comparison questions, because that's where a recommendation turns into a demo. Pull wording from sales call notes, support tickets, G2 reviews and Search Console queries of eight words or more, since long queries there are often people typing AI-style questions into Google.
Choose the engines your buyers use
Track ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews at minimum, and add AI Mode if your buyers search on Google. The mix is shifting fast. Kevin Indig's H1 2026 report found ChatGPT's share of AI assistant use fell from 78% to 56% in a year while Gemini's rose from 15% to 30%.
Control the run conditions
Run every prompt in a fresh session, logged out or with memory off, from the same country. Record the reasoning mode too. Semrush and Kevin Indig found that only 25.6% of cited domains overlapped between minimal and high reasoning in ChatGPT on the same prompt.
That one setting can swing your citation rate more than a month of content work. So pick a mode per engine, write it down, and don't change it mid-quarter.
Track AI referral traffic in GA4
In GA4, build an exploration filtered by session source matching this regex: chatgpt\.com|chat\.openai\.com|perplexity\.ai|gemini\.google\.com|claude\.ai|copilot\.microsoft\.com. OpenAI's publisher FAQ confirms ChatGPT adds utm_source=chatgpt.com to the links it shows, so most ChatGPT clicks are labeled.
Google's AI traffic is harder to see. Google's documentation says clicks and impressions from AI Overviews and AI Mode are counted inside the "Web" search type in Search Console, with no separate filter. In our experience, some AI app traffic also arrives with no referrer and lands in Direct.
Check AI crawler hits in your logs
Tracking answers is pointless if the engines can't read your pages. Search your server or CDN logs for OAI-SearchBot, PerplexityBot, ClaudeBot and Googlebot each month. OpenAI's crawler documentation says sites that block OAI-SearchBot won't be shown in ChatGPT search answers.
How to Run AI Search Tracking Without a Tool
You can run a small panel in a spreadsheet for free. It takes one person about four hours a month for 20 prompts across four engines, which is enough to learn what matters before you buy software. Past 100 answers a month, the copying and pasting stops being worth the time.
Spreadsheet columns for a manual prompt panel
| Column | What to record | Example |
|---|---|---|
| Date and engine | When and where the prompt ran | 2026-09-14, Perplexity |
| Prompt ID and run | Which prompt and which of the 5 runs | P07, run 3 |
| Mentioned / cited / recommended / accurate | Four yes-or-no cells | Y / N / N / Y |
| Brands named | Every brand in the answer, in order | Gusto, Rippling, Deel |
| Sources cited | Each URL the answer linked to | g2.com/categories/payroll |
| Settings | Logged out, country, reasoning mode | Logged out, US, default |
The "sources cited" column is the one people skip and the one you'll use most. After two months it shows which third-party pages each engine leans on for your category, and that list is worth more than the rates themselves.
AI Search Tracking Tools Compared (2026)
You can run the panel by hand in a spreadsheet, but past 100 answers a month a tool pays for itself. The differences that matter are which engines a tool covers, how often it runs each prompt, whether it scrapes the real chat interface or calls the API, and whether you track your own prompts or a vendor's database.
AI search tracking tools compared, September 2026
| Tool | Engines covered | How it collects answers | Starting price |
|---|---|---|---|
| Otterly.AI | ChatGPT, Perplexity, AI Overviews, Copilot (Claude, Gemini, AI Mode as add-ons) | Your prompts, daily runs | $29/month for 15 prompts |
| Surfer AI Tracker | ChatGPT, Perplexity, Gemini, AI Overviews, AI Mode | Scrapes the real interface | $82/month billed yearly |
| Ahrefs Brand Radar | ChatGPT, Perplexity, Gemini, AI Overviews, Copilot, Claude | Your prompts, or an index of 448M prompts | $50/month custom prompts, $199/month index |
| Mangools AI Search Watcher | ChatGPT, AI Overviews, AI Mode, Gemini, Claude, Grok, Mistral, Llama | Runs each prompt several times | Free start, paid plans |
| Profound | Nine engines incl. ChatGPT, Perplexity, Gemini, Claude, AI Mode | Your prompts, daily runs | Custom, free 7-day trial |
| GA4 + Search Console | Referral traffic only | Your own site data | Free |
Pricing and engine lists came from each vendor's site in September 2026 and change often. For hands-on reviews of 12 platforms, see our roundup of LLM SEO tools.
For most SaaS teams under 50 prompts, Otterly's entry plan or Ahrefs Brand Radar's custom prompts cover the basics. Brand Radar makes more sense if you already pay for Ahrefs, since the data sits next to your backlink and keyword reports. Enterprise teams tracking 500+ prompts across regions tend to end up on Profound.
How to Tell If a Tracker's Numbers Are Real
Ask every vendor two questions before you pay. Does it collect answers from the real chat interface or from the model's API? And how many times does it run each prompt before reporting a number? A tool that can't answer both clearly is selling a dashboard, not data.
The API question matters more than most buyers think. Surfer's September 2026 study ran 1,000 prompts and collected 13,779 answers across ChatGPT, Perplexity, Gemini, AI Mode and AI Overviews. Brands named in API answers overlapped with the real interface only 15.5% to 23.8% of the time, and ChatGPT's page-level source overlap was as low as 4.8%.
“Before you spend a dime tracking AI visibility, make sure your provider answers the questions we've surfaced here and shows their math.”
Our view: one screenshot of ChatGPT naming your brand is an anecdote, not a metric. Any AI visibility report that doesn't say how many times each prompt was run and under which settings shouldn't go in a board deck.
Ask for the raw answers too. A good tool lets you open the exact response behind every data point and see the sources it cited. If you can only see a score, you can't check it, and you can't act on it either.
What Doesn't Work in AI Search Tracking
Checking ChatGPT by hand once a month is the most common method and the least reliable. Your own account's memory, location and past chats shape the answer, and one run of a prompt that changes almost every time tells you nothing about the other 99.
Reporting an "AI rank position" fails for the same reason. When the same list in the same order comes back less than 0.1% of the time, "you're #2 in ChatGPT" describes one roll of the dice. Report the share of runs where you're recommended instead.
Daily alerts on single-run changes are the third trap. They fire constantly, teams learn to ignore them, and the one real drop gets missed. Monthly rates with a 10-point threshold catch less noise and more signal.
Turning AI Search Tracking Data Into Citations
AI search tracking only pays off when it tells you which pages to get onto. For every prompt where a competitor is recommended and you aren't, open the answer and list the sources it cited. Those listicles, review pages, forum threads and YouTube videos are your outreach list, ranked by how often each one is cited.
That's how the method earns its keep. A B2B HR software client tracked 60 prompts with SAASY LINKS and found one competitor's lead traced back to six pages ChatGPT cited again and again. After the client was added to five of those pages, its ChatGPT mention rate across the panel went 12% → 38% in 10 weeks.
The fixes differ by engine. Getting into ChatGPT answers leans on third-party mentions, which we cover in how to rank in ChatGPT. Google's answers lean harder on your own rankings, so see our guide on how to rank in AI Overviews for that side.
If you'd rather have someone run the panel and chase the cited pages for you, our AI SEO services include monthly prompt tracking across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews.
Frequently Asked Questions
Can AI searches be tracked?
Yes, but only as rates, not positions. Run a fixed set of prompts several times per engine on a schedule and record how often your brand is named or cited, because single answers change on almost every run.
How do you track AI search visibility?
Build a panel of about 40 buying prompts, run each one 5 times per engine every month and score each answer for mentions, citations, recommendations and accuracy. Add AI referral traffic from GA4 and crawler hits from your server logs.
What is the best AI search tracking tool?
It depends on scale and stack. Otterly.AI starts at $29 a month for 15 prompts, Ahrefs Brand Radar suits teams already on Ahrefs, and Profound fits enterprise panels, but check whether any tool scrapes the real interface or only calls the API.
How much do AI search tracking tools cost?
Entry plans run from about $29 to $82 a month for 15 to 50 prompts as of September 2026. Enterprise platforms such as Profound use custom pricing.
How often should you run AI search tracking?
Monthly is enough for decisions, with each prompt run at least 5 times per engine. Daily single runs mostly measure noise, since AI Overviews change 70% of the time between checks.
Does Google Search Console show AI Overview traffic?
Partly. Google counts AI Overviews and AI Mode clicks and impressions inside the "Web" search type, but you can't filter them out separately, so pair Search Console with a prompt tracker.
What's the difference between AI search tracking and rank tracking?
Rank tracking reports a keyword's position on Google, which stays fairly stable. AI search tracking reports how often an AI assistant names or cites you across repeated runs of a prompt, because the answer itself keeps changing.

Written by
Co-founder & CEO of SAASY LINKS, the B2B SaaS link building and AI visibility agency. 10+ years in SEO and growth marketing for SaaS brands. Mentor at 500 Startups and Techstars. Runs the Backlink Masterminds community for link builders.
Get Your SaaS Ranked on Google & Cited by AI
Senior-led AI SEO combining strategy, content, editorial links, and AI optimization in 90-day sprints. Built exclusively for B2B SaaS.
Book a Free Consultation CallBuilt exclusively for SaaS.



