Key takeaways
- The same prompt run 100 times through ChatGPT or Google AI has less than a 1 in 100 chance of returning the same list of brands twice, according to SparkToro's 2026 consistency study, so a single dashboard snapshot tells you almost nothing on its own.
- Roughly 61.7% of AI citations are "ghost citations": your page gets linked as a source, but your brand name never appears in the visible answer. Semrush and Growth Memo both landed on nearly identical figures independently.
- Vendors collect data two fundamentally different ways, API polling or UI scraping, and the two methods can return different citation counts for the identical prompt run at the identical minute.
- Citation behavior itself moves under your feet. Promptwatch's tracking of the GPT-5.3 rollout showed average citations per ChatGPT response drop about 27% overnight across every model variant, with no recovery a month later.
- You can't outsource trust entirely to a vendor's dashboard. Run your own manual spot-checks monthly and ask pointed questions about collection methodology before you sign a contract.
Why this question even needs asking
Somewhere around $100 million a year is now being spent on AI visibility tracking, per Search Engine Land's estimate cited in SparkToro's research. That's a real budget line for a category of software that, until recently, nobody had bothered to test for basic consistency. Marketers bought dashboards that show "brand mention rate: 42%" and treated the number the way they'd treat a Google Search Console impression count: as ground truth.
It isn't ground truth. It's a sample from a system that is, by design, probabilistic. I don't say that to be alarmist. I say it because once you understand the mechanics, you stop panicking over a week-over-week dip that's actually just noise, and you start asking the right questions of any tool you're paying for.
The core problem: AI answers aren't deterministic
Large language models generate text through token sampling, not lookup. Even at a temperature setting of zero, ties in token probability mean you won't reliably get identical output twice, according to discussion on OpenAI's own developer community. Add live retrieval (Perplexity searches the web on nearly every response), personalization, chat memory, and geographic context, and you get a system where the same question produces meaningfully different answers depending on who asks, when, and from where.
SparkToro and Gumshoe.ai put this to an actual test in 2026: 600 volunteers ran 12 prompts through ChatGPT, Claude, and Google AI a combined 2,961 times. The finding that should worry anyone relying on a single-prompt brand check: there's less than a 1 in 100 chance that ChatGPT or Google's AI returns the same list of brands in any two runs of an identical prompt, and closer to 1 in 1,000 that two lists appear in the same order. The number of brands listed also varies run to run, sometimes two or three names, sometimes ten.
That's not a bug in any particular tool. It's the nature of the thing being measured. Visiblie's research corroborates the practical effect: a brand can show a 40% mention rate one week and 25% the next with zero content changes, purely from sampling drift. If your monitoring tool doesn't disclose sample sizes or run multiple queries per prompt per check, you're staring at noise dressed up as a trend line.
The ghost citation problem: citations aren't mentions
Here's the accuracy gap that surprised me most in the research. Kevin Indig coined the term "ghost citation" for when a domain gets linked as a source in an AI answer, but the brand's name never shows up in the text the user actually reads. Growth Memo's data puts the ghost citation rate at 61.7%. Semrush ran its own independent analysis and landed on almost the identical number, 62%. Writesonic's study of roughly 16 million brand appearances found about 40% of citations didn't name the source brand, and Perplexity dropped the brand name in 52% of its cited appearances specifically.
Aggregator sites take the biggest hit here. Semrush's dataset showed medium.com cited sixteen times and named zero times. If your monitoring platform reports "citations" and "mentions" as the same metric, it's overstating your real visibility by something like 40 to 60%. Ask any vendor directly: does your dashboard separate "we linked to your page" from "we said your brand's name"? If they can't answer clearly, that's the answer.
API vs. UI scraping: the methodology gap nobody talks about
This is the single biggest hidden difference between competing AI visibility tools, and almost none of them lead with it in their marketing. There are two ways to collect AI answer data: hit the model's API directly, or scrape the rendered user interface the way a real person sees it.
API responses often strip out citations, source URLs, and formatting that appear in the actual ChatGPT or Perplexity web interface. Superlines documented a concrete case: identical prompts returned zero citations via the API while the web UI showed five inline citations with source links for the same query. That's not a rounding error. That's a tool potentially telling you "no citations found" for a query where your brand is, in fact, cited in what real users see.
Source-metadata completeness also varies by platform even among tools using the same collection method. Martin Kulawik's technical breakdown found OpenAI at 99.7% metadata completeness and Perplexity at 99.5 to 100%, but Gemini only hit 84.9% until a stricter inclusion rule was applied.
There's also a live vendor dispute worth knowing about, even though it's one-sided: Peec AI's comparison page alleges Profound may inject geographic identifiers directly into prompts to fake localized results in markets where its proxy infrastructure is thin, rather than running real geo-distributed sessions. Take that with appropriate skepticism since it's a competitor making the claim, but it illustrates exactly the kind of question you should be asking any vendor about how they actually generate "location-based" data.

Questions to ask every vendor before you buy
- For each AI platform you track, do you use API polling or UI scraping, and does that choice vary by platform?
- Do you separate citation counts from brand-mention counts, or report them together?
- How many times do you sample a given prompt before reporting a mention rate, and over what window?
- How do you verify crawler traffic isn't spoofed? Ask whether they match source IPs against published operator ranges, the way Promptwatch's crawler-log methodology does.
- Can I export raw response text, not just aggregated scores, so I can spot-check manually?
Citation behavior itself keeps shifting
Even a perfectly honest, methodologically sound tool is measuring a moving target. Promptwatch's tracking around the GPT-5.3 rollout on March 4, 2026 found average citations per ChatGPT response fell from about 6.4 the week before to 4.7-4.9 by late March, a roughly 27% drop across every model variant simultaneously, with no recovery a month later. That's a platform-level change that had nothing to do with any brand's content and everything to do with OpenAI adjusting how many sources it shows.
Similarly, Promptwatch's data on average sources per response shows Perplexity holding almost exactly ten sources per answer day after day, while Microsoft Copilot has swung from under two to nearly seventeen sources per response within a matter of weeks. If a monitoring dashboard shows your citation count dropping and you haven't cross-referenced whether the whole platform's citation behavior shifted that week, you'll waste time chasing a problem that doesn't exist.
Query fanouts compound this. ChatGPT doesn't run a single search per prompt, it fans out into multiple queries, and Promptwatch's fanout data shows the average character length of those fanout queries collapsed from roughly 117 characters in December to about 53 by April, meaning ChatGPT's actual searches are getting far more terse and keyword-like. A tool that only tests the exact sentence a user might type, without replicating fanout behavior, isn't capturing what's actually happening behind a single AI response.
Sentiment scores deserve the same skepticism
Academic benchmarks report 85 to 94% accuracy for LLM-based sentiment classification, but those numbers come from clean test datasets. ZipTie's caution here is worth repeating: production environments are noisy, and if a vendor claims 99% sentiment accuracy, treat that as a red flag rather than a selling point.
Visiblie's proprietary dataset across 200+ brands found something that reframes how "mention rate" should be read entirely: the average brand receives outright endorsement in only 28% of category prompts where it appears, with 41% neutral, 19% cautious, and 12% actually hallucinated. A brand that shows up in 8 of 10 prompts, an apparently strong 80% mention rate, might have an effective positive positioning rate closer to 20%. Raw visibility and favorable visibility are not the same metric, and a lot of dashboards blur them.
AI answers get facts about brands wrong, too
Separate from consistency and citation accuracy, there's the plain question of whether AI answers are factually correct about your brand at all. Oumi's 2026 study for the New York Times, using OpenAI's SimpleQA dataset and an LLM-as-judge pipeline, found Google AI Overviews correct about 91% of the time post-Gemini 3, but only 39% of overviews were both correct and fully backed by their cited sources. At the individual claim level, only 67% of claims were actually supported by the sources cited alongside them. Roughly a third of claims are unsupported even in an answer that's technically "correct" overall.
That's the mechanism behind the 40%+ of users who, per Exploding Topics data cited by Meltwater, report encountering inaccurate or misleading content in AI Overviews. If outdated pricing or a discontinued product feature made it into a model's training data or a frequently cited source, users accept it as fact, and no traditional alert system flags it.
A practical verification workflow
You don't need to trust any single source, including a vendor dashboard. Here's a workflow that mixes manual checking with tooling.
- Run a manual baseline. Pick 20-30 prompts that mirror what your actual buyers type, and run each one 3-5 times across ChatGPT, Perplexity, and Google AI Overviews. Record brand mentioned yes/no, position in the answer, sentiment, and whether the description is accurate.
- Cross-check your monitoring tool's numbers against that manual baseline monthly. If the gap is large and doesn't shrink, something in the tool's methodology, not your brand's actual visibility, is the variable.
- Separate citations from mentions in every report you read. If your tool reports one blended number, ask for the split, or assume the real mention rate is meaningfully lower than what's shown.
- Watch for platform-wide shifts before reacting to your own trend line. A citation count drop that coincides with a known model update (like GPT-5.3) is a platform story, not a content story.
- Don't take a tool's "top competitors" or "top sources" tab at face value. Cross-reference against brands that actually rank for the same commercial-intent queries in traditional search.
Comparing how leading tools handle accuracy
| Tool | Collection method (per vendor claims) | Ghost-citation vs. mention split | Sample depth | Notes |
|---|---|---|---|---|
| Promptwatch | UI-based monitoring across 12+ AI platforms, plus verified crawler logs | Tracks citations, offsite mentions, and crawler-to-citation path separately | Prompt volumes with difficulty and citation-rate scoring per prompt | Also runs Agent Analytics on real crawler IPs, not just self-reported user agents |
| Profound | Mixed API/UI, disputed by a competitor on geo-accuracy | Not clearly disclosed publicly | Starter tier limited to ChatGPT only, 50 prompts/day | Growth tier adds Perplexity and AI Overviews |
| Peec AI | Dedicated UI-scraping infrastructure in 80+ countries per vendor | Not clearly disclosed publicly | Starter covers 3 engines, unlimited regions and seats | Positions itself against Profound's alleged prompt injection |
| Otterly.AI | API-based, lower price point | Basic citation tracking, no explicit mention-split reporting | 15-400 prompts depending on tier | No crawler log or content generation layer |
| Semrush AI Visibility | Bundled with existing SEO data pipeline | Not the primary focus of the toolkit | Domain-level, tied to Semrush's broader keyword index | Best if you're already deep in the Semrush ecosystem |
Of the tools above, Promptwatch is the one built specifically around the idea that monitoring alone isn't the point, understanding why a number moved matters just as much. Its Agent Analytics logs show which AI crawlers actually visited your pages, whether they hit errors, and the crawl-to-citation path per page, which gives you a way to sanity-check a citation-rate change against real server-side evidence rather than trusting a single dashboard metric in isolation.

What this means for your monitoring setup
None of this is a reason to abandon AI visibility tracking. It's a reason to treat any single number from any single tool as directional, not absolute, the same way good analysts already treat individual keyword rankings. Use multiple sampling runs, separate citation from mention data explicitly, and keep a recurring manual spot-check on the calendar so you have an independent reference point that no vendor's methodology choices can quietly distort.
If you want to browse other options across this crowded category, from monitoring-only trackers to fuller GEO platforms, the directory at bestgeosoftware.com is a reasonable place to compare feature sets side by side before you commit budget to any one tool.
The honest summary is this: AI brand monitoring data is real, but it's a sample of a noisy, shifting system, not a clean readout. Treat it that way and you'll make better decisions with it.