Key takeaways
- ChatGPT recommends only 1.2% of brand locations tested, versus a 35.9% appearance rate in Google's local 3-pack for the same brands, per SOCi's 2026 Local Visibility Index. Local AI visibility numbers need far more scrutiny than classic local SEO rankings.
- A single prompt run proves almost nothing. Platform-wide shifts (like the 27% drop in ChatGPT citations after the GPT-5.3 rollout) can swing your "score" overnight for reasons that have nothing to do with your content.
- Personalization, IP geolocation, and account memory all contaminate location-based prompt testing. If you're not controlling for these, you're measuring noise.
- Crawler logs need IP verification, not user-agent matching. Spoofed AI crawler traffic is common enough that raw log counts are unreliable without checking source IPs against published ranges.
- Reddit, social proof, and third-party review sites carry wildly different weight depending on the AI engine you're testing, so "AI visibility" as a single blended number hides more than it reveals.
Every local business, multi-location brand, and agency running AI visibility reports in 2026 is dealing with the same problem: the data looks clean, but the methodology underneath it is often shaky. You open a dashboard, see a visibility score, and it feels authoritative. It isn't, not automatically. Before you present a number to a client or make a budget decision off it, run through these 10 checks. Some take five minutes. A couple take longer. All of them have burned someone I know.
1. Did you test from a clean, logged-out session?
This is the one almost everyone skips. ChatGPT's memory system is on by default for paid accounts and pulls "useful information" from past conversations even when you never asked it to remember anything. OpenAI's own memory FAQ confirms this, and as of the January 2026 release notes, the free ad-supported tier uses past chats for ad personalization too. That means a logged-in test account with chat history can return a completely different answer than a fresh session for the identical prompt.
If you're testing brand visibility on an account you use daily for other work, you're not testing AI visibility. You're testing your own search history reflected back at you. Use incognito sessions, or better, a dedicated test account that's never been used for anything else.
2. Are you controlling for location, or just assuming IP geolocation works?
City-level IP geolocation accuracy is meaningfully worse than country-level accuracy, and that gap matters a lot for local visibility testing. If you're trying to find out whether a Chicago location shows up in AI answers versus a Denver location, and you're relying on a VPN or proxy to simulate location, you should treat the results as directional at best, not a precise score.
The fix is to use explicit location parameters where the platform supports them, rather than inferring location from an IP address. Perplexity is notably more transparent than most about letting users control whether precise location is shared at all, which actually makes it one of the more reliable engines to test against for this exact reason.
3. Did you separate language from location?
Practitioners on r/aeo have reported the same brand, same question, answered differently depending on which country they asked from, and then differently again when they just switched the prompt's language, independent of location. Those are two separate variables. If your checklist conflates "testing from Germany" with "testing in German," you can't tell whether a bad result is a location problem or a language problem, and you'll fix the wrong thing.
4. Did you run each prompt more than once, on more than one day?
A single prompt run is a snapshot, not a trend. Model sampling introduces randomness even with identical inputs, and platform-wide changes can happen between Tuesday and Wednesday. The clearest example: around the GPT-5.3 rollout on March 4, 2026, Promptwatch's data on the ChatGPT citation drop showed average citations per response falling roughly 27% across every model variant, GPT-5.3, GPT-5.4, and GPT-5-Mini, in a single day, with no recovery a month later. That's a platform behavior change, not something any one site did wrong. If you'd run your audit the week before versus the week after, you'd have drawn opposite conclusions about your own performance.
At minimum, re-run each prompt three times across different days before you trust the result enough to act on it.
5. Are you tracking each AI engine separately, not blending them into one score?
This is the single most common mistake I see in local visibility reports. Different engines behave independently of each other, sometimes dramatically so. Reddit held a steady ~3.8% share of ChatGPT citations through late July and early August 2026, then collapsed to under 1% in a single day on August 14, an 86.4% relative drop. Over that same window, Google AI Overviews and AI Mode only declined gradually, no cliff at all. That's documented in Promptwatch's research on Reddit citations dropping in ChatGPT.
Similarly, social platforms carry wildly different weight by engine: ChatGPT cites social media (Reddit, YouTube, LinkedIn, Facebook, Instagram, TikTok, X) in only about 0.46% of responses, while AI Overviews pulls roughly 12.2% of its citations from social, according to Promptwatch's social media citation breakdown by AI model. If your "AI visibility" number averages ChatGPT and AI Overviews together, you've buried a real, actionable difference.

6. Does your citation slot count match what the engine actually gives out?
Sources vary a lot by platform, and it changes how much any single citation matters. ChatGPT hands out only about 5 citations per web-search-enabled response on average. Google AI Overviews cites roughly double that, about 10 sources. Perplexity is close to AI Overviews, also around 10, and notably the most stable of the engines tracked day over day. Microsoft Copilot is the wildcard: its average has swung from under 2 to nearly 17 sources per response within a few weeks, which tells you Microsoft is still re-architecting retrieval and attribution underneath the hood. These figures come from Promptwatch's average sources per response data.
The practical implication: getting one citation on ChatGPT is worth more, proportionally, than getting one citation on AI Overviews. Don't weight every "mention" equally across platforms in your scoring.
7. How many unique domains actually compete for your prompt type?
This is where local and brand-specific queries diverge sharply from generic ones. Across organic prompts, 52.5% of ChatGPT responses cite 10 or more different domains, meaning the competitive set is wide. But on brand-specific prompts, the kind that matter most for "is [brand] good" or "best [local category] near me" queries, only 21.3% cite 10+ domains, and 13.8% cite 3 or fewer, three times the rate seen on organic prompts. This comes from Promptwatch's data on unique domains per ChatGPT response.
For a local business, that's a small, high-stakes source set. Your Yelp profile, your Google reviews, and your own site each carry outsized influence when only three or four sources are in the mix. Audit those three or four, not the top 50 generic results.
8. Did you check domain authority bias, or assume bigger sites always win?
It's tempting to assume a local business can't compete for AI citations against massive review aggregators. The domain rank data doesn't support that assumption as strongly as you'd think. In August 2026, ChatGPT citation share by domain rank showed DR46-60 and DR61-75 sites together capturing about 46% of citations, the single biggest chunk, while DR91-100 sites, the biggest domains on the web, fell to just 3-4.4% by the end of the month, down from about 7% earlier. Even DR0-30 sites captured around 14%. That's from Promptwatch's citation share by domain rank report. You don't need a massive-authority domain to get cited. Mid-tier sites are where the real competition actually happens.
9. Are you verifying AI crawler traffic by IP, not by user-agent string?
If part of your audit includes checking server logs for AI crawler activity, be skeptical of raw hits. Spoofing is common: OpenAI's ChatGPT-User agent is spoofed at roughly a 1:5 ratio (one real request to five fakes), MistralAI-User spoofing runs closer to 1:37, and Perplexity-User spoofing is around 1:88. GreyNoise has documented threat actors actively forging six different AI crawler names across OpenAI, Anthropic, Google, and Perplexity. Counting user-agent strings without checking source IPs against the operators' published ranges (OpenAI, Perplexity, and Anthropic all publish these) will inflate your numbers, sometimes dramatically.
Separately, the mix itself shifts fast. OpenAI's share of verified AI crawler requests fell from 94.8% in early June 2026 to 79.8% by early September, per Promptwatch's AI crawler traffic data. A robots.txt rule or CDN block that seemed harmless a year ago can now be blocking a crawler that's grown into a much bigger share of your traffic. Compare your own logs against the published mix. If a provider is large in that mix but near zero in your logs, you're probably blocked or unreachable.
10. Do your pricing, reviews, and basic facts match across every platform AI actually pulls from?
Controlled testing on restaurant-related prompts found that Apple Maps and Bing Places were cited as text sources zero times across 210 usable ChatGPT answers, while review and business-card data in the UI came from a separate licensed feed rather than the cited sources themselves. That doesn't mean Apple's listing is irrelevant (it still feeds Maps, Siri, and Spotlight outside ChatGPT's reach), but it does mean you shouldn't assume your Apple or Bing listing is driving ChatGPT citations just because it's accurate. Because Bing Places data feeds Bing Maps, Windows Search, Edge, and Copilot, keeping that listing claimed and current arguably matters more for AI exposure than for classic Bing search share.
More broadly, AI engines skew toward higher-rated locations than Google's local pack tolerates. SOCi's 2026 Local Visibility Index found locations recommended by ChatGPT averaged 4.3 stars, versus 3.9 on Gemini and 4.1 on Perplexity. In traditional local search, a middling-rated business can still rank on proximity or category relevance. AI engines are far less forgiving. If your rating sits in the 3.5-4.0 range, don't expect strong AI recommendation rates even with perfect technical setup.
Putting it together: a quick-reference table
| Check | What it catches | Fix |
|---|---|---|
| Clean session testing | Personalization inflating or distorting results | Test from incognito or a dedicated account |
| Location controls | City-level IP geolocation noise | Use explicit location parameters, not VPN guesses |
| Language vs location | Conflated variables, wrong root cause | Test each variable separately |
| Multiple runs over time | Single-snapshot false conclusions | Minimum 3 runs across different days |
| Per-engine tracking | Blended scores hiding real differences | Report ChatGPT, AI Overviews, Perplexity, Copilot separately |
| Citation slot awareness | Overweighting a citation on a high-slot engine | Weight mentions by typical slots per engine |
| Domain competition count | Misjudging how crowded a prompt type is | Check unique domains per response by prompt type |
| Domain authority bias | Assuming only giant sites get cited | Audit mid-tier (DR46-75) competitors, not just the top names |
| Crawler log verification | Spoofed bot traffic inflating logs | Match source IPs against published crawler ranges |
| Cross-platform fact consistency | Rating and NAP mismatches killing recommendation odds | Audit ratings and listings everywhere AI actually sources from |
Where tooling actually helps
You can run all ten checks manually with a spreadsheet and a lot of patience, and plenty of local SEO teams do exactly that for a quarterly audit. But continuous monitoring is where this gets hard to do by hand, especially the platform-wide shifts (check 4) and crawler verification (check 9), both of which change week to week in ways a one-off audit will always miss.
Promptwatch is built around exactly this problem. It monitors the actual user interfaces of ChatGPT, Gemini, Perplexity, Claude, AI Overviews, AI Mode and more rather than relying on APIs that can behave differently from what users see, tracks prompt-level citation data with volumes and difficulty scores, and includes AI crawler log analysis (Agent Analytics) that shows exactly which bots are hitting your site and whether they're hitting errors, with IP-level verification baked in rather than raw user-agent counting. It also separates results by engine instead of blending them into a single score, which, per check 5 above, is the difference between a useful report and a misleading one.

For businesses specifically managing multi-location listings and local review data, tools like Yext and Birdeye focus on keeping NAP data and reviews consistent across the feeds AI engines actually pull from, which maps directly to check 10.
If you're comparing across the broader category of AI visibility platforms, the directory at bestgeosoftware.com is a reasonable place to see how different tools stack up on engine coverage and crawler monitoring before you commit to one.
The bottom line
Local AI visibility data is easy to generate and surprisingly easy to get wrong. The gap between what a dashboard shows and what's actually true can be wide, not because the tools are lying, but because the underlying methodology (session state, location handling, snapshot timing, crawler verification) introduces errors that compound quietly. Run the ten checks above before you act on a number, and re-run them quarterly. The AI search engines themselves aren't holding still, so your verification process can't either.

