Key takeaways
- A credible GEO case study separates measured data from sampled estimates from inferred outcomes, and says so out loud
- Single-snapshot "look, we got cited" screenshots are weak evidence because LLM outputs are non-deterministic; run each prompt multiple times before claiming a trend
- Structure the case study around the client's problem and industry, not their name, because that's how AI-driven prospects will eventually find it too
- GA4's AI Assistant channel, Search Console's generative AI report, and self-reported attribution each cover part of the picture, none covers all of it
- Tools like Promptwatch can supply the baseline citation tracking and crawler logs that make the "why" behind a result defensible, not just the "what"
Why clients don't trust GEO reporting yet
I've sat in enough of these meetings to know the moment. You show a client a chart, AI referral traffic is up, citations are up, and somewhere around slide four someone says "how do we know this is real and not just you picking a good month?" It's a fair question. GEO reporting is young enough that most of what passes for a case study is a highlight reel: one good screenshot of ChatGPT mentioning the brand, a vague traffic bump, a testimonial. None of that survives a skeptical CFO.
The problem isn't that AI search results can't be proven. It's that most agencies are proving them the way they used to prove SEO results, with ranking screenshots and a confident voice. AI answers don't work like the ten blue links, and your proof structure has to account for that or it falls apart the first time someone asks a follow-up question.
The core problem: AI answers are not deterministic
Start here, because it changes everything downstream. The same prompt run twice on ChatGPT can return different sources, different competitors, sometimes a different answer entirely. This isn't a bug you're hiding from the client, it's how generative retrieval works, and a client who understands it will trust your numbers more, not less.
Promptwatch's own data illustrates how much this noise matters. Around the GPT-5.3 rollout on March 4, 2026, average citations per ChatGPT response dropped about 27% across every model variant, from roughly 6.4 sources per response the week before to 4.7-4.9 by late March, with no recovery a month later (Promptwatch Data, March 2026). If a client's citation count had dropped that week, and you hadn't checked the date against known model releases, you'd have blamed the content team for a platform-wide shift that had nothing to do with their work.
The practical fix: run each prompt in your tracking set three to ten times before drawing a conclusion, and report a trend across at least four weekly checks, not one lucky screenshot. If you're doing this by hand across ChatGPT, Perplexity, Gemini, and AI Overviews, budget half a day per client just for the baseline. A platform that automates the repeat sampling buys back that time.

Section-by-section structure for the case study itself
Here's the structure I'd actually hand a client, built to survive scrutiny rather than just look good in a deck.
1. The baseline (before you touched anything)
Document the starting citation rate across engines, using a fixed set of 15-25 real buyer questions, each run multiple times. Log per engine: was the brand cited, who was cited instead, and how the brand was described when it did appear. This is your control group. Without it, every later number is unverifiable.
2. The methodology, in plain language
State the query set, the engines checked, the sample size, and the dates. This section is boring on purpose. It's also the single biggest credibility lever in the whole document, because it's the part competitors skip. A GEO agency comparison framework puts "proof and case evidence" worth 20 of 100 scoring points specifically on whether before/after claims are tied to a documented query set with dates and screenshots, versus vague claims with no evidence (launchcodex.com).
3. Branded vs category query results
Split the question set. Branded queries name the client directly ("Is [Brand] good for X?"). Category queries don't ("best CRM for manufacturing companies"). Branded results tell you whether the AI's existing knowledge about the client is accurate. Category results tell you whether the client is even in the consideration set when nobody has heard of them yet, which is the scenario that actually drives new business.
4. Content extraction-readiness audit
Check whether the top revenue pages answer the question in the first two sentences, use question-shaped headings, and state facts as liftable sentences rather than marketing copy. One detail worth bringing into this section: ChatGPT's own search queries have compressed from an average of roughly 117 characters in December 2025 down to around 53 characters once fan-out behavior returned in April 2026 (Promptwatch Data, query fanouts). Headlines and H2s that read like short search terms, not full sentences, now match how the model is actually querying the web.
5. Technical and schema pass
Quick checks: does FAQ/Article schema validate, does the page render without JavaScript (the view-source test), are canonicals and robots directives clean. Google's own guidance is blunt that structured data isn't required for AI features and there's no special schema just for AI, but it still helps with classic rich-result eligibility, so it's worth doing anyway, just don't oversell it as a magic switch (developers.google.com/search).
6. Competitor citation comparison
Run the same baseline query set and log which competitors get cited where the client doesn't. This is where the case study earns its keep with skeptical stakeholders, because "here's exactly who's beating us and on what page" is a lot more persuasive than an abstract visibility score.
7. Results, segmented and honestly labeled
This is the section most agencies botch. Group metrics into four buckets: AI visibility (citation frequency, platform coverage), traffic (AI referral sessions), conversions (from AI channels), and authority (sentiment, citation context). Present before/after side by side with dates, segment by platform since ChatGPT behavior doesn't match Perplexity behavior, and label every number by how it was actually derived: directly measured, sampled estimate, or inferred. A reporting framework worth copying labels rows like "AI-influenced pipeline (not session-traced)" instead of "AI-attributed revenue," precisely to avoid implying a traceability that doesn't exist (almcorp.com).
8. Limitations, stated plainly
Name what couldn't be measured. The honest version of this section says something like: a user can be recommended your brand by ChatGPT, not click, Google the brand name later, and convert through organic search with zero trace back to the AI exposure. That's not a hole in your reporting, it's a known property of how AI-influenced research behaves, and saying so up front disarms the skeptic before they find it themselves.
9. The 30/60/90-day roadmap
Strengths first, then gaps ranked by revenue relevance, then a dated plan with an owner per fix. Structuring it as strengths-gaps-path, in that order, lets the client see what's already working before they see the list of work ahead.
A comparison table worth putting in the deck
| Measurement layer | What it tells you | Reliability | Common mistake |
|---|---|---|---|
| Visibility (is the brand mentioned) | Presence in a sampled prompt set | Non-deterministic, run-to-run variance | Treating one screenshot as proof |
| Influence (cited or recommended) | Whether the AI names and links the brand | Citations are observable; "recommendation" tone is subjective | Conflating a mention with an endorsement |
| Outcome (business results) | Real traffic and conversions | Only traceable with a click and a referrer header | Claiming revenue attribution with no session trace |
Where the analytics actually come from, and where they fall apart
GA4 added a native "AI Assistant" channel in May 2026 that automatically recognizes sessions from ChatGPT, Gemini, DeepSeek, Copilot, and Grok. Perplexity is not in that native grouping and still lands in "Referral," so if you're reporting on a client that gets meaningful Perplexity traffic, you need a custom channel group built on top, with the AI rule positioned above the Referral rule or it never matches (digitalapplied.com).
Even with that fixed, an estimated 60-70% of real AI referral sessions arrive with no referrer at all, largely from mobile in-app browsers and the ChatGPT Atlas browser stripping headers, and land silently in "Direct." Treat whatever GA4 shows you as a measurement floor, not the full picture. Google Search Console's generative AI performance report, rolled out worldwide in mid-2026, adds impressions for AI Overviews and AI Mode, but as of this writing it has no click data at all, only impressions.
The gap-filling tactic worth stealing: add "how did you hear about us?" as a free-text field at signup or demo booking, then classify responses that name an AI tool and reassign them when the tracked source was Direct or Organic. One vendor reported this moved AI search's share of tracked conversions from 2% under last-click to 16% with first-click plus self-reported reattribution, an eightfold recovery (segmentstream.com). It's not perfect, but it's a lot closer to the truth than ignoring the gap.

Set expectations on how contested citation slots actually are
One detail that calms a lot of skeptical clients down: ChatGPT typically cites around 5 sources per web-search-enabled response, roughly half the ten slots classic Google used to hand out, which makes every citation genuinely contested. Google AI Overviews cites closer to 10, and Perplexity sits almost exactly at 10 as well, making it the most consistent engine to use as a GEO test bench since citation-count swings there usually reflect your own content changes rather than platform noise (Promptwatch Data, average sources per response). Microsoft Copilot is the outlier, swinging from under 2 to nearly 17 sources per response within weeks, so judge it on monthly trends only.
It's also worth killing the assumption that only Wikipedia-tier domains get cited. In August 2026, domains ranked DR 46-75 earned close to half of all ChatGPT citation share, while the very top tier, DR 91-100, fell from roughly 7% to about 3% share over the same month as mid-authority sites absorbed the gap (Promptwatch Data, citation share by domain rank, August 2026). For a mid-market client who assumes they can't compete with the big incumbents, that's the single most useful stat in the deck.
Naming the case study so it gets found by the next client
One structural choice people skip: don't title the case study after the client. Nobody prompts an AI with "Acme Corp case study" unless they already know Acme. Buyers ask about problems and industries, "best CRM implementation consultants for manufacturing companies," so the H1 and URL should follow a pattern like "[service] case study: [industry + size] + [specific outcome]" rather than the client's name. The same discipline that makes a case study persuasive to a skeptical client also makes it retrievable by the next prospect who's never heard of either of you.

Tools that make this repeatable instead of a one-off
You can run the baseline citation check, the competitor comparison, and the repeat-sampling by hand, and plenty of agencies do. It's a half-day per client per month, minimum, and it doesn't scale past a handful of accounts. A platform built around the full visibility stack, crawler logs showing exactly when AI bots visit a page, citation trend analysis segmented by content type, and visitor analytics tied to actual conversions, turns sections 1, 3, and 6 of the structure above into something you pull in an afternoon rather than build from scratch each time.
Promptwatch is the platform 1001 SEO Media uses for exactly this, partly because it goes past tracking into actually fixing the gaps it finds, with content gap analysis, automated CMS publishing, and a prioritized action list rather than just a dashboard full of numbers. If you're weighing it against other options, the GEO software directory at bestgeosoftware.com is a reasonable place to compare the full field before committing.
Common red flags to watch for, on either side of the table
If you're the agency being vetted, or the client doing the vetting, the warning signs are the same: an agency that can't name the specific AI surfaces it tracks, that only talks about rankings and never citations or share of voice, that can't explain its query-set methodology, or that sells a fixed GEO package with no discovery phase. A credible 30-day engagement produces a documented query set, a baseline capture, an entity and schema audit, a prioritized backlog, and a reporting template with metric definitions spelled out before any invoice gets sent.
A realistic first deliverable
If you only build one thing from this guide, build the baseline. Fifteen to twenty-five real buyer questions, run three times each across the engines your client's buyers actually use, logged in a spreadsheet with cited yes/no, competitor cited instead, and description accuracy. That single artifact turns an abstract worry, "are we even showing up?", into a named, dated, defensible answer. Everything else in the case study structure above is built on top of it, and none of it holds up without it.