Google AI Overviews crawler behavior in 2026: what Google-Extended and GoogleOther actually do

Google-Extended isn't a crawler, GoogleOther isn't Googlebot, and blocking either one won't remove you from AI Overviews. Here's what actually controls your visibility in Google's AI search features in 2026.

Key takeaways

  • Google-Extended is not a crawler. It never visits your server. It's a robots.txt control token that tells Google whether content Googlebot already collected can be used to train and ground Gemini, including Google AI Overviews.
  • AI Overviews and AI Mode are built from the regular Google Search index, the one Googlebot builds. Blocking Google-Extended does not remove your pages from AI Overviews.
  • GoogleOther is a genuine, separate fetcher Google uses for research and internal testing, distinct from both Googlebot and Google-Extended, and it's your call whether to allow it.
  • A March 2026 BuzzStream analysis found 92.3% of news sites blocking Google-Extended were still cited in AI answers, which tells you almost everything about how much that token actually controls.
  • Googlebot sees roughly 1.7 to 1.8 times more unique URLs than ClaudeBot or GPTBot, and over 160 times more than PerplexityBot, according to Cloudflare's crawler traffic analysis, which is a big part of why Google's AI features have so much raw material to pull from.

The naming problem that keeps tripping people up

I get why this confuses people. "Google-Extended" sounds like a crawler. It has "Google" in the name, it shows up in your robots.txt right next to Googlebot, and Google's own documentation lists it alongside other user agents. Every instinct says: treat it like a bot you can block.

But it isn't one. No request ever hits your server from Google-Extended. There's no IP address, no log line, nothing to inspect in your access logs. It's a permission flag, nothing more. Regular Googlebot does all the fetching. Google-Extended just tells Google's systems, after the fact, whether that already-fetched content is fair game for training Gemini and grounding its answers.

That distinction matters more in 2026 than it did when Google introduced the token in September 2023, because AI Overviews now appear on a large and growing share of search queries, and site owners are understandably anxious about controlling what feeds them.

Overview of Google's crawlers and fetchers documentation page listing Googlebot, GoogleOther, Google-Extended and other user agents

What Google-Extended actually controls

Here's the mechanism in plain terms. Googlebot crawls your site for Search, same as it always has. Separately, Google's AI training pipeline asks: is this URL allowed for Google-Extended? If you've disallowed it, that page's content is excluded from training data for Gemini models and from the grounding corpus Gemini Apps and Vertex AI draw on. If you haven't, it's eligible.

What it does NOT do is pull your pages out of AI Overviews or AI Mode. Those features are generated from the standard Search index, the one built entirely by Googlebot. Google has said this directly, and the independent data backs it up. PPC Land's timeline cites a March 2026 BuzzStream finding that 92.3% of news sites blocking Google-Extended were still showing up in AI answers. That's not a rounding error. That's basically everyone.

So if your goal is "keep my content out of AI Overviews," Google-Extended is the wrong lever. The only domain-level setting that actually governs inclusion in AI Overviews, AI Mode, and generative features in Discover is a separate control, and it has nothing to do with the training-opt-out token everyone fixates on.

A quick worked example

Say you run a news site and want your archive protected from AI training but you still want new articles eligible, maybe because you've negotiated something, or you just don't care about older content. Google's own documentation example shows exactly this kind of split: disallow Google-Extended for most of a directory while allowing a single archived page. You get granular control over training eligibility. You get zero control over AI Overviews inclusion through this same mechanism, because that's a different system entirely.

Where GoogleOther fits in

GoogleOther is a different animal again, and it gets lumped in with the Google-Extended confusion way too often. Unlike Google-Extended, GoogleOther is a real crawler. It makes actual HTTP requests, it shows up in your logs, and you can block it at the server level if you want to.

Google uses GoogleOther for internal research and testing purposes, separate from both the main Search indexing crawl and the AI-training opt-out mechanism. The practical guidance from most crawler-tracking resources is blunt about it: GoogleOther doesn't feed Search rankings directly, so blocking it is "your call" rather than something with an obvious right answer either way. If you're trying to reduce load from Google's various fetchers without touching your Search visibility, GoogleOther is a far safer target than Googlebot and a more concrete one than the Google-Extended token.

Comparison table distinguishing Googlebot from Google-Extended, showing one is a crawler and the other is a training opt-out control

The table everyone needs pinned above their robots.txt editor

User agentWhat it isWhat blocking it doesSafe to block?
GooglebotReal crawler, indexes for SearchRemoves you from Google Search entirelyNo, almost never
Google-ExtendedNot a crawler, a training opt-out tokenOpts pages out of Gemini training/grounding only. Does not remove you from AI OverviewsYes, with no visibility cost in Search or AI Overviews
GoogleOtherReal crawler, used for internal research/testingReduces server load from Google's test fetches, no confirmed ranking impactYour call
Google-Agent and similar user-triggered fetchersFetchers acting on a direct user request (e.g. inside Gemini apps)These officially ignore robots.txt because they're responding to an explicit user actionNot really controllable via robots.txt

The pattern across this table is consistent: anything tied to AI training or testing is low-stakes to block. Anything tied to indexing or verification is not. That's the whole decision tree, really.

Why Google's position here is structurally different from other AI crawlers

This is the part that should bother publishers more than it seems to. Cloudflare's analysis of crawler traffic found that Googlebot accessed roughly 1.7 to 1.8 times more unique URLs than ClaudeBot and GPTBot over a two-month observation window, about 3 times more than Meta-ExternalAgent, and more than 166 times more than PerplexityBot. Almost no website explicitly blocks Googlebot in full, which makes sense given how dependent most sites are on Search referrals, but it also means Google has a volume advantage in raw web content that competing AI companies simply can't match by asking nicely.

Cloudflare's point, made in their blog post on the UK Competition and Markets Authority's proposed conduct requirements, is that Google can use its search crawler to gather data for a wide range of AI functions with comparatively little risk of being blocked, because doing so would cost a site its Search traffic too. Other AI companies don't get that cover. Their crawlers can be, and increasingly are, blocked outright at the network edge without any collateral damage to a site's Search presence. Data from Cloudflare's AI Crawl Control customers between July 2025 and January 2026 showed sites blocking crawlers like GPTBot and ClaudeBot at nearly seven times the rate they blocked Googlebot or Bingbot.

That asymmetry is the actual story behind the Google-Extended debate. The opt-out token exists, technically, but publishers who've tried to use it to limit Google's AI reach have found it doesn't touch the thing they actually care about, which is appearing (or not appearing) inside AI Overviews. Cloudflare has said plainly that feedback from its own customers is that Google's proprietary opt-out mechanisms, including Google-Extended and the nosnippet directive, have failed to give publishers meaningful control over how their content ends up in generative features.

What this means for your robots.txt strategy

A few practical conclusions, in order of how often I'd actually act on them:

First, don't touch Googlebot unless you genuinely want to vanish from Google Search. That one's non-negotiable advice, and I'd be surprised if anyone reading this didn't already know it, but it's worth restating because Google-Extended's confusing name makes people nervous enough to second-guess the basics.

Second, treat Google-Extended as a training-data decision, not a visibility decision. If you have a philosophical or business objection to your content training Gemini, block it. It costs you nothing in Search or AI Overviews. If you don't care, leave it alone. Either way, stop expecting it to change whether you show up in an AI Overview panel, because it won't.

Third, consider other AI crawlers separately. Google-Extended is one opt-out among roughly 88 AI-related crawlers tracked across the web as of 2026. You can allow Google-Extended while blocking GPTBot, or vice versa. There's no rule that says you have to treat every AI company's training crawler the same way. Some sites decide AI Overviews visibility is worth more than keeping content out of Gemini's training set; others decide the opposite. Either is defensible, just make the decision deliberately rather than by copying a robots.txt template you found online.

Fourth, if GoogleOther is generating noticeable load and you don't see a reason to allow it, block it. It's not tied to ranking in any confirmed way, and unlike Googlebot, there's no documented downside.

Monitoring what's actually happening, not just what robots.txt says

Here's the uncomfortable truth underneath all of this: robots.txt is a request, not an enforcement mechanism. Nothing stops a crawler from ignoring it, and plenty of site owners discover that their carefully written disallow rules didn't change their AI Overviews presence at all, because that presence was never governed by the rule they edited.

If you want to know what's actually happening rather than what you've asked to happen, you need visibility into real crawl behavior and real citation outcomes. Tools like Promptwatch track AI crawler activity against your site (which bots hit which pages, what errors they encounter) alongside whether those pages actually get cited in AI Overviews, AI Mode, ChatGPT, Perplexity, and the rest. That crawl-to-citation path is the thing robots.txt audits alone can't show you.

Favicon of Promptwatch

Promptwatch

Track and optimize your brand's visibility in AI search engines
View more
Screenshot of Promptwatch website

For AI Overviews specifically, Promptwatch's data on citation types shows product pages overtaking listicles as the most-cited format in late July 2026, which is worth knowing if you're deciding what to protect or promote. If you want the raw numbers, Promptwatch's AI Overviews citation types report for July 2026 breaks down the daily mix of listicles, how-tos, news, and product pages.

A few other tools worth a look if you're building out a broader AI crawler monitoring setup: DarkVisitors specializes in tracking AI agents and bot traffic against your logs, and Botify pairs enterprise-scale crawl analysis with GEO-specific signals.

Favicon of DarkVisitors

DarkVisitors

Track AI agents, bots, and LLM referrals visiting your websi
View more
Screenshot of DarkVisitors website
Favicon of Botify

Botify

Enterprise SEO and GEO platform with AI agents for search vi
View more
Screenshot of Botify website

Where this is heading

The UK's Competition and Markets Authority designated Google as having Strategic Market Status in general search in October 2025, a designation that explicitly covers AI Overviews and AI Mode. The CMA has since opened consultation on conduct requirements aimed at giving publishers more meaningful control over how their content feeds Google's generative features. Cloudflare's position, laid out in their blog post on the proposal, is that the CMA's current remedies don't go far enough, and that nothing short of forcing Google to separate its AI crawling from its Search crawling at the technical level will give publishers real choice. Whether that separation ever happens is a regulatory question, not a technical one you can solve today. For now, the practical reality hasn't changed: Googlebot runs your Search presence, Google-Extended only governs AI training eligibility, GoogleOther is a research crawler you can mostly ignore or block at will, and none of robots.txt's levers will pull you out of AI Overviews if Google decides to cite you.

If you're trying to figure out which of the dozens of AI visibility and crawler-monitoring tools fit your setup, the directory at bestgeosoftware.com is a reasonable starting point for comparing options side by side.

Share:

© 2026 AI Search Tools · Best AI search tools and platforms · RSS

AI Search Tools is an affiliate review site. When you click links to vendors or buy through links on our site, we may earn an affiliate commission at no extra cost to you.

AI Search Tools is a review website based on user reviews on Reddit and G2, and on publicly available information. We keep everything as up to date as possible, but pricing and features can change. Always confirm the details with the vendor before purchasing.