Key takeaways
- Cloudflare's AI Crawl Control dashboard has three main sections: Overview, Crawlers, and Metrics, plus a newer Robots.txt tab, and every Cloudflare domain gets it by default.
- The single most common mistake marketers make is blocking all AI bots at once, which kills ChatGPT, Perplexity, and AI Overviews visibility along with training scrapers you actually wanted blocked.
- GPTBot and OAI-SearchBot are not the same crawler. One trains models, the other powers live ChatGPT search citations. Treat them differently.
- A huge crawl-to-refer ratio (Claude's has been reported north of 11,000:1) is not automatically a red flag. Some AI apps just don't send referrer headers.
- Crawl volume in Cloudflare's dashboard tells you who's visiting. It doesn't tell you who's citing you. Pair it with citation-tracking data to see the full picture.
Why marketers suddenly need to care about crawler logs
For most of SEO history, nobody outside the technical team looked at server logs. That's changed. In June 2026, Cloudflare CEO Matthew Prince shared Radar data showing automated requests had, for the first time, overtaken human traffic on the open web: 57.5% of HTML requests now come from bots, not people. AI crawlers alone made up roughly 20.3% of verified bot traffic that same month, according to Cloudflare Radar.
That shift means the question "is my content getting crawled by AI?" is no longer a developer's problem. It's a visibility problem, and visibility is a marketing job. If GPTBot, ClaudeBot, or PerplexityBot can't reach your pages, you simply don't exist in ChatGPT or Perplexity answers, no matter how good your content is. Cloudflare's AI Crawl Control (originally called AI Audit) gives every Cloudflare customer, on every plan, a dashboard that shows exactly which AI systems are hitting their site, what they're reading, and whether they're being let in. Most marketers have never opened it. This guide walks through what you'll actually see.
Finding the dashboard
Log into the Cloudflare dashboard, pick your account and domain, then go to AI Crawl Control. If you've never touched it, it's almost certainly already running in the background, since Cloudflare turned this on for every site by default back in September 2024. You're not setting something up from scratch. You're reading data that's already been collecting.

The overview tab: your five-second health check
The Overview tab is designed to be skimmed, not studied. It shows total request volume, how that volume has changed, the most common status code your server returned to AI crawlers, and the single most popular path AI bots hit. There's also a quick check on whether Cloudflare's managed robots.txt is enabled.
Below that, crawlers are grouped by operator, OpenAI, Microsoft, Google, ByteDance, Anthropic, Meta, with allowed requests and percentage change per operator. Paid plans also get referral counts here, meaning how many actual human visits came after a click from an AI answer.
If you only have five minutes, this is the tab to check. Look for two things: a status code that's mostly 403 (meaning you're blocking crawlers, intentionally or not), and an operator whose activity dropped sharply with no explanation. Both are worth investigating further in the Metrics tab.
The crawlers tab: who's visiting, and are you letting them in
This is a simple table: crawler name, category (search, training, agent, and so on), request volume with a trend sparkline, any robots.txt violations, and an Allow/Block toggle. Click into any row and you can view that crawler's detailed metrics or jump to its profile on Cloudflare Radar.
The categories matter more than the raw numbers. Cloudflare tags verified bots with behavior labels including Search, Agent, Training, Transact, Data Collection, and several others. A bot tagged Training is scraping your content to improve a model's general knowledge. A bot tagged Search or Agent is fetching your page in real time to answer a specific user question, and might cite you. Those are fundamentally different relationships with your content, and they should get different treatment.
The metrics tab: where the real analysis happens
This is the dense one, and it's where most of the useful signal lives.
Requests over time lets you toggle between all requests, allowed requests only, or data transfer, and group by crawler, category, operator, or host. This is the best view for spotting a sudden spike or drop. If a single operator's volume doubled overnight, something changed, either your site or their crawler behavior.
Status code distribution breaks down 2xx, 3xx, 4xx, and 5xx responses over time. Pay close attention to 403s here. A wall of 403 responses to a crawler you thought was allowed usually means a WAF or Bot Management rule is blocking it at the network edge, before robots.txt is even checked. Robots.txt rules do not override firewall rules. This order-of-operations issue trips up a lot of people who assume an "Allow" in robots.txt settles the matter.
Content Format compares what AI crawlers are asking for (via the Accept header) against what your server actually serves (via Content-Type). A mismatch here can mean a crawler requesting JSON or plain text and getting a bloated HTML page back, which isn't necessarily broken but is worth knowing.
Top referrers and referrals over time show which AI platforms are actually sending human traffic to you, chatgpt.com, perplexity.ai, and so on, and which pages or path patterns those referrals land on.
Most popular paths is arguably the most actionable table in the whole dashboard. It lists exact paths and hostnames AI crawlers hit most, with a Patterns view that groups similar URLs (like /blog/* or /api/*) up to two levels deep. If your pricing page is getting hammered by training crawlers but your help docs aren't getting touched at all, that's a content strategy signal, not just a technical curiosity.
Every chart and table in the Metrics tab respects a shared filter bar (date range, crawler, operator, hostname, path) and can be exported as a CSV or image, which makes it easy to drop into a monthly report.
The robots.txt tab: don't skip this one
Added in early 2026, this tab shows whether your robots.txt file is healthy and how AI crawlers are actually interacting with it, as opposed to how you think they are. Cloudflare's managed robots.txt layers its own Content-signal directives on top of whatever rules you've already written. The default signal is roughly "search=yes, ai-train=no, use=reference," with a list of named crawlers individually disallowed for training. If you've never checked this tab, there's a decent chance Cloudflare's managed rules are doing something slightly different from what you assume your own robots.txt says.
Cloudflare's own audit found that only about 7.8% of robots.txt files it checked explicitly disallowed GPTBot, and fewer than 5% addressed anthropic-ai, PerplexityBot, ClaudeBot, or Bytespider by name. Most sites simply haven't made a deliberate choice here. That's either an opportunity or a liability depending on how your content gets used.
The mistake almost everyone makes: blocking everything at once
Cloudflare offers a one-click "Block AI Scrapers and Crawlers" toggle, and it's tempting. The problem is that it typically blocks training bots (GPTBot, ClaudeBot, Bytespider, Google-Extended) and live search/citation bots (OAI-SearchBot, PerplexityBot's search crawler) at the same time. Flip that switch and you've likely made yourself invisible in ChatGPT and Perplexity answers, not just protected your archive from training scrapers.
The fix is to handle crawlers individually rather than with a blanket rule. GPTBot and OAI-SearchBot, for example, are both operated by OpenAI but serve completely different purposes: one trains models, the other fetches pages in real time to ground a ChatGPT Search answer. A robots.txt pattern that blocks training but preserves search visibility looks like this:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
Same logic applies across Anthropic (ClaudeBot training crawler vs. its citation crawler) and others. Check the Crawlers tab category labels before you flip anything to Block.
Making sense of the crawl-to-refer ratio
Cloudflare Radar's AI Insights page tracks a metric worth understanding: crawl-to-refer ratio, how many times a crawler visits your site for every one referral click it sends back. Some of these numbers look alarming out of context. Anthropic's Claude has been reported at ratios in the tens of thousands to one in some measured periods; a third-party aggregator put ClaudeBot at 11,122:1 for the week of May 25 to June 1, 2026.
Don't panic at a big number by itself. A lot of this gap exists because native AI apps, Claude's app being one example, often don't send a referrer header at all, which means Cloudflare undercounts real referral traffic even when a human did click through. The ratio is directionally useful for spotting trend changes (is it getting worse month over month?) but shouldn't be read as a literal measure of how little value you're getting back.
Crawl volume isn't citation share: pair Cloudflare with visibility data
Here's the gap Cloudflare's dashboard can't close on its own: it tells you who's knocking on the door, not who's actually quoting you once they're inside. Promptwatch's crawler traffic data shows OpenAI accounted for roughly 79.8% of verified AI crawler requests in the week of August 31 to September 6, 2026, down from 94.8% in early June, a useful benchmark to compare your own Cloudflare logs against. If OpenAI has a huge share of the overall AI crawler landscape but is barely showing up in your site's logs, something on your end is likely blocking or redirecting it.
Similarly, Promptwatch's citation-share-by-domain-rank data for August 2026 found that DR 61-90 domains held roughly 42% of ChatGPT citations all month, while top-tier DR 91-100 domains actually fell from about 7% to 3% mid-month. That matters for strategy: chasing only the biggest, most authoritative domains for PR and backlinks isn't where ChatGPT is pulling most of its answers from right now.

Tools like Promptwatch track AI crawler traffic by provider alongside actual citation data, Reddit and YouTube mentions, and content gap analysis, closing the loop that Cloudflare's dashboard leaves open: crawl visibility plus citation visibility plus what to do about it.
Quick reference: Cloudflare tabs and what they tell you
| Tab | What it shows | When to check it |
|---|---|---|
| Overview | Total volume, top status code, top path, operator breakdown | Weekly health check |
| Crawlers | Per-crawler table with category and Allow/Block toggle | Before changing any block rule |
| Metrics | Deep charts: requests, status codes, content format, referrers, top paths | Monthly reporting, investigating anomalies |
| Robots.txt | Health of your robots.txt vs. Cloudflare's managed rules | After any CDN or robots.txt change |
A basic weekly workflow for marketers
Check the Overview tab for anything that jumps out, an operator with activity that dropped or spiked, or a spike in 403 responses. If something looks off, go to Metrics and filter by that operator or crawler to see the status code trend over the same window. Cross-check against the Robots.txt tab to see whether a managed rule changed underneath you. Export the Most Popular Paths table once a month so you can track which sections of your site AI systems actually read, and compare that against which pages show up in citations using a dedicated tracking tool, since Cloudflare's logs alone won't tell you that part.
If you manage more than one site, or you're reporting this up to a client or a CMO, tools like ZipTie or Promptwatch can combine crawler data with citation and traffic data in one dashboard rather than forcing you to stitch together Cloudflare exports by hand.
When the data points to a bigger problem
If your Cloudflare reports show you're letting AI crawlers in freely but you're still invisible in ChatGPT, Perplexity, or AI Overviews, the issue usually isn't access. It's content structure, lack of citable facts, or thin pages that don't give an AI system anything worth quoting. That's a different fix, and it's where a GEO-focused content process, or an agency that does this work full-time, tends to help more than another dashboard tweak. If you're trying to figure out whether that's a DIY problem or one worth handing off, 1001 SEO Media works on exactly this kind of technical-plus-content gap for brands trying to show up in AI answers.
For broader comparisons of GEO and AI visibility platforms beyond what's covered here, the directory at bestgeosoftware.com is worth a browse before you commit to one tool.
