Key takeaways
- There is no separate "Copilot bot." Copilot's retrieval runs on Bingbot and Microsoft's Prometheus orchestration layer, so your logs will only ever show
bingbot/2.0, never a dedicated Copilot user agent. - Bingbot's crawl footprint is small compared to Googlebot (Cloudflare found Googlebot reaches 3.26x more unique URLs than Bingbot) and tiny compared to OpenAI's crawler volume, so don't panic if Bingbot hits look sparse next to GPTBot.
- UA strings are spoofable. Always verify Bingbot hits against the published IP ranges at bing.com/toolbox/bingbot.json or Bing's Verify Bingbot tool before trusting a log line.
- A recurring 2026 problem: Bingbot connections to dynamic pages ending in HTTP 499 or timing out, even when Googlebot fetches the same URL fine. If your sitemap status flips between "Success" and "Sitemap could not be fetched" in Bing Webmaster Tools, this is almost always the cause.
- Bingbot hits confirm indexing, not citation. Because Copilot answers from Bing's pre-built index rather than crawling live per query, a healthy crawl log tells you Bing can find you, not that Copilot is about to cite you today.
Why Bingbot is the only name you need to learn
I'll save you some time digging through user-agent databases: Microsoft has not shipped a distinct Copilot crawler. Every retrieval operation behind a Copilot answer runs through Bingbot feeding Bing's index, and a separate orchestration layer called Prometheus that scores and retrieves from that index at query time. There's no CopilotBot/1.0 token to grep for, no agentic-action UA to block or allow separately. If you see "bingbot" in your logs, that's Copilot's supply chain. If you don't, Copilot has nothing to cite.
This trips people up because every other major AI lab seems to be shipping a new crawler every few months. OpenAI has GPTBot, OAI-SearchBot, and ChatGPT-User, each with a different job. Bing consolidated instead. Microsoft currently operates five documented crawlers:
- Bingbot - the main crawler, responsible for the vast majority of what you'll see
- AdIdxBot - crawls sites for Bing Ads quality checks
- BingPreview - generates page snapshots
- MicrosoftPreview - snapshots for other Microsoft products
- BingVideoPreview - video-specific previews
All five come in desktop and mobile variants. One real pitfall worth flagging early: a generic "block all bots" rule on your CDN or WAF often takes out Bingbot while leaving AdIdxBot, BingPreview, and the others running completely unaffected under their own UA strings. You can think you've blocked Bing and be wrong.
The current Bingbot user-agent string
The desktop version looks like this:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36
The Chrome version number tracks whatever rendering engine Bing's using internally (Microsoft has cited something like 80.0.345.0 as an example), so don't hardcode a specific version number in any filtering logic. There's also a mobile variant that spoofs a Nexus 5X on Android 6.0.1, which catches people out if they're filtering by device string instead of the bingbot/2.0 token buried in the UA.
Don't trust the UA string, verify the IP
Microsoft says this plainly in its own documentation: a user-agent string is trivially spoofable, and scrapers impersonating Bingbot are common enough that you shouldn't take any log line at face value. Two ways to check:
- Bing's public Verify Bingbot tool at bing.com/webmasters/verifybingbot
- Cross-referencing the request IP against the published ranges at bing.com/toolbox/bingbot.json
I'd make this step one of any log audit, not an afterthought. If you're running automated alerts on crawler activity, build the IP check into the pipeline rather than trusting the User-agent header alone. The same logic applies to every AI crawler, not just Bing's. ClickRank's crawler guide and Human Security's 2026 bot report both make the same point from different angles: logs show you who's crawling, but only verified IPs tell you who's actually telling the truth.
What Bingbot actually looks like next to everything else in your logs
Here's where it helps to zoom out. Promptwatch's crawler-traffic data shows OpenAI alone accounted for 79.8% of verified AI crawler requests in the first week of September 2026, down from a staggering 94.8% share back in June. Promptwatch's AI crawler traffic report tracks OpenAI, Anthropic, Google, Perplexity, and Mistral as distinct provider segments, and Bingbot doesn't show up as a tracked line item there at all, because it predates the "AI crawler" framing and is treated as a legacy search crawler rather than a training or agentic bot.

That matters for expectations. If you're used to seeing GPTBot hammer your logs daily and then go looking for the Copilot equivalent, you'll probably be underwhelmed. Cloudflare's January 2026 crawler analysis found Googlebot reached 3.26x more unique URLs than Bingbot over a two-month window, and 167x more than PerplexityBot. Bingbot's crawl footprint trails even Google's "legacy" index crawler, let alone the training-focused bots racking up volume right now. A quiet Bingbot section in your logs isn't necessarily a problem. It's just Bing being Bing.

Crawl activity vs. citation: two different things
This is the part people get backwards most often. Copilot doesn't crawl your site live when someone asks it a question. It queries Bing's pre-built index, which Prometheus then scores for relevance and hands to a GPT-class model for synthesis. So the Bingbot hit timestamped in your logs reflects indexing activity. It tells you Bing found and fetched the page at some point. It does not tell you the exact moment, or even the same week, that a Copilot answer cited that page to a user.
Practically, this means a healthy stream of 200-status Bingbot requests is necessary but not sufficient. ai-advisors.ai puts it bluntly in their Copilot citation guide: if Bingbot can't reach your site, nothing downstream matters, full stop. But a well-crawled site with zero Copilot citations is still a completely normal and common state, because citation depends on relevance scoring inside Prometheus, not just crawl access.
The 499 problem nobody warns you about
A pattern that's shown up repeatedly in Microsoft's own community forums through 2026: Bing Webmaster Tools reports a sitemap status that flips between "Success" and "Sitemap could not be fetched," with no code changes on the site owner's end. Pull the server or CDN logs for the exact timestamps and you'll often find Bingbot's connection ending in an HTTP 499 (client closed request) or a straight timeout, specifically on dynamic or JS-rendered directory pages. Googlebot fetches the identical URL in the same window without issue.
Microsoft's own community volunteers have been candid that they have no visibility into Bing's internal crawler telemetry on these threads, and the standard advice is to open a direct support ticket with log evidence: the sitemap URL, exact timestamps, and the 499 log lines themselves. If your dynamic pages take a beat longer to render than static ones, this is worth checking before you assume your content strategy is the problem.
Separately, and a little unsettling: there are documented cases of IPs that verify as legitimate Bingbot, matching published ranges like 20.15.133.160/27 or 40.77.167.0/24, sending malformed query parameters completely unrelated to the site's content, things like query=wells+fargo+zelle showing up on an unrelated site's logs. Whether that's Bing repurposing verified ranges for other traffic or some kind of spoofing riding on legitimate IP blocks isn't fully clear, but several site owners have reported these IPs landing on bot blacklists as a result. Worth knowing before you auto-block an IP range that's technically "real" Bingbot.
A practical log-reading checklist
Here's the sequence I'd run through on any site that isn't showing up in Copilot the way it should:
- Filter logs for
bingbot/2.0in the user-agent string, across both desktop and the Nexus 5X mobile variant - Cross-check the source IPs against bingbot.json before trusting any of the results
- Check status codes on those verified hits. Anything other than 200, especially repeated 403, 429, or 499, is worth investigating
- Look at request paths. Is Bingbot reaching your highest-value pages, or stalling on a handful of URLs?
- Run a direct test with curl against your own server using a Bingbot-style UA string, to rule out a server-side block before blaming Bing:
curl -I -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/80.0.345.0 Safari/537.36" https://yoursite.com
- Submit your sitemap directly to Bing Webmaster Tools. Google Search Console submission does nothing for Bing's pipeline, they're entirely separate systems
- Ping changed URLs through IndexNow for faster notification, but don't mistake notification for indexing. Microsoft's own Q&A is explicit that IndexNow guarantees Bing is told promptly, not that the page gets crawled or indexed any faster
A quick reference table
| Crawler | Primary job | Where it shows in logs | Notes |
|---|---|---|---|
| Bingbot | Builds Bing's index, feeds Copilot | bingbot/2.0 in UA | Verify via bingbot.json, respects robots.txt |
| AdIdxBot | Bing Ads quality checks | Separate UA string | Often missed when blocking "Bing" generically |
| BingPreview / MicrosoftPreview | Page and product snapshots | Separate UA strings | Runs independently of Bingbot |
| OAI-SearchBot | Feeds ChatGPT Search | OAI-SearchBot/1.4 | Does not feed Copilot directly, but surfaces in Edge sidebar/Windows search |
| GPTBot | Training corpus collection | GPTBot/1.4 | Unrelated to search retrieval, blocking it doesn't affect Copilot |
Where log analysis fits next to AI visibility tracking
Server logs tell you whether Bing can reach your content. They don't tell you whether Copilot, ChatGPT, Gemini, or Perplexity are actually citing you in answers, or how often, or for which prompts. That's a different layer entirely, and it's the gap most GEO-focused platforms exist to fill.
Promptwatch tracks citation-side data across ChatGPT, Gemini, Claude, Perplexity, Copilot and Google's AI surfaces, alongside its own crawler-log ingestion (via Cloudflare, Vercel, AWS CloudFront, and other CDNs) so you can line up a Bingbot crawl event with an actual citation outcome, rather than guessing. Its crawl-to-citation tracking is the kind of thing that turns "Bingbot showed up in my logs" from a trivia fact into an actionable signal.
If you're building out a broader AI-crawler monitoring stack, it's worth browsing the GEO software directory at bestgeosoftware.com to compare platforms that handle both the log and citation sides of this problem, since most tools on the market only do one or the other.
The bottom line
Bingbot is the whole story for Copilot. There's no secondary crawler to chase, no separate agentic UA to allow, just one crawler feeding one index that one orchestration layer queries at inference time. The work is in verification (don't trust the UA alone), health checks (watch for 499s on dynamic pages), and managing expectations (a quiet Bingbot log doesn't mean you're broken, it means Bing crawls less aggressively than everyone else right now). Get those three right and you've done everything the crawl side of Copilot visibility actually requires.