Scope: this table covers provider-documented agents that materially affect AI search visibility, model-training controls, or user-directed retrieval. It is not a list of every automated agent on the web - entries based only on third-party evidence are labeled and grouped separately, not mixed in with the provider-documented rows. A handful of adjacent, product-specific agents (ad-safety crawlers, single-product fetchers) are listed in a short appendix rather than the main table; a few more, like Google's newer Vertex AI Agent crawler, are named but not detailed here because we couldn't independently confirm their current documentation in this pass.

Every string below was pulled directly from the operator's own published documentation, not copied from another list. Where a detail couldn't be confirmed first-hand - Anthropic's exact cosmetic user-agent string, ByteDance's official documentation for Bytespider, several operators' robots.txt compliance stance for their training/search crawlers specifically - that's stated outright instead of guessed at. One finding worth knowing before you use this: Google-Extended and Applebot-Extended aren't crawlers at all. Neither has a user-agent string of its own; both are robots.txt-only controls layered on top of existing crawl traffic. More on that below.

Core AI Visibility Agents

These are the crawlers and fetchers most likely to affect whether your content is retrieved, cited, or trained on. Version numbers in the user-agent examples are operator-supplied snapshots, not stable identifiers - filter your logs on the bolded product token, not the full string, since OpenAI and Google both document that version numbers change over time.

BotOperatorCurrent Example User-AgentWhat It Doesrobots.txt TokenDocumented robots.txt Behavior
GPTBotOpenAIGPTBot/1.4Training - crawls to build/update OpenAI's modelsGPTBotNot stated explicitly - OpenAI's guidance is to disallow the token to exclude a site from training, which implies compliance, but no sentence in OpenAI's docs directly affirms GPTBot honors robots.txt the way it does for ChatGPT-User's exception below.
ChatGPT-UserOpenAIChatGPT-User/1.0User-triggered - fetches a page live when a user asks ChatGPT to browse itChatGPT-UserMay not apply - OpenAI states directly: "Because these actions are initiated by a user, robots.txt rules may not apply."
OAI-SearchBotOpenAIOAI-SearchBot/1.4Retrieval/search - powers ChatGPT's search-result indexingOAI-SearchBotNot stated explicitly - OpenAI's guidance is to keep it unblocked so it "has access," which implies compliance, but the docs don't contain a direct affirmation.
ClaudeBotAnthropicName confirmed; exact version string not published by AnthropicTraining - "collecting web content that could potentially contribute to their training," per AnthropicClaudeBotHonors - Anthropic states its bots "respect 'do not crawl' signals by honoring industry standard directives in robots.txt."
Claude-UserAnthropicName confirmed; exact version string not published by AnthropicUser-triggered - "supports Claude AI users" when someone asks Claude to access a siteClaude-UserHonors - same statement above; notably, unlike every other operator on this page, Anthropic makes no user-triggered exception. Worth confirming in your own logs if that surprises you.
Claude-SearchBotAnthropicName confirmed; exact version string not published by AnthropicRetrieval - "navigates the web to improve search result quality"Claude-SearchBotHonors - same statement above.
PerplexityBotPerplexityPerplexityBot/1.0Retrieval/search indexingPerplexityBotNot stated explicitly - Perplexity's docs recommend allowing it "to ensure your site appears in search results," which implies compliance, but don't directly affirm it the way they do for Perplexity-User below.
Perplexity-UserPerplexityPerplexity-User/1.0User-triggered - live fetch when a user asks Perplexity to check a pagePerplexity-UserGenerally ignores - Perplexity states: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules."
Meta-ExternalAgentMetameta-externalagent/1.1Training - crawls "for use cases such as training foundation AI models or improving products by indexing content," per Metameta-externalagentNot stated explicitly - Meta's crawler documentation doesn't include a direct robots.txt compliance statement for this agent.
Meta-WebIndexerMetameta-webindexer/1.1Retrieval/search - "navigates the web to improve Meta AI search result quality for users," per Metameta-webindexerNot stated explicitly - same gap as Meta-ExternalAgent.
Meta-ExternalFetcherMetameta-externalfetcher/1.1User-triggered - "fetches individual links at a user's request," including agentic-AI product functions, per Metameta-externalfetcherMay bypass - Meta states this crawler "may bypass robots.txt rules" because it performs fetches requested by the user.
GooglebotGoogleMozilla/5.0 ... (compatible; Googlebot/2.1; +http://www.google.com/bot.html) - full string varies by desktop/smartphone variantSearch indexing; renders JavaScript (see the pillar page for why this makes Googlebot an exception to the JS-rendering problem)GooglebotHonors - Google classifies Googlebot as a "common crawler," and states common crawlers "always obey robots.txt rules when crawling automatically."
Google-AgentGoogleMozilla/5.0 ... (compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent) ...User-triggered - a newer agent, added to Google's crawling documentation March 20, 2026, for "Google agents hosted on Google infrastructure" that "navigate and act on user requests"Google-AgentGenerally ignores - Google-Agent is documented alongside Google's other user-triggered fetchers, which as a category "generally ignore robots.txt rules" because the fetch was requested by a user; Google's stated rationale doesn't call out Google-Agent by name specifically, but the blanket rule covers it.
ApplebotAppleMozilla/5.0 ... (compatible; Applebot/version; +http://www.apple.com/go/applebot) - varies by desktop/mobile variantPowers search features in Spotlight, Siri, and Safari; may also train Apple's foundation models. Renders pages in a browser context - blocking JS/CSS resources can cause incomplete renderingApplebotHonors - Apple states: "Applebot respects standard robots.txt directives in general search crawls that are targeted at Applebot."

Control Tokens (Not Crawlers)

These two tokens never appear as their own identity in a server log. They're robots.txt-only controls layered on top of crawl traffic that already happened under a different, ordinary user-agent.

TokenOperatorWhat It Controlsrobots.txt Behavior
Google-ExtendedGoogleWhether content Google has already crawled under an ordinary Google user agent may be used for training future Gemini models, and for Gemini/Vertex AI grounding (surfacing that content to the model at prompt time). Does not affect Search inclusion or ranking.Control token only - see the section below for the exact scope.
Applebot-ExtendedAppleWhether content Applebot has already crawled may be used to train Apple's generative AI models. Apple states it directly: "Applebot-Extended does not crawl webpages." Publishers can disallow it while remaining indexed by Applebot itself.Control token only.

Third-Party-Attributed (Not Provider-Verified)

This entry is presented separately from the table above on purpose. Everything in the Core Agents table is confirmed against the operator's own published documentation. Bytespider is not - no official ByteDance crawler documentation could be located, so what follows is vendor-observed, not provider-verified, and should be weighted accordingly.

BotOperatorUser-Agent (vendor-reported)What It's Reported to Dorobots.txt Tokenrobots.txt Behavior
BytespiderByteDanceBytespider (+http://www.bytedance.com) - sourced to a third-party bot-management vendor (DataDome), not ByteDance's own docsPer DataDome: "collects web data to enhance search functionalities and content recommendations" across ByteDance's platforms. Cloudflare's own traffic classification (used in our pillar page) buckets Bytespider under training-purpose crawling - the two framings aren't necessarily contradictory, just different lenses (stated use vs. observed classification). DataDome also rates its own identification confidence for this bot as "medium," not high.BytespiderNot respected - per DataDome's own bot-management data, Bytespider does not honor robots.txt directives.

A Bytespider claim in your logs cannot be authenticated against a provider-published IP range the way GPTBot or Googlebot can - see the verification section below. Treat it as an unverified claim of identity, not a confirmed one.

Adjacent and Product-Specific Agents (Appendix)

These agents are real and provider-documented, but narrower in scope than the core table above - either limited to a specific product (ad review, a research tool) or, in Common Crawl's and Amazon's case, not AI-answer-engine crawlers in the same sense as the rest of this page. Listed here rather than the main table so the core table stays focused on what actually moves AI search visibility.

BotOperatorUser-AgentWhat It Doesrobots.txt Token
OAI-AdsBotOpenAI...compatible; OAI-AdsBot/1.0; +https://openai.com/adsbotValidates the safety of web pages submitted as ads on ChatGPT - not a general content or training crawlerOAI-AdsBot
Google-GeminiNotebookGoogleDocumented as a user-triggered fetcher; specific example string not itemized in Google's fetcher overview. Originally added as Google-NotebookLM (October 2025), renamed Google-GeminiNotebook July 16, 2026 - Google's own changelog notes "the old value will continue to be supported" during the transitionFetches URLs a user has specifically supplied as a source in a Gemini Notebook (formerly NotebookLM) research projectGenerally ignores robots.txt, per Google's blanket statement for user-triggered fetchers
CCBotCommon Crawl (nonprofit)CCBot/2.0 (https://commoncrawl.org/faq/)Builds the open, publicly downloadable Common Crawl dataset - not an AI-answer-engine crawler itself, but widely used as training data by third partiesCCBot
AmazonbotAmazonMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36Improves Amazon's products and services and may train Amazon AI models; honors robots.txt, rel=nofollow, and meta robots tags, but not crawl-delayAmazonbot

One more agent, Google-CloudVertexBot, is real and reported in current search-industry coverage as governing crawls that site owners specifically request when building Vertex AI Agents, with no effect on Search or other Google products. It's excluded from the table above because this pass couldn't independently confirm a stable first-party Google documentation page for it the way every other row on this page is sourced - worth checking Google's crawlers-and-fetchers documentation directly if it's relevant to your stack.

The Google-Extended Exception

Google-Extended has no separate HTTP request user-agent string of its own. Google's own documentation states it plainly: Google-Extended "doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity." In practice, that means you'll never see "Google-Extended" in a log line - the requests still arrive under whatever ordinary Google user agent did the crawling, not necessarily Googlebot specifically. The Google-Extended token in robots.txt exists purely so a site owner can opt out of having that already-crawled content used for two related but distinct purposes: "training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini," and separately, "providing content from the Google Search index to the model at prompt time to improve factuality and relevancy" - grounding, not training. Google states directly that "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." If a UA-reference list shows Google-Extended crawling your site under its own identity, or collapses the training/grounding distinction into a single "AI opt-out," that list is imprecise.

How to Actually Verify a Hit Is Real

A user-agent string is just a text field the requester sends - anyone can put "GPTBot" in it. Matching on the string alone tells you what a request claims to be, not what it is. Each operator publishes a way to check the source IP instead. Treat these as living endpoints to query at audit time, not values to hard-code into a permanent list - an operator can rotate its ranges without notice:

  • OpenAI: published IP ranges at openai.com/gptbot.json, openai.com/chatgpt-user.json, and openai.com/searchbot.json.
  • Anthropic: a combined IP list at claude.com/crawling/bots.json.
  • Perplexity: published lists at perplexity.com/perplexitybot.json and perplexity.com/perplexity-user.json.
  • Meta: no JSON endpoint - Meta's own docs describe checking whether the source IP falls within AS32934 (Meta's autonomous system number), noting the range "changes often," and point to Meta's Peering page for current data.
  • Google: JSON IP-range files at developers.google.com/static/crawling/ipranges/common-crawlers.json (Googlebot and other common crawlers), a separate file for special-case crawlers like AdsBot, and, for the newer user-triggered fetchers, dedicated files including user-triggered-fetchers.json and - specifically for Google-Agent - user-triggered-agents.json, all in CIDR format.
  • Apple: reverse-DNS lookups resolving to the *.applebot.apple.com domain, or a published JSON CIDR file at search.developer.apple.com/applebot.json.
  • Common Crawl (CCBot): reverse-DNS lookups resolving to *.crawl.commoncrawl.org, plus a published JSON range list at index.commoncrawl.org/ccbot.json.
  • ByteDance: no published IP-verification method was located for Bytespider. Treat any traffic claiming to be Bytespider as unverifiable by IP until ByteDance publishes one - string-matching alone is not a security or measurement control for this bot.

One distinction worth being precise about: sending a request with a spoofed or copied user-agent string and seeing what a server returns tests what content that server serves to a claimed identity - it does not prove a specific historical log entry was actually the provider. IP-range matching against the published JSON files above is what establishes identity; a user-agent string, spoofed or genuine, never does on its own. Where no provider authentication source exists - Bytespider today - label that traffic "claimed" or "unverified" in any reporting, not as confirmed crawler traffic.

The step-by-step version of this - pulling logs, filtering, cross-checking IPs, confirming what a server actually serves - is on the self-test page. This page is the lookup table you'll reference while you're doing it.

What to Do Next

← Back to the full pillar page
Which of these bots actually matters most for your traffic? →
Use this table to run the log-based self-test →


Sources: GPTBot, ChatGPT-User, OAI-SearchBot, and OAI-AdsBot strings and robots.txt notes from OpenAI's official bot documentation and OpenAI's Publishers and Developers FAQ. ClaudeBot, Claude-User, Claude-SearchBot names, purposes, and robots.txt compliance from Anthropic's own crawler documentation - exact version string not published there. PerplexityBot and Perplexity-User from Perplexity's official crawler documentation. Meta-ExternalAgent, Meta-WebIndexer, and Meta-ExternalFetcher strings, purposes, and IP-verification method from Meta's developer documentation. Bytespider string, stated purpose, verification status, and robots.txt behavior from DataDome - a third-party bot-management vendor, used because no official ByteDance crawler documentation could be located; treat with correspondingly less certainty than the first-party sources above. Googlebot user-agent strings and IP verification from Google's crawler documentation and Google's Googlebot-verification documentation. Google-Extended's exact scope, and Google-Agent's string, purpose, and IP file, from Google's common-crawlers documentation and Google's user-triggered-fetchers documentation (Google-Agent added to that documentation March 20, 2026 and the NotebookLM fetcher renamed Google-GeminiNotebook July 16, 2026, per Google's own crawling-docs changelog). Applebot and Applebot-Extended strings, rendering behavior, robots.txt compliance, and IP verification from Apple's official Applebot documentation. CCBot string, purpose, and verification method from Common Crawl's own CCBot page. Amazonbot string, purpose, and robots.txt handling from Amazon's official Amazonbot page.

About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.