AI crawler registry

The list the AI Visibility Site Audit evaluates your robots.txt against. Registry updated 2026-08-06 — 21 entries, 6 of them counted in the Access score.

What this page is

What this page is not

TokenOperatorTypeScoredDocumented behaviorOfficial sourceVerified
Amazonbot Amazon Training no Used to improve Amazon products and services and may train Amazon AI models; respects robots.txt and honors noarchive/noindex meta tags. developer.amazon.com 2026-08-06
Amzn-SearchBot Amazon Search / retrieval yes Makes content eligible for Alexa search experiences; respects robots.txt; officially states it does not crawl for generative-AI training. developer.amazon.com 2026-08-06
Amzn-User Amazon User-triggered fetch no Supports user actions such as Alexa queries needing fresh information; may not follow all robots.txt directives; not used for generative-AI training. developer.amazon.com 2026-08-06
Claude-SearchBot Anthropic Search / retrieval yes Crawls to improve search result quality; respects robots.txt. support.claude.com 2026-08-06
Claude-User Anthropic User-triggered fetch no Fetches pages when Claude users ask about them; Anthropic states robots.txt is respected. support.claude.com 2026-08-06
ClaudeBot Anthropic Training no Respects robots.txt and Crawl-delay. support.claude.com 2026-08-06
Applebot Apple Search / retrieval yes Crawls for Apple products including Siri and Spotlight suggestions; respects robots.txt. support.apple.com 2026-07-07
Applebot-Extended
robots.txt token only — no crawler of its own
Apple Training no robots.txt control token governing use of Applebot-crawled content for Apple AI training (no separate crawler UA). support.apple.com 2026-07-07
Bytespider ByteDance Training no No official documentation page known; robots.txt behavior undocumented. no official page known 2026-07-07
CCBot Common Crawl Training no Open web archive that feeds many AI training sets; respects robots.txt. commoncrawl.org 2026-07-07
Google-Extended
robots.txt token only — no crawler of its own
Google Training no robots.txt control token only (no separate crawler UA); governs use of crawled content for Gemini training; no impact on Google Search inclusion or ranking. developers.google.com 2026-08-06
Googlebot Google Search / retrieval yes Google Search's main crawler; respects robots.txt. AI experiences built on Google Search (e.g. AI Overviews) draw on its crawl. developers.google.com 2026-08-06
meta-externalagent Meta Training no Crawls for AI training and related purposes; respects robots.txt per Meta's crawler documentation. developers.facebook.com 2026-07-07
ChatGPT-User OpenAI User-triggered fetch no User-initiated fetches; OpenAI states robots.txt rules may not apply and it is not used for search inclusion. developers.openai.com 2026-08-06
GPTBot OpenAI Training no Respects robots.txt; disallowing it opts your content out of OpenAI model training. developers.openai.com 2026-08-06
OAI-AdsBot OpenAI Other no Validates ad landing pages submitted to ChatGPT ads; not used for training. developers.openai.com 2026-08-06
OAI-SearchBot OpenAI Search / retrieval yes Respects robots.txt; sites opted out are not shown in ChatGPT search answers. developers.openai.com 2026-08-06
Perplexity-User Perplexity User-triggered fetch no Official: since a user requested the fetch, this fetcher generally ignores robots.txt rules. docs.perplexity.ai 2026-08-06
PerplexityBot Perplexity Search / retrieval yes Surfaces and links sites in Perplexity search; respects robots.txt; not used for foundation-model training. docs.perplexity.ai 2026-08-06
anthropic-ai Anthropic Legacy token no Token absent from Anthropic's current official documentation (checked 2026-08-06); shown because existing robots.txt files still reference it. support.claude.com 2026-08-06
Claude-Web Anthropic Legacy token no Token absent from Anthropic's current official documentation (checked 2026-08-06); shown because existing robots.txt files still reference it. support.claude.com 2026-08-06