The list the AI Visibility Site Audit evaluates your robots.txt against.
Registry updated 2026-08-06 — 21 entries, 6 of them counted in the Access score.
robots.txt files still reference them.robots.txt.| Token | Operator | Type | Scored | Documented behavior | Official source | Verified |
|---|---|---|---|---|---|---|
Amazonbot |
Amazon | Training | no | Used to improve Amazon products and services and may train Amazon AI models; respects robots.txt and honors noarchive/noindex meta tags. | developer.amazon.com | 2026-08-06 |
Amzn-SearchBot |
Amazon | Search / retrieval | yes | Makes content eligible for Alexa search experiences; respects robots.txt; officially states it does not crawl for generative-AI training. | developer.amazon.com | 2026-08-06 |
Amzn-User |
Amazon | User-triggered fetch | no | Supports user actions such as Alexa queries needing fresh information; may not follow all robots.txt directives; not used for generative-AI training. | developer.amazon.com | 2026-08-06 |
Claude-SearchBot |
Anthropic | Search / retrieval | yes | Crawls to improve search result quality; respects robots.txt. | support.claude.com | 2026-08-06 |
Claude-User |
Anthropic | User-triggered fetch | no | Fetches pages when Claude users ask about them; Anthropic states robots.txt is respected. | support.claude.com | 2026-08-06 |
ClaudeBot |
Anthropic | Training | no | Respects robots.txt and Crawl-delay. | support.claude.com | 2026-08-06 |
Applebot |
Apple | Search / retrieval | yes | Crawls for Apple products including Siri and Spotlight suggestions; respects robots.txt. | support.apple.com | 2026-07-07 |
Applebot-Extendedrobots.txt token only — no crawler of its own |
Apple | Training | no | robots.txt control token governing use of Applebot-crawled content for Apple AI training (no separate crawler UA). | support.apple.com | 2026-07-07 |
Bytespider |
ByteDance | Training | no | No official documentation page known; robots.txt behavior undocumented. | no official page known | 2026-07-07 |
CCBot |
Common Crawl | Training | no | Open web archive that feeds many AI training sets; respects robots.txt. | commoncrawl.org | 2026-07-07 |
Google-Extendedrobots.txt token only — no crawler of its own |
Training | no | robots.txt control token only (no separate crawler UA); governs use of crawled content for Gemini training; no impact on Google Search inclusion or ranking. | developers.google.com | 2026-08-06 | |
Googlebot |
Search / retrieval | yes | Google Search's main crawler; respects robots.txt. AI experiences built on Google Search (e.g. AI Overviews) draw on its crawl. | developers.google.com | 2026-08-06 | |
meta-externalagent |
Meta | Training | no | Crawls for AI training and related purposes; respects robots.txt per Meta's crawler documentation. | developers.facebook.com | 2026-07-07 |
ChatGPT-User |
OpenAI | User-triggered fetch | no | User-initiated fetches; OpenAI states robots.txt rules may not apply and it is not used for search inclusion. | developers.openai.com | 2026-08-06 |
GPTBot |
OpenAI | Training | no | Respects robots.txt; disallowing it opts your content out of OpenAI model training. | developers.openai.com | 2026-08-06 |
OAI-AdsBot |
OpenAI | Other | no | Validates ad landing pages submitted to ChatGPT ads; not used for training. | developers.openai.com | 2026-08-06 |
OAI-SearchBot |
OpenAI | Search / retrieval | yes | Respects robots.txt; sites opted out are not shown in ChatGPT search answers. | developers.openai.com | 2026-08-06 |
Perplexity-User |
Perplexity | User-triggered fetch | no | Official: since a user requested the fetch, this fetcher generally ignores robots.txt rules. | docs.perplexity.ai | 2026-08-06 |
PerplexityBot |
Perplexity | Search / retrieval | yes | Surfaces and links sites in Perplexity search; respects robots.txt; not used for foundation-model training. | docs.perplexity.ai | 2026-08-06 |
anthropic-ai |
Anthropic | Legacy token | no | Token absent from Anthropic's current official documentation (checked 2026-08-06); shown because existing robots.txt files still reference it. | support.claude.com | 2026-08-06 |
Claude-Web |
Anthropic | Legacy token | no | Token absent from Anthropic's current official documentation (checked 2026-08-06); shown because existing robots.txt files still reference it. | support.claude.com | 2026-08-06 |