Reference
The AI crawler directory
42 crawlers, indexers and agent fetchers you will see in your logs, with the exact user-agent string, the operator behind it, whether it honours robots.txt, how to verify it is genuine, and what you lose by blocking it. Training crawlers and answer indexers are different things — blocking the wrong one removes you from AI answers entirely.
AI agent
6 botsChatGPT-User
OpenAIFetches a page on demand because a user or a ChatGPT action asked for that specific URL.
honours robots.txt
Claude-User
AnthropicFetches a page because a Claude user's request required it.
honours robots.txt
Google-CloudVertexBot
GoogleFetches sites on behalf of Vertex AI customers building grounded agents.
honours robots.txt
Perplexity-User
PerplexityFetches a specific page because a user's query required it.
ignores robots.txt
meta-externalfetcher
MetaFetches specific URLs on behalf of Meta AI user requests.
honours robots.txt
MistralAI-User
Mistral AIFetches pages on behalf of Le Chat users.
honours robots.txt
AI answer
6 botsOAI-SearchBot
OpenAIBuilds the search index ChatGPT uses to cite live sources in answers.
honours robots.txt
Claude-SearchBot
AnthropicCrawls the web to improve the quality of Claude's search-grounded answers.
honours robots.txt
PerplexityBot
PerplexityIndexes pages so Perplexity can cite them.
honours robots.txt
YouBot
You.comIndexes pages for You.com's AI search answers.
honours robots.txt
DuckAssistBot
DuckDuckGoSupports DuckDuckGo's AI-assisted answers.
honours robots.txt
Nostimates canary
NostimatesWe collect AI answers from the engines themselves. We do not crawl customer or third-party websites.
honours robots.txt
AI training
16 botsGPTBot
OpenAICrawls public pages to build training corpora for OpenAI foundation models.
honours robots.txt
ClaudeBot
AnthropicCrawls public content for Anthropic model training.
honours robots.txt
Google-Extended
GoogleA robots.txt control token governing whether your content trains Gemini and grounds Vertex AI, without affecting Search.
honours robots.txt
Applebot-Extended
AppleRobots.txt token controlling whether Apple uses your content to train Apple Intelligence models.
honours robots.txt
meta-externalagent
MetaCollects public web data to train Meta AI models.
honours robots.txt
Bytespider
ByteDanceCollects training data for ByteDance models including Doubao.
honours robots.txt
TikTokSpider
ByteDanceCrawls web content for TikTok search and recommendation systems.
honours robots.txt
Amazonbot
AmazonCrawls public content to improve Amazon products and services, and may be used to train Amazon's AI models. (Alexa's on-demand answering uses a separate agent, Amzn-User.)
honours robots.txt
cohere-ai
CohereCrawls content for Cohere model training and retrieval products.
honours robots.txt
cohere-training-data-crawler
CohereDedicated training-corpus crawler.
honours robots.txt
CCBot
Common CrawlBuilds the Common Crawl corpus, which underpins a large share of every major model's training data.
honours robots.txt
Omgilibot
Webz.ioCollects web data resold as training and market-intelligence datasets.
honours robots.txt
Diffbot
DiffbotBuilds a structured knowledge graph sold to AI and data customers.
honours robots.txt
DeepSeekBot
DeepSeek (unconfirmed)Third-party bot lists attribute this user agent to DeepSeek, but DeepSeek publishes no official crawler documentation, declared user agent or IP ranges — treat it as unverified.
honours robots.txt
AI2Bot
Allen Institute for AICollects openly licensed content for open research datasets.
honours robots.txt
Timpibot
TimpiBuilds a decentralised index sold for AI use.
honours robots.txt
Search
10 botsGooglebot
GoogleGoogle's primary crawler. Feeds classic Search, AI Overviews and AI Mode from the same index.
honours robots.txt
Bingbot
MicrosoftMicrosoft's search crawler, feeding Bing and by extension Copilot answers.
honours robots.txt
MSNBot
MicrosoftLegacy Microsoft crawler, largely superseded by Bingbot but still seen in logs.
honours robots.txt
Applebot
ApplePowers Siri, Spotlight and Safari suggestions.
honours robots.txt
Baiduspider
BaiduBaidu's primary crawler, feeding both classic results and ERNIE AI answers.
honours robots.txt
PetalBot
HuaweiCrawls for Petal Search and Huawei assistant features.
honours robots.txt
YisouSpider
UCWeb / Alibaba (Shenma)Crawls to build the index for Shenma (SM.CN), the mobile search engine from UCWeb and Alibaba.
honours robots.txt
Sogou web spider
Tencent / SogouCrawls for Sogou search and Tencent-adjacent AI features.
honours robots.txt
Naverbot / Yeti
NaverNaver's crawler, feeding Korea's default search and its AI answer surfaces.
honours robots.txt
YandexBot
YandexCrawls for Yandex Search and its Alice assistant answers.
honours robots.txt