Nostimates

Reference

The AI crawler directory

42 crawlers, indexers and agent fetchers you will see in your logs, with the exact user-agent string, the operator behind it, whether it honours robots.txt, how to verify it is genuine, and what you lose by blocking it. Training crawlers and answer indexers are different things — blocking the wrong one removes you from AI answers entirely.

AI training

16 bots

GPTBot

OpenAI

Crawls public pages to build training corpora for OpenAI foundation models.

honours robots.txt

ClaudeBot

Anthropic

Crawls public content for Anthropic model training.

honours robots.txt

Google-Extended

Google

A robots.txt control token governing whether your content trains Gemini and grounds Vertex AI, without affecting Search.

honours robots.txt

Applebot-Extended

Apple

Robots.txt token controlling whether Apple uses your content to train Apple Intelligence models.

honours robots.txt

meta-externalagent

Meta

Collects public web data to train Meta AI models.

honours robots.txt

Bytespider

ByteDance

Collects training data for ByteDance models including Doubao.

honours robots.txt

TikTokSpider

ByteDance

Crawls web content for TikTok search and recommendation systems.

honours robots.txt

Amazonbot

Amazon

Crawls public content to improve Amazon products and services, and may be used to train Amazon's AI models. (Alexa's on-demand answering uses a separate agent, Amzn-User.)

honours robots.txt

cohere-ai

Cohere

Crawls content for Cohere model training and retrieval products.

honours robots.txt

cohere-training-data-crawler

Cohere

Dedicated training-corpus crawler.

honours robots.txt

CCBot

Common Crawl

Builds the Common Crawl corpus, which underpins a large share of every major model's training data.

honours robots.txt

Omgilibot

Webz.io

Collects web data resold as training and market-intelligence datasets.

honours robots.txt

Diffbot

Diffbot

Builds a structured knowledge graph sold to AI and data customers.

honours robots.txt

DeepSeekBot

DeepSeek (unconfirmed)

Third-party bot lists attribute this user agent to DeepSeek, but DeepSeek publishes no official crawler documentation, declared user agent or IP ranges — treat it as unverified.

honours robots.txt

AI2Bot

Allen Institute for AI

Collects openly licensed content for open research datasets.

honours robots.txt

Timpibot

Timpi

Builds a decentralised index sold for AI use.

honours robots.txt