Nostimates
← Glossary

Definition

robots.txt

robots.txt is a root-level file that grants or denies crawler access per user-agent token. In the AI era it carries a distinction most sites get wrong: training crawlers and answer-engine indexers are separate agents.

Blocking GPTBot opts you out of OpenAI training but leaves ChatGPT citations intact; blocking OAI-SearchBot removes you from ChatGPT's citable index entirely.

Some agents — notably user-initiated fetchers — state that they are not governed by robots.txt, so enforcement there requires CDN rules.

Nostimates measures robots.txt across twelve AI answer surfaces. Request access →