Every day, a small fleet of automated visitors reads the public web on behalf of AI systems. Some gather material for training; others fetch pages for search, retrieval or user-requested access. The independently published PTODA C01 study analysed 2,699 sampled rows across five national cohorts, representing 2,652 distinct domains against 21 crawler identities. The research is published through the independent PTODA research programme. Digital Dominator applies related crawler-access analysis commercially but does not own the research findings.
This article is the field guide. Who the 21 are, what each one actually does, and which ones decide whether your business can appear in an AI answer.
The distinction that matters more than any name
AI crawlers do two fundamentally different jobs, and your robots.txt can treat them differently.
Training crawlers collect content that may be used to train future models. GPTBot is the famous one. Blocking a training crawler is a legitimate business decision about how your content is used - and it has no effect on whether today's AI assistants can cite you.
Retrieval crawlers fetch pages at answer time, so an assistant can read current information and cite its source. OAI-SearchBot, Claude-SearchBot and PerplexityBot live here. Block these and you are choosing invisibility: the assistant literally cannot read you, so it cites someone who allowed it.
A third group, on-demand user agents like ChatGPT-User, Claude-User and Perplexity-User, fetch a page because a human asked about it in a conversation right now. Blocking these tells an interested customer's assistant "no".
In our research, a large share of sites blocking "AI" had blocked all three kinds with one rule - usually a copy-pasted template - while intending only to opt out of training. That is the accident we see most.
The OpenAI family
GPTBot - training. The most-blocked AI crawler on the web, and the one most robots templates target. OAI-SearchBot - retrieval for ChatGPT search; this is the one that determines whether ChatGPT can cite you. ChatGPT-User - on-demand fetches during conversations. Three agents, three different decisions - and OpenAI honours each separately.
The Anthropic family
ClaudeBot - training. Claude-SearchBot - retrieval for Claude's search features. Claude-User - on-demand fetches when someone asks Claude about your site. The older tokens anthropic-ai and Claude-Web still appear in many robots files; they are legacy names, but grading them tells you how old your template is.
Perplexity
PerplexityBot - retrieval; Perplexity is an answer engine, so this is its lifeblood. Perplexity-User - on-demand. Perplexity cites aggressively, which makes it one of the easiest places for a well-structured site to earn visible citations - if the crawler can get in.
Google and Microsoft: the special cases
Googlebot is the crawler almost nobody blocks - it feeds classic search and AI Overviews and AI Mode. You cannot block Google's AI features without leaving search entirely. What Google offers instead is Google-Extended: a control token that opts your content out of Gemini training without touching your rankings. Blocking Google-Extended is the training opt-out; blocking Googlebot is self-removal from Google.
Bingbot is the same story for Microsoft: it feeds Bing search and Copilot together. Block it and you leave both.
The training-heavy long tail
CCBot (Common Crawl) supplies the public dataset many models train on - blocking it is the broadest single training opt-out that exists. Bytespider (ByteDance) has a reputation for ignoring robots directives, which is worth knowing when you interpret your logs. Applebot-Extended is Apple's training control token; meta-externalagent and FacebookBot cover Meta's AI collection; Amazonbot feeds Alexa and Rufus; MistralAI-User and DuckAssistBot serve Mistral's Le Chat and DuckDuckGo's answers respectively.
What your robots.txt is probably doing right now
Across PTODA’s five national cohorts, whole-site exclusion of AI retrieval crawlers remained uncommon at 3.5–5.0% of policy-observed domains, while restriction reaching primary public content remained a minority position at 6.7–11.7%. Policy is still only one evidence class: firewalls and CDN bot protection can alter what an identified crawler actually receives.
Which is why the practical question is never "should websites block AI?" It is: what is my site saying to each of these 21, and does my infrastructure agree?
How to find out
You can read your own robots.txt today - it is at yoursite.com/robots.txt - and check it against the names above. That covers policy. It will not show you what your firewall does to these crawlers when they actually connect, and it will not tell you whether the blocking you find is deliberate or inherited. For the evidence-grade answer, our AI Access Audit grades all 21 identities at both the policy and enforcement layers and hands back the exact fixes. We wrote up how the audit works here, and the wider context of ranking in AI systems in our guide to AI search.