Short answer: AI agents fall into four roles, and the role determines the cost of blocking. Training crawlers cost you nothing in today's answers. Retrieval agents cost you the answers themselves. User-triggered fetchers cost a real person's request. Usage controls do not crawl at all.
The four roles
| Role | Examples | Cost of blocking |
|---|---|---|
| Training corpus | GPTBot, ClaudeBot, CCBot, meta-externalagent | Nothing in today's answers — a defensible policy position |
| Search / answer index | OAI-SearchBot, PerplexityBot, Claude-SearchBot, Applebot | You disappear from that engine's answers |
| User-triggered fetch | ChatGPT-User, Claude-User, Perplexity-User | A request a real person made fails; several are documented as generally not following robots.txt |
| Usage control | Google-Extended, Applebot-Extended | Bounded and specific — e.g. grounded Gemini answers, with Google Search explicitly unaffected |
Why one rule per operator is the wrong shape
Every major operator runs more than one agent, deliberately, so the decisions can be separated. A wildcard on the operator collapses four decisions into one and almost always produces an outcome nobody chose. The most common version: declining training use and removing yourself from search results in the same line.
Two things the table can't tell you
- Whether your server honours it. A rule is an intent; only a fetch is evidence. CDN and WAF rules routinely override robots.txt for anything that looks like a bot.
- Whether an undocumented agent exists. Some vendors publish no name you can allow by. Where that is the case the honest answer is to say so — inventing a user-agent name is worse than a gap, because site owners will write rules against a name nothing honours and believe they have allowed something.
Keep the list dated
Vendor documentation changes. Any crawler table without a "checked on" date is a snapshot of an unknown moment, and the one thing worse than no policy is a policy based on last year's names.