Short answer: AI agents fall into four roles, and the role determines the cost of blocking. Training crawlers cost you nothing in today's answers. Retrieval agents cost you the answers themselves. User-triggered fetchers cost a real person's request. Usage controls do not crawl at all.

The four roles

RoleExamplesCost of blocking
Training corpusGPTBot, ClaudeBot, CCBot, meta-externalagentNothing in today's answers — a defensible policy position
Search / answer indexOAI-SearchBot, PerplexityBot, Claude-SearchBot, ApplebotYou disappear from that engine's answers
User-triggered fetchChatGPT-User, Claude-User, Perplexity-UserA request a real person made fails; several are documented as generally not following robots.txt
Usage controlGoogle-Extended, Applebot-ExtendedBounded and specific — e.g. grounded Gemini answers, with Google Search explicitly unaffected

Why one rule per operator is the wrong shape

Every major operator runs more than one agent, deliberately, so the decisions can be separated. A wildcard on the operator collapses four decisions into one and almost always produces an outcome nobody chose. The most common version: declining training use and removing yourself from search results in the same line.

Two things the table can't tell you

  • Whether your server honours it. A rule is an intent; only a fetch is evidence. CDN and WAF rules routinely override robots.txt for anything that looks like a bot.
  • Whether an undocumented agent exists. Some vendors publish no name you can allow by. Where that is the case the honest answer is to say so — inventing a user-agent name is worse than a gap, because site owners will write rules against a name nothing honours and believe they have allowed something.

Keep the list dated

Vendor documentation changes. Any crawler table without a "checked on" date is a snapshot of an unknown moment, and the one thing worse than no policy is a policy based on last year's names.