Short answer: the file matters less than the discipline around it. Name each agent explicitly, write the reason beside the rule, verify by fetching rather than by reading, and date the whole thing. Most AI-crawler robots.txt files in the wild fail all four.

Why explicit beats wildcard

Every major operator runs several agents on purpose — one for training corpus, one for search indexing, one for user-triggered fetches — so that site owners can decide each separately. A wildcard aimed at the operator collapses those decisions into one, and the result is almost never what anyone chose. The classic outcome is declining training use and deleting yourself from an answer engine in the same line.

Write the reason next to the rule

This sounds like housekeeping and is actually the highest-value habit here. An unexplained Disallow gets copied forward through every migration, forever, because nobody dares remove a line they do not understand. Two words of context prevents a decade of cargo cult.

# Training corpora: declined. Position reviewed 2026-09.
User-agent: GPTBot
Disallow: /

# Answer surfaces: allowed. Blocking these removes us from answers.
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

# Admin surfaces: never indexed anywhere.
User-agent: *
Disallow: /admin/
Disallow: /api/

Sitemap: https://example.com/sitemap.xml

Verify by fetching

A robots.txt reading tells you your intent. Only a request tells you your behaviour. CDN rules and WAFs routinely return 403 to anything that looks like a bot while robots.txt reads as fully permissive — and a robots.txt checker will report everything as fine. Request your own pages once per user agent, compare status codes and served HTML, and treat "could not fetch" as different from "page not present".

Two things not to do

  • Do not invent agent names. Some vendors publish nothing you can allow by name. Writing a rule against a name nothing honours is worse than a gap, because you will believe you have allowed something.
  • Do not block user-triggered fetchers expecting it to work. Several are documented as generally not following robots.txt, because the request came from a person. If you need to stop them, that is an edge-rule decision with a real cost.
Date the file. Vendor documentation changes, and a policy built on last year's agent names is worse than no policy, because it looks current.