Short answer: OpenAI documents GPTBot for training-corpus crawling, OAI-SearchBot for search indexing, and ChatGPT-User for fetches triggered by a person's request. Blocking the first is a defensible policy position. Blocking the second removes you from ChatGPT's search results. A wildcard aimed at "OpenAI" does both.

The mistake, precisely

A site decides it does not want its content used for model training. Reasonable. It then writes one broad rule, and quietly deletes itself from the surface where buyers are asking questions today. The cost is invisible because nothing breaks — you simply stop appearing, and there is no error to notice.

Name each agent explicitly in robots.txt. Never use a wildcard on the operator, and write the reason next to the rule.

There is also OAI-AdsBot for ad safety review, which is unrelated to whether you appear in answers.

Verify by fetching, not by reading

Reading your own robots.txt tells you what you intended. It does not tell you what your server does. A CDN rule or a WAF that blocks by user agent will return 403 to a named agent while robots.txt says "Allow" — and a robots.txt reader will report everything as fine. Request your own pages once per user agent and compare status codes and served HTML.

Then the extraction half

Access gets you considered. Extraction gets you used. ChatGPT assembles replies from passages, so:

  • One question per heading, in the words a person would use.
  • Answer in the first sentence under it, subject named, no dependency on the heading above.
  • Facts in tables with units and dates; hedges cannot be quoted.
  • Nothing important behind an accordion, tab or modal — if it isn't in the server-rendered HTML, it does not exist.

What not to measure

There is no rank in a ChatGPT answer. Tools that report one are usually reporting the order brands happened to appear in a single sampled answer, which is not stable enough to trend. Measure presence — mentioned, recommended, cited — on a fixed question set, and report the three separately.