Short answer: for Gemini there is no crawler to allow and no log line to check. The lever is Google-Extended, a usage control, and everything else you do is ordinary extraction work on pages Google already has.

Declared, not fetched

Google documents Google-Extended as determining whether already-crawled content may be used to train Gemini models and to ground answers in Gemini Apps and Vertex AI, while explicitly not affecting Google Search inclusion or ranking. Because it controls usage rather than fetching, no request appears in your logs when it is honoured. There is nothing to verify on your side.

The practical consequence: record the decision and its reason next to the rule in robots.txt. A year from now, whoever reads that file will not remember which line was deliberate. Then measure the effect where it is actually observable — mention and citation rates in Gemini over a fixed question set, compared against engines where no such control was applied.

A separate agent worth knowing about

Google also documents Google-CloudVertexBot, which crawls sites at the site owner's own request when building Vertex AI Agents. It is only relevant if you are building an agent on your own content — it is not a general retrieval agent, and blocking it costs you nothing otherwise.

What moves Gemini answers

Since the index is shared with Search, the optimisation surface is extraction rather than access:

  • Headings phrased as the questions people ask, one question each.
  • The answer in the first sentence under the heading, naming its subject.
  • Structured data for facts you want stated exactly — it is still the cheapest way to make a fact survive extraction.
  • Explicit dates, because grounding favours content that can be placed in time.
With Gemini you control the declaration and the page. You do not control, and cannot inspect, what happens in between — so build the measurement on the answer side or you are flying blind.