Short answer: a composite score is a change detector, not a diagnosis. Ours weights mention presence and citation presence into one deterministic number precisely so it moves when something real moves — and we show the components next to it, because the components are what you act on.

What a composite is good for

One thing: telling you to go look. A score that drops eight points this month is a prompt to open the evidence, not a finding in itself. Used that way it earns its place, because nobody reviews forty questions across five engines every week by hand.

Where it stops being useful

The moment someone asks "so what do we fix?" Two brands can share a score with completely opposite problems — one cited everywhere and recommended nowhere, the other recommended in a handful of answers with no links at all. Same number, different work. Any score that cannot be decomposed on screen is a number you have to take on faith.

Three ways scores get inflated

  • Self-referential questions in the set. Questions containing the brand name answer themselves. Thirteen of fifteen questions naming the brand produced a 13% score for us once, where the honest figure was zero.
  • Retrieval pool counted as citations. The candidate pool an engine assembles is much larger than what it cites in the answer body — counting the pool overstated our own citation numbers by two to four times before we fixed it.
  • Unmeasured rendered as zero, or worse, dropped. Dropping failed runs quietly raises the average. Rendering them as 0% lowers it. Either way the number stops meaning what it says; the only honest option is a visible gap.

What to ask a vendor

Show me the question set. Show me which of these runs failed. Show me one answer where I was counted as cited, with the source marker.

Three requests, all cheap to satisfy if the measurement is sound. If any of them produces friction, the score is doing work it cannot support.