In one sentence: share of voice in AI answers is only meaningful if every brand the answers named is counted the same way, on the same questions, in the same window.
The metric is easy to compute and easy to corrupt. Every mistake below produces a number that looks reasonable, moves in the direction leadership hopes for, and cannot be spotted by looking at the result.
Mistake one: counting your brand generously and competitors strictly
The common implementation matches your brand loosely, catching spelling variants and product lines, while matching competitors only against a fixed list. Every brand the assistant named that is not on the list falls out of the denominator, so your share rises without anything improving.
The rule: extract every brand named in the answer, count them all, and use your competitor list only to highlight rows in the report. Report the total number of distinct brands next to your share, because being 1% of 117 brands is a completely different situation from 1% of five.
Mistake two: counting questions that contain your brand name
Ask an assistant whether a named company is any good and it will of course name that company. These self-referential questions cannot measure whether a brand surfaces on its own, but they raise mention rate and share of voice alike.
We have seen a real case where a reported 13% visibility came entirely from one brand-named question, and the true unprompted figure was zero. Exclude them at generation time and again at scoring time, and keep them in a separate reputation view.
Mistake three: treating a failed answer as a zero
When a model times out or a request fails, no evidence was produced in either direction. Counting that cell as a non-mention drags the percentage down, makes a normal month look like a decline, and can fire an alert about a drop that never happened.
Keep unmeasured cells null all the way to the interface, render them as a dash, and report attempted and completed counts side by side so a poor completion rate is visible rather than absorbed into the metric.
Mistake four: comparing across a changed ruler
Turning web search on, adding a model, or extending the question set all change the measurement, not the brand. In one real case, enabling retrieval moved a site's mention rate from 3% to 13% while the site did nothing at all.
Stamp every stored run with a methodology version and refuse to compare across versions. If you must change the ruler, either re-measure history on the new one or mark the break in the chart.
Two things worth reporting alongside
Report share of voice per engine rather than as a cross-model average, because the models differ enough that an average describes no real situation. And report how many distinct questions your brand appeared in, not only the percentage: appearing once in each of ten questions is a stronger position than appearing ten times in one.
A minimum honest report
A share-of-voice report that cannot be misread contains six things: the question set and its version, the engines and whether retrieval was enabled, attempted and completed answer counts, the total number of distinct brands named, your share and rank among all of them, and the date. Anything shorter is a number without a ruler.