Short answer: pick a fixed set of buyer questions that contain no brand names, run them on a fixed set of engines at a fixed interval, and detect every brand in your competitor set with the same code path. Everything else is refinement; skip any of those four and the number is decorative.

Step 1 — Build the question set

Questions should be the decisions buyers actually ask AI to help with: comparisons, recommendations, "which X for Y", "what should I use if". Broad category questions measure awareness; comparison and purchase questions expose commercial gaps.

The trap is self-referential questions. If the question names your brand, the answer will too. An LLM asked to generate a prompt set drifts toward these on its own — ours once produced fifteen questions of which thirteen contained the brand name, and the resulting score read 13% across five engines when the true value was zero. Filter them at generation time, not just at scoring time, or you burn the quota on questions you are going to discard.

Step 2 — Fix the engines and the cadence

Record which engines ran, every time. Adding an engine is a methodology change and the series restarts. Same interval, same time of day, and anchor results to the run's start timestamp rather than when the job finished — otherwise a slow run lands in the wrong period and your trend develops a kink nobody can explain.

Step 3 — Detect symmetrically

One pass, one code path, every brand in the set including yours. Ties resolve against you. Normalise names consistently: if "Acme" and "Acme Systems" fold into one entity for you, they fold for the competitor too.

Step 4 — Report mention and recommendation separately

An answer saying "Acme is often criticised for X, so most buyers choose Y" contains your brand. Whether that counts depends on which question the number answers. Two columns, always.

Step 5 — Keep the evidence

Scores summarise; evidence explains what to do. Store the prompt, the answer, the engine, the market, the cited URLs and the competitor context for every result.

Without the answer text you cannot tell a detection bug from a real change, and in this category you will have both.