Picture next quarter's board meeting. A director who just watched her own buying behavior change asks the question that is coming for every revenue team: when buyers ask AI about our category, how often are we the answer? Around most tables, silence. An anecdote, maybe. A guess dressed as an estimate.
There is a rigorous way to answer, and it starts with naming the metric. Citation share: of the answers engines give to your buyers' real questions, the fraction that cites you. Beside it sit two companions. Share of voice: how often you are mentioned at all, relative to your peer set. Sentiment: whether the citing answer recommends you, hedges, or warns. Simple to say. Easy to measure badly. Here is how to measure it well.
Build the prompt set from real buying questions
The measurement is only as honest as the questions, so the prompt set cannot come from a keyword tool or a brainstorm. It comes from buying language: questions asked on first sales calls, phrases pulled from closed-lost notes, objections your champions field internally, support tickets, community threads. Not vanity prompts like who leads the category, but the awkward, specific questions evaluations actually run on. Integration constraints. Security posture. How pricing behaves at scale. Direct comparisons against the alternatives you actually face.
Then treat the prompt set like code. Version it. Log every addition and retirement with a date and a reason. When citation share moves, the first hard question will be whether reality changed or the questions did, and only a versioned set can answer. An unversioned prompt set produces trends you cannot defend in front of a skeptical CFO.
Measure by engine, and read the sentiment
ChatGPT, Claude, Perplexity, and Google AI Overviews are different systems with different retrieval habits, different source preferences, and different refresh rhythms. An aggregate number across them hides exactly what you need to know: where you are winning, where you are absent, and which engines your segment actually uses. Report citation share and share of voice per engine, per topic cluster, and let the differences direct the work.
Then read the answers, not just the tallies. A citation that says you are frequently criticized for slow onboarding counts toward citation share and works against you in every deal it touches. A coarse sentiment read per answer, positive, neutral, negative, turns a visibility metric into a reputation metric. Being cited is table stakes. Being cited well is the objective.
The Snapshot Fallacy
Here is where most first attempts fail. Answer engines are probabilistic: the same prompt, on the same engine, on the same day, can return different answers with different citations. A single-day measurement is not a baseline. It is one draw from a distribution, and a decision made on it is a decision made on noise. We call this The Snapshot Fallacy, and it is why a one-time AI visibility audit is a photograph of weather sold as a climate report.
The cure is cadence. Weekly runs. Multiple samples per prompt. Results tracked as rolling windows rather than points, so movement means something before anyone reacts to it. A snapshot is weather. The cadence is climate. Boards should be shown climate.
Tie movement to pipeline without lying
Now the RevOps question: does any of this connect to revenue? Yes, if you refuse to overclaim.
Three defensible instruments. First, correlation windows: log when citation share moves on a topic cluster, then watch qualified conversations, demo requests, and self-reported attribution over the following weeks against the preceding baseline. Second, self-reported attribution done properly: found us through an AI assistant as an explicit option on every form, asked out loud on every first call, never buried under other. Third, holdout thinking: leave some topic clusters or segments deliberately unworked and compare trajectories over time. None of this is a clinical trial, and you should say so out loud. It is disciplined observation, which beats confident fiction every quarter.
What you must not do is claim causal revenue credit from citation share alone. It is a leading indicator of presence in decisions you cannot otherwise observe, and it should be presented with the humility of brand search volume: directionally vital, causally modest. Underclaiming is what makes the number durable. It is how a Series D identity verification company we work with could state, rather than feel, that it moved from trailing to the most-cited leader on its highest-value topics across ChatGPT, Perplexity, and Google AI Overviews in a single quarter. Versioned prompts, weekly cadence, methodology written before the first run. The claim survived scrutiny because the method preceded the result.
That is the standard we hold AI search optimization to, and it is the point of the whole exercise: one number the board trusts, refreshed on a cadence, with the methodology one click away.
When the board finally asks the question, and it will, would you rather offer an anecdote or a number you can defend?
Get your defensible baseline with the AI Search Diagnostic, or start the conversation with Danton.
