Playbooks

AI visibility KPIs for executive reporting

Updated August 1, 2026

Five KPI families make AI visibility legible to an executive team. Share of answers: the percentage of unbranded buying questions where AI engines mention you. Recommendation rate: how often the engine actively recommends you rather than merely listing you. Citation share: how often your domain is among the sources the engine cites. Trend deltas: the movement in each of these against last period and against named competitors. Revenue proxies: AI referral sessions and their conversion rate, which Semrush measured at 4.4 times standard organic in 2026. Reachroller reports the first four from repeated sampled runs with raw-answer receipts, so the number an executive sees can survive the question that follows it.

Why AI visibility needs its own KPI language

Executives already have a language for search: rankings, traffic, conversion rate. AI visibility breaks that language in two places. First, the win event changed: being named inside a composed answer replaces holding a position on a results page, and a name often produces no click at all. Second, the measurement changed: answers are probabilistic, so every KPI is a rate over repeated samples rather than a position you check once. Report AI visibility in ranking language and you will either mislead the room or lose it.

The stakes justify the new vocabulary. G2 found 51 percent of B2B software buyers start research with an AI chatbot more often than Google, and 33 percent bought from a brand they had never heard of before an AI named it. Forrester's 2026 survey of 18,000 buyers found 55 percent compared vendors inside AI tools during their latest purchase. The conversations these KPIs measure are the conversations that build shortlists now.

This playbook defines the five KPI families, gives the formula and cadence for each, and closes with the reporting mistakes that get GEO programs defunded. It assumes your measurement itself is honest; if you have not yet pressure-tested how the underlying score is computed, read what an AI visibility score actually measures first, because no KPI survives a rotten data layer.

KPI one: share of answers

Share of answers is the headline: the percentage of sampled AI answers to your tracked unbranded buying questions that mention your brand. If you track 25 questions, sample each 8 times in a month across engines, and 58 of the 200 answers name you, your share of answers is 29 percent. It is the closest thing this channel has to market presence, and it is the number that belongs at the top of the executive slide.

Two construction rules keep it honest. The question set must be unbranded, because questions containing your name produce mentions by construction and inflate the rate; Reachroller excludes branded questions from the headline number for this reason. And the rate must come from repeated runs, because single checks measure luck. The volatility is documented: SparkToro found under a 1 percent chance that two identical ChatGPT runs return the same brand list, which is why measuring AI visibility is a sampling discipline rather than a lookup.

Keep the question set stable and versioned, because the KPI is only comparable across periods if the denominator holds still. Questions should change when the product or market changes, and each change belongs in the report notes for the period it happened. A question set that quietly sheds its losing questions manufactures an uptrend, and an executive who discovers that once will discount every chart that follows.

Present it segmented before blended. Per-engine share of answers matters because engines disagree with each other about sources and brands; a 40 percent share on ChatGPT and 5 percent on Perplexity is a strategy insight that a blended 24 percent hides. Per-intent segmentation earns its place too: winning how-to questions while losing best-of questions points the content plan at a specific gap.

KPI two: recommendation rate

A mention is attendance; a recommendation is the trophy. Modern engine answers editorialize, especially on best-of and comparison questions: they list several options, then steer, with phrasing like "for a small team, X is the strongest choice". Recommendation rate measures how often, among answers that mention you, the engine puts you in that steering position. A brand mentioned in 60 percent of answers but recommended in 5 percent of them has a positioning problem wearing a visibility costume.

The KPI matters because buyers act on the steer. The mechanics of how engines pick their recommended option, and what shifts it, are unpacked in how ChatGPT recommends brands. For reporting purposes the operational rule is simpler: score it from stored raw answers with a consistent rubric, so that "recommended" means the same thing in March as it does in August. A rubric that drifts produces a KPI that flatters.

Executives should see recommendation rate next to share of answers as a two-by-two: invisible, listed but not chosen, chosen where listed, or dominant. Each quadrant implies a different quarter of work, from earning any presence at all to converting presence into the verdict.

KPI three: citation share

Citation share tracks how often your domain appears among the sources an engine cites, across your sampled answers. It is the supply-side KPI: it measures whether your content is feeding the answers, rather than whether the answers name you. The two decouple more than most teams expect. McKinsey found a brand's own website accounts for only 5 to 10 percent of the sources AI platforms reference, and University of Toronto research found 91 percent of AI answers cite third-party content. You can be mentioned entirely on the strength of other people's pages, and cited without being named.

Track citation share in two rings. The inner ring is your own domain: answers citing pages you control, which your content team can move directly. The outer ring is earned coverage: answers citing third-party pages that mention you favorably, a review profile, a comparison listicle, a community thread. The outer ring usually dwarfs the inner one, and growing it is PR work rather than publishing work, which affects whose quarterly goals the KPI lands in. Reporting the rings separately keeps that ownership question answered before it is asked.

This is why citation share belongs in the executive set even though it sounds tactical: it predicts durability. A mention grounded in your own cited page is repeatable; a mention grounded in one Reddit thread can vanish when retrieval shifts. Tracking which third-party domains carry your losses also converts the KPI into a work queue, since earning presence on those domains is the fix. The distinction and its strategy implications are covered in AI mentions vs citations.

KPI four: trend deltas and competitive share of voice

A rate without a delta is a fact; a delta is a story. Every KPI above should ship with two comparisons: against last period, and against named competitors on the same question set. The competitive frame matters because category mention share is concentrating: a March 2026 U.S. brand visibility report found the top three brands in a category capture 68 percent of AI mentions, up from 54 percent in Q3 2025. Standing still in a concentrating market is losing slowly.

Competitive share of voice, your mentions divided by all mentions of tracked brands, is the cleanest single expression of the trend. It also disciplines interpretation of your own movement: if your share of answers rose from 20 to 25 percent while the category leader rose from 40 to 60, the honest slide says both things. The metric's construction details and its traps are covered in AI share of voice.

One presentation rule saves credibility: attach cause markers to the trend line. Mark the dates fixes were published, pages got indexed, or a PR placement landed. A delta with annotated causes reads as a program; a bare delta reads as weather.

KPI five: revenue proxies

Executives fund pipeline, so the KPI set must end in money, and the honest way to get there is proxies rather than invented attribution. The available ones: AI referral sessions from your analytics, segmented by source (chatgpt.com, perplexity.ai, gemini.google.com and peers), their conversion rate against organic, and the share of signups or demos they produce. Branded search volume lift after visibility gains is a softer fourth, since many buyers hear a name in an answer and search it rather than clicking.

The published benchmarks make the proxy case persuasive on their own. Semrush measured AI-driven visitors converting at 4.4 times standard organic. Ahrefs found AI referrals were 0.5 percent of sessions but 12.1 percent of signups. Adobe Digital Insights measured AI assistant visitors converting 42 percent better than non-AI traffic in March 2026, and Semrush clickstream data shows ChatGPT referral traffic grew 206 percent year over year. The full dataset, with the caveats, lives in AI traffic conversion data.

Revenue proxy benchmarkFindingSource
AI visitor conversion rate4.4x the rate of standard organic trafficSemrush, 2026
Share of signups from AI referrals0.5% of sessions drove 12.1% of signupsAhrefs traffic analysis
Ecommerce conversion vs non-AI traffic42% better, March 2026Adobe Digital Insights, Q1 2026
ChatGPT referral growth206% year over year, Jan 2025 to Jan 2026Semrush clickstream study
ChatGPT referral vs non-branded organic1.81% vs 1.39% conversion across 94 ecommerce brandsVisibility Labs, 2025

Use these as context lines under your own analytics numbers, never as substitutes for them.

The KPI table, ready to steal

KPIFormulaCadenceExecutive question it answers
Share of answersAnswers mentioning brand / total sampled answersWeekly measure, monthly reportDo we exist when buyers ask?
Recommendation rateAnswers recommending brand / answers mentioning brandMonthlyWhen we appear, do we win the verdict?
Citation shareAnswers citing our domain / total sampled answersMonthlyDoes our content feed the answers?
Competitive share of voiceOur mentions / all tracked-brand mentionsMonthlyAre we gaining on the category?
Trend deltaCurrent period rate minus prior period rate, per KPIEvery reportIs the program working?
AI referral conversionConversions from AI referrals / AI referral sessionsMonthly, from analyticsDoes visibility become pipeline?

All rates computed over repeated sampled runs on unbranded questions, per engine first, blended second.

Setting targets executives will accept

A KPI without a target invites the question "so is that good?", and the honest answer starts from how empty the field is. Victorious found 89 percent of brands never appear in AI answers to category research questions, and Wellows measured over 73 percent of brands with zero AI mentions despite page-one Google rankings. A brand starting from the normal baseline of near zero should target tiers rather than perfection: first sustained presence on any tracked question, then a double-digit share of answers, then parity with a named competitor. The cross-industry tier data is laid out in AI visibility benchmarks.

Time-box the first commitment to a quarter, structured as instrument, move, prove. Month one establishes the measured baseline and the competitive gap. Months two and three publish fixes against the ten most winnable lost questions and report the per-question flips. That cadence produces a defensible slide at week thirteen: baseline 4 percent, current 11 percent, six questions flipped, each flip dated against its shipped fix. Small numbers with visible mechanics beat large numbers with none.

Set targets per engine, and resist averaging them. Category concentration is rising, a March 2026 U.S. brand visibility report found the top three brands capture 68 percent of category mentions, so the realistic goal on a leader-dominated engine may be presence on niche intents while another engine offers a path to the shortlist. One line of context for the CFO closes the section: the buyers these rates describe convert at multiples of organic when they do click through, which is why the channel deserves a target at all.

What gets AI visibility programs defunded

Screenshot theater. A slide with one glorious ChatGPT answer is the fastest way to lose the room three months later, when someone reruns the prompt and gets a different answer. Given sub-1-percent run-to-run consistency, single answers prove nothing in either direction. Report rates, and keep the receipts available for anyone who wants to drill in.

Branded inflation. A score juiced with branded prompts starts high and cannot rise, which reads as a stalled program even when real visibility is improving. Executives forgive a low starting number with a rising trend far more readily than a high flat one they eventually learn was hollow.

Blended engine averages. One cross-engine number smooths a 40 percent ChatGPT share and a 3 percent Perplexity share into a 20 percent average that describes neither. The engines cite different sources and reward different work, with citation analyses finding only about 11 percent of domains cited by both ChatGPT and Perplexity, so the blend erases exactly the information a budget decision needs. Lead with the engine your buyers use most, and show the split on the same slide.

KPIs without a work loop. Numbers that no action can move are trivia. Every lost question in the report should map to a fix: a page published, a source earned, a wrong claim corrected. This is the loop Reachroller automates end to end, tracking the questions, generating the publish-ready fix page for each loss, and rechecking until the answer flips, from $29 per month. The KPI story writes itself when the trend line has cause markers on it.

Frequently asked questions

What is the single most important AI visibility KPI?+

Share of answers on unbranded buying questions. It answers the existential question first: when a buyer asks an AI engine for options in your category, do you exist? Every other KPI qualifies this one. Recommendation rate is meaningless at zero mentions, and citation share explains mentions rather than replacing them.

How is share of answers different from share of voice?+

Share of answers is absolute: the percentage of sampled answers naming you, regardless of who else appears. Share of voice is relative: your mentions as a fraction of all mentions earned by tracked brands in the category. Executives need both, because share of answers can rise while share of voice falls if a competitor grows faster.

Why report recommendation rate separately from mentions?+

Because engines increasingly editorialize. An answer can list six tools and recommend one, and buyers hear the recommendation. Forrester found 55 percent of buyers compared vendors inside AI tools during their most recent purchase, so the verdict inside the answer does shortlist work. Tracking only mentions hides whether you win or merely attend.

What revenue proxies work for AI visibility?+

AI referral sessions segmented by source, their conversion rate against organic, and share of signups attributed to AI referrals. The benchmarks favor the channel: Semrush measured 4.4x organic conversion for AI visitors, and Ahrefs found AI referrals were 0.5 percent of sessions but 12.1 percent of signups. Pair these with mention trends to argue causation honestly.

How often should AI visibility KPIs go to executives?+

Monthly, built on weekly sampled measurement. AI answers are volatile run to run, SparkToro measured under 1 percent consistency between identical ChatGPT runs, so weekly numbers wobble in ways that waste executive attention. A monthly report on repeated samples smooths the noise while staying fast enough to show cause and effect from published fixes.

What baseline should I expect on my first AI visibility KPI report?+

Low, and that is normal rather than alarming. Victorious found 89 percent of brands never appear in AI answers to category research questions, and Wellows measured over 73 percent of page-one Google brands with zero AI mentions. Frame the first report as the instrumented baseline and set quarter-long targets from it, because the trend from an honest low number is the asset.

Can I compute these KPIs manually?+

For a handful of questions, yes: ask each question in fresh sessions several times per engine, log mentions in a spreadsheet, repeat weekly. The workload compounds quickly at real coverage, 25 questions across engines with repeated runs is hundreds of answers monthly. Reachroller automates the loop at $29 per month and keeps the raw answers as receipts.

Sources referenced

  • Semrush, AI visitor conversion analysis, 2026
  • Semrush, ChatGPT clickstream study, 1B+ lines of U.S. clickstream data, October 2024 to February 2026
  • Ahrefs, AI referral traffic and signup share analysis
  • Adobe Digital Insights, Q1 2026 ecommerce AI traffic analysis
  • Visibility Labs, ChatGPT referral conversion study of 94 ecommerce brands, 2025
  • Forrester, 2026 Buyers' Journey Survey (18,000 global business buyers)
  • G2, B2B buyer AI research, 2026
  • SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
  • McKinsey, AI Discovery Survey, August 2025

KPIs that survive the follow-up question

Three days, 50 credits, every feature, no card. Share of answers, per-question receipts, and a trend line with causes on it.

Check my brand free