Design an AI SEO Program · Part 5 of 8
Design the Scoreboard Before the Work
Last reviewed:
Traditional SEO measurement came with the platform: Search Console told you what ranked and what got clicked. AI search measurement mostly does not exist until you build it. That changes its status in the program — measurement is not a reporting task at the end, it is an instrument you design at the start.
What the platforms actually give you
Take inventory before buying anything.
- Microsoft provides the only platform-native citation data: the AI Performance report in Bing Webmaster Tools (public preview since February 2026) shows total citations, cited pages, and the grounding queries behind them, across Copilot and Bing’s AI summaries. Free, first-party, and the calibration reference for everything else.
- Google reports AI Overviews and AI Mode inside overall Search totals in Search Console — included in the “Web” search type, not broken out. You cannot isolate AI-feature clicks or impressions from Google’s own reporting.
- Everyone else offers nothing first-party. Third-party visibility tools fill the gap by sampling prompts against the assistants — which makes their prompt sets, not their dashboards, the thing to evaluate.
The prompt corpus is the instrument
Every AI visibility number is an answer to “when we asked these questions, this happened.” The questions are the measurement instrument, and they deserve instrument discipline: versioned, staged across the buyer journey, split by market and product line, and reviewed on a schedule. An unrepresentative corpus does not produce wrong numbers — worse, it produces plausible ones.
Sample breadth is part of validity. Semrush’s reasoning-mode study found ChatGPT’s two modes cited mostly different domains for identical prompts — with citation rates and source mixes shifting between modes. A corpus that samples one mode on one surface reports a fraction of reality with full confidence. Reading AI Visibility Metrics covers the interpretation traps; the measurement checklist verifies the setup.
The metric hierarchy
Build metrics in layers, cheapest first:
- Presence — are you mentioned at all for corpus prompts?
- Citation — is your URL the source, and which pages earn it?
- Share of voice — you versus named competitors, per surface, per market. This is the headline KPI leadership will actually track.
- Accuracy and sentiment — what do assistants say? Wrong pricing, invented specifications, and repeated complaint themes are visibility incidents, and only you hold the ground truth to detect them. Diff assistant answers against your own product data on a schedule.
- Referral and cohort quality — AI-referred sessions, tagged and tracked through conversion and retention, because the channel’s value claim will eventually need revenue attached.
Report everything on rolling windows with volatility bands. Single-month readings of a non-deterministic system are noise wearing a suit.
The ROI conversation
Here is the uncomfortable mechanism: AI answers satisfy many queries without a click. Pew measured users clicking a result in 8% of visits when an AI summary appeared versus 15% without, and clicking links inside the summary in 1% of visits. The channel’s influence is real and mostly invisible to last-click analytics.
So agree the ROI model with finance before budget season, not after the first disappointing traffic report:
- Blend four evidence types: AI referral revenue (small but attributable), share-of-voice movement against competitors (the visibility claim), answer accuracy (the risk-avoidance claim), and surveyed influence — “how did you research this purchase?” — to size the zero-click effect analytics cannot see.
- Set the counterfactual honestly. The alternative to the program is not “current traffic continues”; it is competitors occupying the answers your buyers read. Share-of-voice versus named competitors captures that; raw traffic does not.
- Pre-commit to the timeline. Tie each metric to the clock it obeys (sequencing guide), so nobody judges model-layer work on a retrieval-layer schedule.
Decision rules
- Stand up free first-party measurement (Bing AI Performance, analytics referral segments, server logs) before evaluating any paid tool — it is your calibration data.
- Buy sampling breadth, not dashboard polish: engine coverage, market and language coverage, corpus control, and export access are the procurement criteria that matter.
- Treat the prompt corpus as owned infrastructure. Tools change; the corpus and metric definitions must survive the change.
- If a metric cannot survive the question “sampled how, on which mode, over what window?” it does not go in front of leadership.