Standardize Definitions Before Discussing Numbers

The most common problem in GEO discussions is two parties reading the same data under different definitions. Before looking at any number, confirm four things: what the denominator is, how the numerator is judged, how many samples were taken, and which platforms are covered.

The denominator is usually total sampled prompts, but it can also be valid responses. If some prompts returned no result due to platform errors, counting them in the denominator understates visibility. The numerator depends on judgement rules: does a bare mention of the brand name count, or must it be described positively?

Those two definitional gaps alone can swing a single site's mention rate between twenty and forty-five percent. So the first step before any horizontal comparison is confirming the definitions match.

The Four Core Metrics Defined

Mention rate: the share of sampled answers in which the brand is mentioned in any form. This is the loosest metric and serves as a water-level indicator, reflecting whether AI knows you exist.

Recommendation rate: of those mentions, the share where the brand is directly recommended as an action for the user. Its numerator is always a subset of the mention numerator, so it can never exceed mention rate. A large gap between the two usually means the brand is known but not trusted.

Citation rate: the share where your own pages are cited as source links. This is the strictest metric and the only one that directly generates clicks; it measures whether your content is treated as citable evidence.

Share of Voice: across comparable prompts, your mention count as a share of total mentions across all competitors. It removes industry-wide movement and is the correct metric for relative position.

Read all four together. Mention rate alone overstates results; citation rate alone understates brand-building progress.

What Counts as Healthy

A caveat first: cross-industry absolute benchmarks are of limited use because competitive intensity varies enormously. The ranges below apply to B2B and content sites under moderate competition and are meant to set initial expectations.

Mention rate below ten percent usually indicates a technical access or entity recognition problem, so audit before optimizing content. Ten to thirty percent is the starting band, meaning partial recognition. Thirty to fifty percent is a reasonable band after sustained investment. Above fifty percent generally appears only for category definers or highly authoritative sites.

The ratio of recommendation rate to mention rate is more informative. Above 0.6 means the brand is usually recommended positively when mentioned and the content persuades. Below 0.3 means it is known but not trusted, pointing to content credibility rather than visibility.

Citation rate sits lower: most sites fall between five and fifteen percent. Above twenty percent means your content has become a significant evidence source for the topic.

Why a Single Sample Cannot Support a Conclusion

AI answers are stochastic: the same prompt at different times, in different sessions, or even in adjacent calls to the same model can produce different answers. GEO metrics are therefore observations of a random variable, not fixed values.

The correct approach is repeated sampling of the same prompt set with uncertainty expressed as a confidence interval. For ratio metrics, the Wilson interval is more robust than a normal approximation at small samples: with thirty samples and six mentions, the point estimate is twenty percent, but the ninety-five percent interval spans roughly nine to thirty-eight percent.

That width matters. If you measured twenty percent last week and twenty-six percent this week on thirty samples, that is entirely consistent with random variation and cannot be called improvement. Detecting a trend requires either substantial separation of intervals or raising sample count above one hundred to narrow them.

Broccoli AI GEO samples core intents repeatedly and reports variance precisely to avoid reading noise as signal. In any report, look at intervals before point values.

Five Common Misreadings

One, treating rising mention rate as business growth. Mention rate reflects visibility, not conversion. It becomes meaningful only when recommendation and citation rates move with it.

Two, comparing two rounds that used different prompt sets. Changing the prompt set makes results incomparable; the set must be fixed.

Three, watching only your own brand. When industry-wide visibility rises, your mention rate can rise while your relative position falls. Track Share of Voice alongside.

Four, extrapolating one platform to all AI. Retrieval strategies and corpora differ substantially, so divergence between ChatGPT and Perplexity is normal rather than anomalous.

Five, ignoring the time distribution of sampling. Collecting all samples within one hour introduces time correlation, so samples are not independent. Spread sampling across several days and times of day.

Building a Reproducible Measurement Process

Step one, fix the prompt set: ten to fifteen prompts tightly tied to the business, spanning category, scenario, and constraint types. Keep it as a versioned document and log every change with time and reason.

Step two, fix the sampling configuration: sample count, platform list, and time window all stay constant. Run weekly or biweekly and spread samples across different days and times.

Step three, fix judgement rules: define explicitly what counts as a mention, a recommendation, and a citation. Write it down so human judgement does not drift between rounds.

Step four, establish a control group: track three to five key competitors over the same window to separate industry movement from your own change.

Step five, review quarterly: watch the direction in which the intervals move across all four metrics, not a single round's value. Genuine improvement shows as upward movement of intervals across at least three consecutive rounds.