AEO Measurement 12 min read

How do you know if AEO is actually working?

Seeing your brand mentioned by ChatGPT is a useful signal. It is not a measurement strategy. If AEO is going to earn budget, it needs to be measured from the first recommendation through to the commercial outcome. That means tracking more than rankings or screenshots. It means understanding whether your brand appears, how prominently it appears, what sources support the answer, whether people visit your site afterwards, and whether any of those people become customers. That is how I measure AEO performance.

The Short Answer

How should AEO performance be measured?

AEO performance should be measured across five layers: AI visibility, recommendation prominence, citation strength, website behaviour and commercial outcomes.

At minimum, a useful measurement system should answer:

  1. Are AI engines mentioning the brand?
  2. Are they recommending it, or merely mentioning it?
  3. Which sources are influencing those answers?
  4. Is AI visibility creating measurable website activity?
  5. Is that activity becoming leads, pipeline or revenue?

A single "AI visibility score" can be useful as a summary. It should never replace the underlying numbers.

The Problem

A screenshot is evidence. It is not a KPI.

A typical AEO success report looks something like this: "ChatGPT now mentions your company for 14 prompts." Good. What happened next? Did the company move from sixth in the answer to first? Was it actually recommended? Did the answer describe the product correctly? Which sources caused the change? Did anybody visit the website? Did those visitors convert? Did pipeline increase?

Without those answers, we know visibility changed. We do not yet know whether the business benefited. That distinction matters. AEO lives somewhere between brand, search, reputation, content and acquisition. Trying to squeeze all of that into one number usually hides more than it reveals. I prefer to measure the chain.

A screenshot is evidence. It is not a KPI.

The AEO Measurement Chain

01AI visibility
02Recommendation
03Citation
04Visit
05Lead
06Pipeline
07Revenue
What to Measure

The five layers of AEO performance

1. Visibility

The first question is the simplest: Does the brand appear at all? Build a fixed set of commercially relevant prompts and test them consistently across the AI and search environments that matter to your market. Do not build the prompt set around your company name. Build it around buyer behaviour. Instead of "What is Quantumflux Digital?" test "Who can help a South African business improve its visibility in ChatGPT?" or "Which AEO consultants should I consider in South Africa?".

The measurement then becomes Brand Mention Rate = responses mentioning the brand ÷ eligible responses. If the brand appears in 32 out of 100 relevant responses, its observed mention rate is 32%. Simple. Repeatable. Much more useful than "we asked ChatGPT a few questions and found you."

2. Recommendation prominence

Not every mention has equal value. There is a substantial difference between "Other providers include Company X." and "I would recommend Company X for this requirement." So I separate presence from prominence. Useful measures: first recommendation rate; average recommendation position; percentage of prompts where the brand appears in the top three; percentage of responses where the brand is explicitly recommended; competitor recommendation frequency.

One useful metric is First Recommendation Rate = responses where the brand is recommended first ÷ eligible responses. This gets much closer to the commercial question: when the buyer asks for help, are we actually in the consideration set?

3. Citation strength

Next, look behind the answer. Why does the AI believe what it believes? When citations or source references are available, record them. You are looking for patterns: Is the brand's own website being used? Are third-party publications being used? Are old pages influencing current answers? Do comparison sites appear repeatedly? Are competitors supported by stronger external sources? Are review sites affecting recommendations? Is the answer drawing from inaccurate or outdated material?

Citation analysis often exposes the real AEO problem. A company may have excellent website content but a weak information footprint outside its own domain. Another may have very little content, but be repeatedly supported by respected third-party sources. That matters. Because AEO is not only about what you publish. It is also about what the information environment can verify.

4. Traffic and behaviour

Eventually, some AI interactions become website visits. That traffic should be separated from the rest of your referral traffic wherever the available source information allows it. At a minimum, monitor: sessions from identifiable AI referrers; engaged sessions; landing pages; key events; lead starts; lead completions; assisted conversions; conversion rate.

The mistake here is assuming this tells you the full story. It does not. A buyer may discover your brand through an AI answer and visit later through branded search, direct traffic or another device. Some referral information may also be unavailable. So AI referral traffic should be treated as observable AI traffic, not the total economic influence of AI search. That distinction is important. I go deeper on this in AI Search Attribution.

5. Leads, pipeline and revenue

This is the layer I care about most. Once identifiable AI visitors reach the site, the measurement job becomes familiar. Track: AI visit → key action → lead → qualified lead → opportunity → customer → revenue. If the business uses a CRM, acquisition-source information should travel into the lead record wherever the stack makes that practical.

That allows us to ask better questions. Not "How many ChatGPT sessions did we receive?" but "How many qualified opportunities started with identifiable AI referral traffic?" And eventually: "What revenue did those customers generate?" This will never capture every buyer influenced by an AI recommendation. Neither does most marketing attribution. The objective is not perfect omniscience. It is better evidence. This is the work I do in the Revenue Measurement Architecture.

The Scorecard

What I would actually put on an AEO dashboard

Ten metrics, mapped to the layer each one answers. No headline numbers here. The numbers only mean something once you have a baseline and a fixed method behind them.

LayerMetricWhat it tells you
AI visibilityBrand Mention RateResponses mentioning the brand divided by eligible responses.
RecommendationFirst Recommendation RateResponses where the brand is recommended first divided by eligible responses.
Competitive positionAI Share of VoiceHow often the brand appears or is recommended relative to competitors.
Citation footprintCitation ShareWhich sources influence the answers, and how much of that support is yours.
AccuracyAnswer Accuracy RateWhether the AI describes the brand, products and pricing correctly.
StabilityRepeatabilityWhether the same prompt returns consistent results across runs.
TrafficIdentifiable AI Referral SessionsObservable sessions arriving from identifiable AI referrers.
EngagementAI Visitor Conversion RateWhether identifiable AI visitors take the key actions that matter.
PipelineAI-Sourced OpportunitiesQualified opportunities that started with identifiable AI referral traffic.
RevenueAI-Sourced or AI-Influenced RevenueThe commercial value connected to that acquisition.
Share of Voice

A note on "AI share of voice"

Share of voice can be useful. It can also become meaningless very quickly. If I test five hand-picked prompts where I already know my client performs well, I can manufacture a fantastic score. The prompt universe matters as much as the result. A credible share-of-voice measure requires: a defined market; a defined buyer; a fixed prompt set; consistent competitors; repeatable testing; consistent platforms; a recorded date and methodology. Then the number has context. Without that, it is marketing decoration.

AI Share of Voice = brand recommendation appearances ÷ recommendation appearances for all tracked brands.

More advanced versions can weight first position, recommendation strength or platform coverage. If weighting is used, publish the weighting. Do not hide the methodology behind a proprietary score.

Don't Ignore the Answer

Ranking first with the wrong information is not a win.

AEO performance is often discussed entirely in terms of visibility. That misses an obvious problem. What is the AI actually saying? A brand could be mentioned prominently while the answer contains: discontinued products; old pricing; incorrect locations; wrong service descriptions; obsolete positioning; inaccurate comparisons. That is not successful visibility. It is scaled misinformation.

For one national-brand engagement, the important outcome was not "more mentions". The major AI engines were already talking about the business. The problem was that they were using product and pricing information that was years out of date. The work was to identify where the stale information came from, correct the entity and source layer, and test whether the answers changed. They did.

Before You Optimise

If you did not measure the starting point, you cannot prove the improvement.

Before implementing AEO work, capture the baseline. At minimum:

Define the prompt universe

Use questions reflecting real commercial discovery and evaluation. Include a spread of intents: recommendation, comparison, problem solving, trust, price and value, and product or service suitability.

Define the platforms

Choose environments relevant to the market. For many brands this could include ChatGPT, Gemini, Perplexity and Google AI experiences.

Define competitors

Lock the relevant competitive set before optimisation begins.

Record the answers

Capture brand presence, recommendation position, citations, factual accuracy, competitors, platform and date.

Record acquisition performance

Before the work starts, understand existing referral traffic, conversion rate, lead volume, pipeline and revenue.

Now there is something to compare against. Without a baseline, every positive answer after launch can conveniently be called a success. That is not measurement.

Measuring Change

The number matters less than the direction.

AEO measurement becomes much more useful when it is longitudinal.

Illustrative
Suppose a brand has 34% recommendation visibility in January and 49% in April. That may be meaningful. But I would still want to know: Did competitors also rise? Did the prompt set change? Did the AI platform change? Did the increase occur across multiple engines? Did citations change? Did referral traffic change? Did conversion behaviour change?

AI systems change constantly. So raw growth alone can be misleading. The goal is not to claim credit for every positive movement. It is to build enough measurement around the change to make a sensible commercial judgement.

Vanity Metrics

Four numbers I would never report on their own

"Number of AI mentions"

Without a fixed prompt universe, the number means very little.

"AI visibility score: 84/100"

Who decided what 84 means? Show the inputs.

"ChatGPT referral traffic increased 200%"

From 10 sessions to 30? Compared with what? Did ChatGPT itself simply become more widely used?

"We now rank #1 in ChatGPT"

For which prompt? On which run? For which user context?

AI recommendation systems are not ten blue links. Treating them that way creates false precision.

Cadence

How often should you measure AEO?

Enough to identify change. Not enough to create noise.

There is no useful universal frequency. It depends on the business and the amount of work being implemented. For an active AEO programme, I generally want:

  • Baseline before implementation.
  • Post-change test after meaningful technical, entity, source or content changes have had time to be discovered and processed.
  • Ongoing pulse, a consistent recurring subset of important prompts.
  • Periodic full benchmark, re-running the broader prompt universe to detect structural movement.

Do not change the test every month because somebody has found a nicer-looking result. Consistency is what makes the data useful.

Measurement Stack

GA4 can show part of the story. Not all of it.

When an AI service sends a user to your website with identifiable referral information, analytics can capture that source. That allows AI traffic to be analysed separately rather than disappearing inside a generic referral bucket. A practical implementation can include: source and medium analysis; a dedicated AI-assistant channel grouping; AI-specific exploration reports; landing-page analysis; key-event tracking; CRM source persistence.

But analytics can only report the information it receives. If referral information is not carried through, the session may be classified elsewhere, including direct. That is why I do not equate AI referral traffic with all demand influenced by AI. One is measurable acquisition data. The other is a broader customer-journey question. Both matter. I unpack the connection in AI Search Attribution.

The Point

The goal is not a better AEO report.

The goal is to make better decisions. If visibility is improving but traffic is not, that tells us something. If traffic is growing but leads are poor, that tells us something else. If the brand appears frequently but is never recommended first, we have a positioning problem. If competitors dominate citations, we have an authority problem. If AI traffic converts at twice the site's average rate, we may have found an acquisition channel worth investing in. If none of it produces meaningful commercial behaviour, the answer may be to spend the next rand somewhere else. That is what measurement is for.

The goal is not to prove the work was successful. Finding out whether it was.

QFD Approach

Visibility first. Revenue eventually.

My AEO/GEO work starts with a measured baseline rather than a list of tactics. I want to know: where the brand appears; how it is described; who appears instead; which sources influence the answers; what information is wrong; what can already be observed in analytics; whether the measurement architecture is capable of connecting acquisition to revenue.

Then we change the system. Then we measure it again. The output is not simply "Your AI visibility improved." The useful answer is: "Here is what changed, here is the evidence, here is what happened commercially, and here is what I would do next."

Start With the Baseline

Before you optimise AI visibility, measure what is already happening.

QFD's AEO/GEO Visibility Audit establishes the starting point across the AI environments relevant to your market. You get the visibility gaps, competitor benchmark, citation sources, answer accuracy and prioritised actions. If measurement underneath it is broken, we find that too.

30 minutes. No pitch deck. No obligation.

FAQ

Frequently asked questions

Measure brand visibility across a fixed commercial prompt set, recommendation prominence, citations, answer accuracy, identifiable AI referral traffic, conversions, pipeline and revenue. The strongest measurement systems track the journey rather than relying on a single AI visibility score.

Useful metrics include brand mention rate, first recommendation rate, AI share of voice, citation share, answer accuracy, repeatability, identifiable AI referral sessions, conversion rate, qualified opportunities and revenue.

GA4 can identify traffic when referral-source information is available. That traffic can be analysed using source and medium dimensions and grouped into a dedicated AI-assistant channel. It should not be assumed to represent every customer influenced by ChatGPT or another AI service.

AI share of voice measures how frequently a brand appears or is recommended relative to competitors across a defined set of AI prompts. The prompt set, platforms, competitors and scoring rules need to remain consistent for the metric to be meaningful.

No. Visibility measures whether a brand appears within relevant AI answers. ROI requires a commercial outcome and a defensible connection between the investment in AEO activity and the value generated. I cover this in AI Search ROI.

Establish a baseline before optimisation, retest after meaningful changes and maintain a consistent recurring prompt set for ongoing monitoring. Full benchmarks can be run periodically to identify broader changes.