Your Google Search Console shows impressions climbing and clicks flat. Your prospects ask ChatGPT for recommendations and never see a results page. Somewhere between those two facts, the AI search performance metrics question stops being academic: the measurement stack the industry built over twenty years no longer answers the only question it was built for. Are you winning?

We build Kuroma, an AI visibility platform. Our team has spent the past year measuring how AI engines talk about brands, and we have helped brands across seven markets act on what the measurements say. This page defines the metrics we believe actually measure success in AI search, what each one can and cannot tell you, and the baselines we have measured across hundreds of brands and hundreds of thousands of real AI answers. Where a number comes from our own data, we say so and date it. Where the answer is "nobody knows yet", we say that too.

Why did rankings stop being the scoreboard?

The old contract was simple. Rank well, get clicks, count them. AI search breaks the contract twice over. First, the answer often ends the journey: the user asks, reads, and leaves without visiting anyone. Second, the connection between ranking and being used as a source is weakening fast. Industry-scale crawls of Google AI Overviews found that the share of cited pages which also appear in the top organic results has been falling steeply over recent quarters. Ranking is becoming one input among many, not the scoreboard.

So what replaces it? Six metrics, in two families: metrics about being KNOWN, and metrics about being USED.

The metrics at a glance (measured baselines, August 2026)

Metric What it measures Our measured baseline
Mention rate Share of relevant AI answers naming the brand 14.4% average across 240+ brands on unbranded discovery questions
AI share of voice Your share of all brand naming in category answers Varies by category; measure per engine
Citation rate Answers naming you that also cite your site 37.2% (so 62.8% of mentions are ghost mentions)
Answer rate Substantive answers vs deflections Engines deflect under 2% of discovery questions
Criteria-first share Substantive answers naming no product at all 63 to 67% of answered discovery questions
Retrieval alignment Semantic closeness of your content to engine queries Strongest single predictor improvement we have measured

Baselines come from our own measurement corpus: 240+ brands, over 37,000 real engine answers, repeated runs, seven engines, collected through August 2026. The full method, including every cohort and test size, is published on our methodology page so the numbers are independently checkable.

What is mention rate in AI search?

Mention rate is the share of relevant AI answers that name your brand at all. Ask the engines a fixed set of questions your buyers actually ask, in your market and language, and count. It is the closest thing AI search has to unaided awareness, and it is the headline number for what we call generative engine optimization, whether you spell it GEO, AISEO, AIO, or AEO.

Two things separate a mention rate you can act on from one that flatters you. The question set must not name your brand, because an answer to "is Kuroma good?" that mentions Kuroma proves nothing. And it must be measured across repeated runs, because AI answers change from run to run far more than people expect. In our measurements, a single check is closer to a coin-flip sample than a reading. Treat any tool that shows you a mention rate from one run per prompt with suspicion, including ours if you configure it that way.

What is AI share of voice?

Share of voice is mention rate made relative: of the brands named in answers to your category questions, what fraction of naming goes to you? It answers a different question than mention rate. A 20 percent mention rate means one thing in a category where AI names nobody, and something else entirely where every answer lists five competitors. Measured per engine, it also reveals where you are strong. The engines differ a great deal in what they cite and whom they name: analysis of roughly 680 million AI answer citations found Wikipedia alone accounting for 7.8% of ChatGPT's citations while other engines lean on entirely different source mixes, and our own seven-engine corpus shows the same divergence. A single blended number hides more than it shows, and the engine where you are weakest is usually the one your dashboard averages away.

What is citation rate, and what is a ghost mention?

Being named is not being used. Citation rate measures whether the engine actually cites your site as a source when it talks about you. Our research measured this directly in August 2026: we analyzed 5,714 AI answers that named a tracked brand and checked them against the sources those same answers cited. We found the brand's own website cited in 37.2% of the answers that named it.

Key finding: in 62.8% of AI answers that mention a brand, the brand's own website is not among the sources. The engine is speaking from memory, not from your site.

Read that the other way and you get the ghost mention rate, a pattern independent research has documented at scale: in roughly two out of three answers that mention a brand, the AI is speaking from what it already knows, not from the brand's site. The engine learned about you somewhere else, and your website was not part of the conversation.

This distinction turns out to be the most useful diagnostic in the whole stack, because the two numbers respond to different work. Our analysis found something sharper: a brand's footprint on the wider web predicts whether AI knows it, but predicts whether AI cites it slightly worse than chance. This is the most significant dissociation in our data. What predicts citation is retrievability: whether your content actually surfaces when the engine goes looking. Brands famous enough to be ghost-mentioned constantly can still be invisible as sources, and small brands with sharply aligned content can punch far above their footprint in citations. If you track one pair of numbers, track this pair.

What is answer rate, and why does answer composition matter?

Before any mention can happen, the engine has to produce a substantive answer. We classify every monitored response as answered, deflected, or errored, and the composition is worth watching. In our recent monitoring data, engines rarely refuse outright, but our research found roughly two thirds of substantive answers to buyer-style discovery questions named no specific product at all: 63 to 67% of the 50,018 answered checks we classified in the first half of August 2026. The engines increasingly open with criteria, what to look for, how to compare, and only sometimes proceed to names. That is not a measurement failure. It is the market moving, and it means the page that teaches the criteria is competing for the same answer your product page used to own.

Retrieval alignment: the leading indicator

Everything above is an outcome. The most promising leading indicator we have measured is retrieval alignment: how semantically close your site's content sits to the queries AI engines actually issue when they research your category. Engines do not search your buyer's question verbatim. They fan out into their own reformulations, and whether anything on your site matches those reformulations is measurable with embeddings. When we analyzed 6,950 brand-question pairs across 193 brands, adding this one feature produced the largest single improvement we have measured in models predicting which brands get named. It is early, and we published the numbers with their limitations, but the direction is consistent with everything else on this page: alignment with how engines retrieve beats decoration of what humans read.

What does NOT measure success in AI search?

A few numbers circulate as AI search metrics that do not survive contact with data. Your organic ranking is a weakening proxy for citation, as above. AI crawler hits in your server logs mostly measure training ingestion, not answer inclusion; a bot reading your site is not a bot citing your site. And one-shot visibility checks are noise wearing a dashboard: independent studies of answer stability, and our own repeated sampling, both find that the sources cited for the same question overlap between runs far less than dashboards imply. Anything measured once is mostly measuring luck.

Content tricks deserve a mention here too. The academic field that named GEO, in the original 2024 paper, reported large visibility gains from content transformations, but the recent replication work, led by the C-SEO Bench study, tells a more sobering story: across engines and domains, most rewriting tricks moved nothing, and several made citation rank worse. What survived replication is substance, citing sources, quoting evidence, carrying real statistics. Which is convenient, because that is also just good writing.

How do you measure AI search performance without fooling yourself?

Four rules our specialists learned the hard way, current as of August 15, 2026:

  • Freeze the question set before you start measuring, and version every change.
  • Sample repeatedly: report rates over runs, never single checks.
  • Split by engine: ChatGPT, Gemini, Perplexity and Google AI Overviews behave like different countries, not different fonts.
  • Demand published validity: a visibility score that never shows whether it predicts real citations is a mood, not a metric.

In longer form: Fix your question set before you start and freeze it, or you will measure your own prompt edits instead of the market. Sample repeatedly and report rates, never single runs. Split every metric by engine, because ChatGPT, Gemini, Perplexity and Google AI Overviews behave like different countries, not different fonts. And demand that any score you rely on publishes its predictive validity: if a vendor sells you an AI readiness number, ask them to show whether it actually predicts citations. We publish ours, including the cohort where it did badly, because a number you only show when it flatters you is not evidence.

AI search is measurable. It is just not measurable with the instruments the last era left behind. Define the metrics, publish the baselines, and let the data argue.

Frequently asked questions about AI search metrics

What are the core AI search performance metrics?

Mention rate, AI share of voice, citation rate, ghost mention rate, answer rate, and retrieval alignment. The first two measure whether AI knows you; the next two measure whether AI uses you as a source.

What is a good mention rate?

Our measured average is 14.4% across 240+ brands on unbranded discovery questions, but the benchmark that means anything is per category and per engine, against your competitors on the same frozen question set.

Why does my brand get mentioned but not cited?

Because models often speak from what they already learned about you. We measured own-site citation in only 37.2% of answers naming a brand. Citation follows retrievability, not fame.

Do SEO rankings still matter for AI search?

As one input, yes; as the scoreboard, no. The overlap between cited pages and top organic results has fallen sharply, and our own models found ranking features add far less than footprint and retrieval alignment.