Google AI Overviews

How to track Google AI Overviews

What to measure, why one check tells you nothing, and how to tell being cited apart from being named. The four states, and what each one means.

Jamie Partridge · Published · Updated · 7 min read

  • Google AI Overviews
  • Measurement
  • Citations
An AI Overview panel sitting above ordinary search results, with citation lines running from the results up into the generated answer and a rising bar chart alongside.

Google AI Overviews answer the question above the results. A reader can finish their search without scrolling to a single link, which means your position in the blue links is now a separate question from whether you were part of the answer at all.

That makes AI Overviews worth tracking. It also makes them awkward to track, because almost every instinct carried over from rank tracking is wrong here. Google's own help page is blunt about what the thing is: an AI-generated snapshot with links to dig deeper, and one where "AI responses may include mistakes", so you should check anything important in more than one place.

Why an AI Overview is not a ranking

A traditional search result is retrieved from an index. Ask the same question twice, ten minutes apart, and you get essentially the same ten links in essentially the same order. That stability is what makes a rank tracker meaningful: position four means something, and position four tomorrow means the same thing.

An AI Overview is generated. Google assembles it fresh, from sources it selects at that moment, and the result varies. Two checks minutes apart can name different brands, cite different pages, or produce no Overview at all.

Google documents the mechanism itself. AI Overviews are grounded by retrieval-augmented generation and a "query fan-out" technique: the model issues several concurrent related searches across subtopics and assembles an answer from what comes back, rather than returning a list it already held.

This has an uncomfortable consequence. A single check of an AI Overview is close to meaningless. If a tool checks once and shows you a green tick, it is reporting the outcome of a coin flip as though it were a fact. The only honest way to describe something that varies is a rate across repeated observations, with the number of observations attached.

We set out the general version of this rule in our measurement methodology: counts always carry their denominators, and below five observations we show a fraction rather than a percentage, because one run is a draw and not a rate.

The four states worth measuring

Most tools collapse AI Overview tracking into one number: are you in it or not. That throws away the distinction that tells you what to do next. There are four outcomes, and they need different work.

No Overview appeared

Google answered with ordinary results and no Overview at all. This is common, and it varies by query, by location and over time. It is also deliberate: Google says AI Overviews are shown only when its systems judge them additive to classic Search, and "as such, often don't trigger".

This is not a query you are losing. There is nothing to win on it yet. If you count it as an absence you will depress your own numbers with queries that were never available, and you will chase a problem that does not exist. We record the absence separately and exclude it from the denominator.

An Overview appeared and you were absent

An Overview ran, named other brands, and did not mention you. This is the real gap, and it is the number most worth watching.

You were cited but not named

Your page is in the source panel. A competitor is named in the sentence a reader actually reads.

This one surprises people, and it is the most commonly misdiagnosed state in the whole category. The instinct is to write more content. That is usually the wrong response: Google already found your page, already judged it relevant enough to retrieve, and already linked it. What it did not do is repeat your brand as the answer. That is a framing problem rather than a coverage problem, and another three articles on the same topic will not fix it.

You were named

Your brand appears in the answer text. Worth recording where in the answer, because a first mention and a mention in the final clause are not equivalent.

What to actually measure

Four things, per query, per day.

Whether an Overview appeared at all. The base rate. If Overviews appear on three of your twenty tracked queries, your addressable surface is three queries, not twenty, and every percentage should be against that denominator.

Whether your brand was named. The headline. As a share of the checks where an Overview appeared, never as a share of all checks.

Whether your domain was cited. Tracked separately from naming, because the gap between the two is the single most actionable signal available, and the reason our measurement methodology keeps the two as separate states rather than one score.

Which competitors were named. A share of voice with nobody to share it with is just a number. Being named in two Overviews out of twenty means one thing if nobody else is named either, and something entirely different if one competitor is named in eighteen.

Location changes the answer

AI Overviews vary by where the search comes from. An Overview served in London can name different brands to the same query in New York, and the citation set can differ even when the named brands do not.

So location is part of the measurement rather than a setting you configure once and forget. A tool that does not tell you which market a check was run against is not telling you what the number means. We measure the United States market and we say so on the page, which is the same discipline that puts a denominator on every figure in how a run works. What we will not do is blend two markets into one average, because the average describes nowhere.

AI Overviews and AI Mode are different surfaces

Both are Google. They are not the same thing and should never be merged.

AI Overviews sit above ordinary results on a normal search. AI Mode is a separate conversational surface with its own retrieval behaviour. We have measured the two naming different brands for the same question on the same day, and Google says to expect exactly that: the two "may use different models and techniques, so the set of responses and links they show will vary".

Averaging them into a single Google score would smooth away exactly the disagreement worth acting on. If AI Mode names you and the Overview does not, that gap is a specific, investigable thing. A blended number makes it vanish. The same argument applies across every engine, which is why we report each platform separately rather than producing one score.

How often to check

Several checks a week is the right default for a set you care about.

Weekly is too slow to separate a real movement from ordinary variance, because the variance between individual runs is large enough that two weekly checks can differ substantially with nothing having changed. Hourly is waste: you pay for volume that tells you nothing a steady series does not.

Give it about two weeks before you read a direction into anything. Below that you are looking at noise with a trend line drawn through it. This is also why we show a fraction rather than a percentage below five checks on the AI Overview tracker: one run is a draw, not a rate.

None of that variance means the Overview is wrong, and the two get confused constantly. An answer can be factually solid and still name a different set of brands on the next run, which is a separate question from whether Google hallucinates — worth reading on its own if you have been asked whether AI Overviews are accurate.

Common mistakes

Checking once and believing it. The single most common error, and the one most tools encourage.

Counting no-Overview queries as losses. Depresses your numbers with queries that were never winnable and sends you chasing a phantom.

Merging citations and mentions into one score. They have different causes and different fixes. Collapsing them destroys the information, which is the same reason we never average the six engines we track into one number either.

Ignoring competitors. Your own number in isolation cannot tell you whether two out of twenty is bad.

Rewriting the query. Wording changes the answer. Change a tracked question and you have started a new series, not continued the old one. If you must change it, treat it as new.

Where to start

Pick ten to thirty questions your buyers actually ask, phrased the way a person would type them. Fix the wording. Settle on one market and stay in it. Run on a fixed schedule, record all four states, and track your competitors alongside yourself.

None of that replaces ordinary search work. Google's own position is that the best practices for SEO continue to be relevant because its generative AI features "are rooted in our core Search ranking and quality systems". Measurement tells you which of the four states you are in. It does not substitute for the content that gets you out of it.

Then wait two weeks before drawing conclusions. The first day tells you where you stand. The first fortnight tells you which way it is going, and only the second of those is worth acting on.

If you want this run for you, that is what our AI Overview tracker does, and how it works covers the mechanics end to end.

Questions about Google AI Overviews

Why does the same query give a different AI Overview each time?
An AI Overview is generated at the moment you search, not retrieved from an index, so Google assembles it fresh from whichever sources it selects that time. Two checks minutes apart can name different brands, cite different pages, or produce no Overview at all. That variance is the reason a single check tells you almost nothing and why a rate across repeated checks is the only honest way to describe it.
Do AI Overviews appear for every search?
No, and the base rate is the first thing worth measuring. Google shows an Overview on some queries and not others, and that changes by query, by location and over time. A query with no Overview is not a query you are losing, so it belongs outside your denominator rather than counted as an absence.
What does it mean if I am cited but not named?
Your page is in the source panel while a competitor is named in the sentence a reader actually reads. Google found your page, judged it relevant enough to retrieve, and linked it, but did not repeat your brand as the answer. That is a framing problem on the page rather than a coverage problem, which is why writing three more articles on the topic usually does not fix it.
How often should I check an AI Overview?
Several times a week on each tracked query, and daily if you can afford it. Answers vary enough between runs that a weekly check cannot separate a real change from ordinary noise, and a trend needs roughly two weeks before a direction becomes visible.
Does location change the AI Overview?
Yes, enough that a check run from the United States and one run from the United Kingdom are different measurements and should never be merged into an average. Pick the market your buyers are in and stay in it, and make sure your tool says which market a number describes rather than leaving you to assume.
Jamie Partridge, Founder of AI Mention Tracker

Jamie Partridge

Founder

Builds AI Mention Tracker. Spends most of his time reading AI answers that name someone else, and working out why.

More about why we built this

More on google ai overviews

See where you stand

Connect a brand, pick the questions your buyers ask, and get your first snapshot in a few minutes. Card required, nothing charged for 7 days.