TDM Insights brand mark TDM Insights SEO & AI Advisory Book a diagnosis
AEO

Every AI Overview Statistic Contradicts the Next

AI Overview statistics disagree because each tracker picks a different keyword set and window. See how to read any prevalence figure.

By David Jubé · Jul 20, 2026 · 13 min read
TDM Insights graphic with the headline 'Same period. Opposite headlines.' and a green background.

AI Overview statistics disagree with each other because of decisions made before any measuring starts, not because Google keeps changing how often the feature appears.

59.73 percent and 15.69 percent describe the same feature in the same market, and both were honestly measured. SE Ranking reached the first by tracking 100,000 keywords across 20 curated US niches, reporting from May 2025 through February 2026. Semrush reached the second across more than 10 million keywords, reporting a climb from 6.49 percent in January 2025 to a July 2025 peak of 24.61 percent, then a fall to 15.69 percent by November 2025.

Those are two keyword universes measured over two differently defined stretches of months, so no arithmetic turns one figure into the other.

The gap has a cause, and Google changing its mind between studies is not it. Sampling frame, start and end dates, and detection method are three choices each vendor makes before a single search runs, and each one moves the headline. Which makes the useful question about any prevalence figure not whether it is accurate, but what it was accurate about.

Key takeaways

  • SE Ranking reports prevalence rising to 59.73 percent in February 2026 across 100,000 keywords in 20 curated US niches, measured from May 2025. Semrush reports 15.69 percent in November 2025 across more than 10 million broadly pulled keywords, measured from January 2025.
  • Those two windows are ten months and eleven months long, they start four months apart and end three months apart, and they overlap for only seven of them, so the two trend lines cannot be laid over each other as a contradiction about one shared period.
  • BrightEdge connects two endpoints twelve months apart (February 2025 against February 2026). Conductor reports a monthly series across 274,524,214 US searches from September 2025 to February 2026. They share one month, February 2026, and disagree about it by roughly 13 points.
  • Query shape moves the number further than any vendor gap does. One Ahrefs sample of 146 million results ran from 0.09 percent on navigational queries to 57.9 percent on question-based ones.
  • Read across a row, never down a column: a prevalence figure is usable only alongside the keyword set, the sample size and the exact window that produced it.

No Prevalence Figure Means Anything Without Its Window

Neither growing nor shrinking, taken as one settled trend. The major trackers publishing AI Overview prevalence right now disagree with each other across overlapping months, and the disagreement traces to how each one measured rather than to a swing in how often Google serves the feature.

Across this series I run each widely repeated AEO claim through the same four steps, claim, test, verdict and limit, set out in What the AEO Evidence Actually Shows Right Now. This claim breaks at the second step, because the tests behind it were never measuring the same thing. iPullRank’s running collection of AI Overview findings shows how far apart they sit.

Activation rates published across 2024 alone ranged from 10 percent to 42 percent, depending on whose sample you took.

Three decisions do most of the damage, and each is made before a single search runs:

  • Which keywords the tracker samples, and how that list was built.
  • Which two dates the headline connects, and how far apart they sit.
  • How the tracker decides a rendered result page counts as carrying an AI Overview.

How AI answer engines choose their sources covers the retrieval side of the feature. Measurement is a separate problem, and it is the one producing the contradictory headlines.

Keyword Set Decides the Headline: SE Ranking Against Semrush

What Each Tracker Sampled

SE Ranking’s tracking runs on 100,000 keywords spread across 20 curated US niches, 5,000 per niche, a set built to represent commercial and informational search behaviour. Prevalence there moved from about 28 percent in May 2025 to 60.85 percent in January 2026, easing to 59.73 percent in February 2026.

Semrush’s AI Overviews study works from a far wider net, more than 10 million keywords pulled broadly rather than curated by niche. Prevalence there rose from 6.49 percent in January 2025 to a peak of 24.61 percent in July 2025, then fell to 15.69 percent by November 2025.

Curation is the whole difference. A set built around the questions an AI Overview exists to answer will report a higher hit rate than an unweighted pool carrying every navigational, branded and long-tail query that was never going to trigger one.

Both figures are correct answers to different questions, which is why neither one settles the argument.

Why the Two Windows Cannot Be Set Against Each Other

SE Ranking reports across ten months, May 2025 to February 2026. Semrush reports across eleven months, January 2025 to November 2025. The two windows start four months apart, end three months apart, and share only seven months in the middle.

Ten months and eleven months offset by a third of a year are not the same window.

  • SE Ranking: May 2025 to February 2026, ten months of reporting.
  • Semrush: January 2025 to November 2025, eleven months of reporting.
  • Shared: May to November 2025, the seven months both vendors cover.

So the two trend lines cannot be laid over each other and read as a contradiction about one stretch of time. SE Ranking’s steepest climb lands in the months after Semrush stops reporting, and Semrush’s decline sits inside a period SE Ranking reports only as part of a longer rise.

Window Length Decides Direction, and the Shared Month Still Disagrees

BrightEdge and Conductor both end their reporting in February 2026 and reach it by very different routes.

BrightEdge’s year-over-year read connects two endpoints twelve months apart, February 2025 and February 2026, and reports growth from about 31 percent to about 48 percent, which it frames as a 58 percent increase.

Conductor’s month-by-month analysis covers roughly six months, September 2025 to February 2026, across 274,524,214 US searches. It reports a January 2026 peak of 47 percent falling to 34.5 percent in February 2026.

Twelve months against six is another mismatch, and it decides the direction on its own. Two endpoints a year apart produce a growth story. A short monthly series produces a growth story with a correction on the end of it.

Three readings sit inside the same feature:

  • Twelve months, two endpoints: growth from about 31 percent to about 48 percent.
  • Six months, monthly series: a 47 percent peak in January 2026, 34.5 percent in February.
  • One shared month, February 2026: about 48 percent against 34.5 percent.

At February 2026, BrightEdge reports about 48 percent and Conductor reports 34.5 percent, a gap of roughly 13 points in the same month on the same market. Window length explains the opposite headlines. It does not explain that 13 points, which leaves sample composition and detection method to account for the rest.

Vendor Prevalence Figures Are Not About Your Market Every tracker above measured a keyword universe assembled for a study, not the queries that bring you customers. A diagnosis starts by fixing a query set to your business and running it on one window, so the trend you read is yours rather than borrowed. Book a free diagnosis

Seven Datasets Side by Side, Each With Its Window Stated

Every figure in the table below carries a link to the source that published it, because a prevalence number without a source and a window is not evidence. Comparison tables and formats LLMs love to cite travel well precisely because they are liftable, which is an argument for keeping the methodology attached to the number.

TrackerSample as reportedWindow as the tracker defines itHeadline figureSource
SE Ranking100,000 keywords, 20 curated US nichesMay 2025 to February 2026About 28% rising to 59.73%SE Ranking
SemrushMore than 10 million keywords, broad poolJanuary 2025 to November 20256.49% to a 24.61% peak, 15.69% at the endSemrush
BrightEdgeNot disclosedFebruary 2025 against February 2026, two endpointsAbout 31% to about 48%BrightEdge
Conductor274,524,214 US searchesSeptember 2025 to February 2026, month by month47% in January 2026, 34.5% in February 2026Conductor
Ahrefs146 million search resultsReported 202520.5% overall, 0.09% navigational to 57.9% question-basedSearch Engine Journal
Pew Research Center900 US adults, 68,879 real searchesMarch 1 to March 31, 202518% of searches returned an AI summaryPew Research Center
Peec AI500,000 prompts, commercial intent onlyApril 202687%Search Engine Journal

Reading the Table Without Stacking the Rows

Rows here are not competitors for one right answer.

Each describes a different keyword universe over a differently defined window, so the only sound use of the table is lateral. Read across a row, not down a column.

Three things belong to every row and travel with it:

  • Sample: how the keyword list or panel was assembled, and how large it was.
  • Window: written out with a start month and an end month, not a season.
  • Detection: how a rendered result page was judged to carry an AI Overview.

Down a column is where the contradiction gets manufactured. Put 87 percent and 18 percent in one sentence with no methods attached and it looks like a scandal. Attach “500,000 commercial-intent prompts, April 2026” to the first and “68,879 real searches by 900 US adults, March 2025” to the second and it turns into two facts about two different populations.

The Range These Seven Rows Support

Published prevalence figures span 18 percent on a real browsing panel to 87 percent on a commercial-prompt sample, with the large keyword trackers landing between 15.69 and 60.85 percent.

That spread is the finding. Averaging it would produce a number describing no real keyword set at all, which is why Similarweb’s zero-click reporting and every other market-level read is worth checking against the query mix behind it. Semrush’s own tracking makes the same point from a different angle: the share of AI Overview queries that are informational fell from 89.03 percent in October 2024 to 57.16 percent a year later, so even a fixed keyword list changes what it is measuring as Google shifts which query types trigger the feature.

Inside One Dataset, Activation Runs From 0.09 to 57.9 Percent

The clearest measurement of how far query shape moves the answer comes from inside a single dataset rather than from the gap between two.

Ahrefs analysed 146 million search results and put overall AI Overview presence at 20.5 percent. Broken out by query characteristic, the same sample ran from 0.09 percent on navigational queries to 57.9 percent on question-based ones, with medical and other YMYL topics at 44.1 percent.

Inside that one sample:

  • Navigational queries: 0.09 percent.
  • Whole sample, all query types: 20.5 percent.
  • Medical and other YMYL topics: 44.1 percent.
  • Question-based queries: 57.9 percent.

One dataset, one detection method, one window, and a spread of nearly 58 points that depends only on the shape of the query.

Independent academic work reports the same effect at a similar magnitude. Xu, Iqbal and Montgomery of Washington University in St. Louis (arXiv:2605.14021, published May 2026) sampled 55,393 trending queries between 13 March and 21 April 2026 and found 13.7 percent overall activation against 64.7 percent on question-form queries.

Nothing about Google changed between those pairs of figures. Only the queries did, which is the mechanism the entire vendor disagreement is built from, isolated and measured in one place.

Four Checks to Run Before Repeating AI Overview Statistics

Methodology literacy here comes down to four checks, and a figure that cannot answer all four is a headline rather than evidence.

  • Keyword set: curated niche sets and broad unweighted pools measure different populations, so their hit rates are not comparable in either direction.
  • Sample size: 900 panellists, 100,000 keywords and 274 million searches carry very different error bars, and Nielsen Norman Group’s nine-participant study is built for behaviour rather than prevalence.
  • Exact window: a ten-month span and a six-month span ending in the same month are not the same measurement, even when both describe the same metric on the same market.
  • Endpoints or series: two dates a year apart hide everything between them, and the choice of the second date decides whether the story reads as growth or as correction.

There is a fifth variable, and almost nobody publishes it. Search Console does not attribute AI Overview impressions separately, so every tracker builds its own detection pass to decide whether a rendered result page counts, and none of the vendors above releases its raw query list or its classification logic. That is the limit on everything in this article: the methodology gap is inferred from what each vendor discloses, not audited directly.

What to Do With a Figure That Fails One

Discard the trend, keep the observation.

A figure with no stated window still tells you that some share of some keyword set triggered the feature, which is worth knowing once and worth nothing as a time series. Measuring AI traffic when there is no referrer hits the same wall from the analytics side. The number exists, the attribution behind it does not, and treating the two as equivalent is where reporting goes wrong.

What Good Looks Like: One Query Set You Hold Constant

Three things get fixed once and then left alone.

  • Write down 50 to 200 queries a real prospect would type, and stop editing the list between pulls.
  • Run it on a 28-day rolling window, which means the same thing on every pull where a calendar month does not.
  • Compare like window against like window, the rule BrightEdge and Conductor would have needed in common.

Record how many of those queries return an AI Overview each time, and read the series rather than any single pull.

That number moves for reasons that belong to you: your queries, your market, your competitors. It does not move because a vendor widened its sample between two studies.

How to show up in Google AI Overviews covers what makes a page eligible in the first place, and SEO and AI visibility tracking tools, compared sets out what each platform measures before you commit budget to one. Once you can judge a statistic on your own terms, the next question is which signals correlate with AI visibility. Brand Mentions Beat Backlinks in Ahrefs’ Own Data holds a well-specified statistic up to the same light.

Frequently Asked Questions

What percentage of Google searches show an AI Overview?

Published estimates run from 18 percent on a real browsing panel to 87 percent on a commercial-prompt sample, with the large keyword trackers landing between 15.69 and 60.85 percent. The right figure depends on which queries were sampled. No tracker measures the whole market evenly.

Why do AI Overview studies disagree with each other?

Each study samples a different keyword universe, defines its own start and end dates, and runs its own detection pass. SE Ranking’s curated 100,000-keyword set and Semrush’s 10 million-plus broad pool were never measuring the same population, so their opposite trends are compatible rather than contradictory.

Which methodology details change a prevalence figure most?

Keyword set moves it furthest, because AI Overviews are distributed unevenly across query types. Window definition comes second, since a twelve-month endpoint comparison and a six-month monthly series can read as growth or as correction. Detection method comes third and is rarely disclosed.

Does query mix explain the gap between AI Overview datasets?

Largely, yes. One Ahrefs sample of 146 million results found 20.5 percent activation overall, 0.09 percent on navigational queries and 57.9 percent on question-based ones. Query shape alone opens a range inside a single dataset wider than the gap between most competing vendors.

How should I read an AI Overview statistic before repeating it?

Check four things: the keyword set and how it was chosen, the sample size, the exact start and end dates, and whether the figure compares two endpoints or a series. Read across those attributes rather than down a list of competing percentages.

Can I rely on a single vendor’s AI Overview tracker?

Not for a market-wide claim, because no vendor publishes its raw query list or classification logic. A tracker is reliable for its own keyword set over its own window. For your business, a fixed query set you control on a constant window beats any vendor aggregate.

Continue Reading

More From This Series

More from TDM Insights

Explore TDM Insights Topics