What the AEO Evidence Actually Shows Right Now
Four AEO claims, four published tests, four verdicts. See where the AEO evidence holds, where it fails, and where it runs out entirely.
AEO evidence exists for the four claims founders act on first, and it comes back negative on three of them.
Four AEO claims carry almost all of the work being done in this space right now, and three of them fail their own tests.
Schema markup produced no citation lift in the controlled testing run so far. Adoption of llms.txt keeps climbing while the evidence behind it stays flat. Published AI Overview prevalence runs from single digits to nearly sixty percent, depending on whose panel you read.
One claim holds. Unlinked brand mentions correlate with AI Overview visibility roughly three times more strongly than backlink count does, measured across 75,000 brands.
Below is the four-step check I put every claim through, the verdict each of the four came back with, and where I land on the argument sitting underneath all of them.
Key takeaways
- Schema markup did not move AI citations, and the reason is mechanical. One experiment planted a fake address inside invalid JSON-LD and got it back in the answer, which is ChatGPT and Perplexity reading structured data as ordinary page text.
- llms.txt adoption is real and its evidence is not. Ship the file anyway, for reasons that have nothing to do with getting cited.
- Prevalence figures are a measurement definition as much as a finding. SE Ranking ran the same 100,000 keywords twice and reported 64 percent and then 8.71 percent, because the second count only included actual AI Overview answers.
- Unlinked brand mentions are the one signal here with a large published correlation behind them, and the gap over backlink count is wide enough to move a marketing budget.
- Fix the technical floor once, then spend the ongoing hours on brand.
Three of Four Claims Fail Their Own Tests
Here is the whole answer before the working: of the four AEO claims a founder is most likely to act on, three come back negative or unreadable once you go to the study behind them, and one comes back strong enough to move a budget.
Those four verdicts, one line each:
- Schema markup drives AI citation: no.
- llms.txt adoption predicts AI citation: no.
- AI Overview prevalence is definitively rising or falling: neither, as the claim is posed.
- Brand mentions beat backlinks for AI visibility: yes, for AI Overviews specifically.
This page is the hub of a six-part set, and each of the four claims gets a full test in its own article, with a fifth taking on what happens after a citation lands.
If you need the ground floor first, what AEO actually is and how AI answer engines choose their sources cover the retrieval and citation mechanism everything below assumes you already have.
Claim, Test, Verdict, Limit: The Check Behind Every Line Above
Every claim on this page runs through the same four steps before I say anything about it, and the fourth step decides whether the other three hold up.
- Claim: the assertion in the words practitioners actually use, not a softened version of it. “Schema helps you get cited by AI.” “llms.txt is now essential.”
- Test: the specific, named, dated study behind the claim, with its sample size, its method, and what it actually measured rather than what the headline implies it measured.
- Verdict: my position given that evidence. Never “it depends” without saying what it depends on, and never a summary of both sides with no landing point.
- Limit: what the test cannot tell you, and what would have to be true before I would push the claim further than the data does.
Step four is where almost all of the damage happens. A correlation gets reported as a causal instruction, a single-vendor model gets reported as an industry finding, and a number measured on one keyword panel gets quoted as a property of Google.
A claim that survives step two and states its limit plainly is one you can act on this quarter.
Scoring the Four Claims Against Their Own Studies
Every figure in the table below appears on the page cited beside it, so each row can be checked at source.
The schema verdict rests on independent testing reported by Search Engine Roundtable, the llms.txt verdict on SE Ranking’s own account of what does and does not move AI visibility, and the prevalence range on SE Ranking’s AI Overview tracking alongside two trackers reading the same month five points apart.
The brand-mentions row comes from Ahrefs’ study of 75,000 brands, whose headline finding Ahrefs’ content lead frames as unlinked mentions having “very little impact on SEO, but a much bigger impact on GEO”.
Where the Three Negatives Agree
The three negative verdicts are not three separate disappointments. They are one finding arriving three times: the things you can add to a page are not the things an answer engine weighs when it decides whom to name.
Schema is the clearest case. The markup is still worth having, because it earns rich-result eligibility and makes a page machine-readable, and the founder’s version of that job takes an afternoon. What it does not buy is a citation.
The file at the centre of the second claim lands the same way for a different reason. It is cheap and harmless, and nothing published so far shows an engine changing its citation behaviour because it found one.
Prevalence is the odd one out, because there is no null result to report. There is a range, and the panel produces it:
- Semrush’s 200,000-keyword study found 82 percent of desktop AI Overviews appearing for keywords under 1,000 monthly searches, and 80 percent for informational intent.
- SE Ranking’s panel is 100,000 US keywords spread evenly across 20 niches, a deliberately different mix.
- Two trackers reading the same month in July 2026 still landed five points apart.
Load a panel with informational long-tail terms and the number climbs. Load it with commercial head terms and it falls.
Both vendors are measuring accurately, and neither number is a property of Google. The datasets side by side takes that apart with each vendor’s own stated window.
Why the One Positive Is Worth Acting On
The brand-mentions finding is the only one of the four with a correlation large enough, and a sample wide enough, to move a budget.
Ahrefs found unlinked web mentions correlating with AI Overview visibility at 0.664 on the Spearman scale, brand anchors at 0.527 and brand search volume at 0.392. Backlink count came in at 0.218, and the full correlation table carries the rest.
The limit matters as much as the number. This is correlation rather than causation, and the dataset covers Google AI Overviews rather than every engine, so the honest reading is that mentions predict visibility, not that mentions cause it.
Underneath the Four Claims Sits One Bigger Disagreement
Two of the more serious writers in this field read the same landscape and reach opposite conclusions about what AEO fundamentally is, and a founder with limited hours has to pick one of them to fund first.
The Engineering Case
Mike King treats AI visibility as a retrieval and information-engineering problem. Writing for iPullRank, he notes that “you still need to make content accessible, indexable, and understood, but the difference is that in classic IR, your content comes out the same way it goes in”, which it does not in a generated answer.
The institutional version of that position is Relevance Engineering, which iPullRank describes as a framework that “treats visibility in modern search and AI systems as an engineering problem because you have to build something rather than tweak something”.
The Brand Case
Eli Schwartz reads it the other way. Writing on 20 August 2026, he puts AEO as “closer to brand marketing than to SEO” from a conversion standpoint, because the payoff shows up as influence rather than as an attributable click.
A week earlier he set out the failure mode directly: “if AEO efforts are built to satisfy a checklist rather than actually finish someone’s job, the model will eventually ignore that site” and pull from whoever solved the real problem instead.
- King describes the floor: crawlability, structure and machine-readability are what an engine needs before it can retrieve you at all.
- Schwartz describes what comes after it: once retrieval is possible, the question is whether a model has any reason to name you.
They only contradict each other if you assume the two layers compete for the same hour, which is the assumption the evidence can settle. If the vocabulary is new, the difference between GEO, SEO, AEO and LLMO is worth ten minutes first.
What the AEO Evidence Says About That Fight
The four tests do not back either writer outright. They split the fight in a specific and useful place: the floor tests flat, and the one signal with a real correlation sits at the brand layer.
The Floor Tests Flat
Every technical tactic in this set came back null or negative. Schema produced no lift, and llms.txt produced no measurable influence on citation selection.
Neither result says the floor is worthless. Both say it has stopped being a differentiator.
The academic work is starting to separate the two things people conflate here. One 2026 framework distinguishes citation selection from citation absorption across 602 controlled prompts and 21,143 search-layer citations, and finds that high-influence pages are longer, better structured and richer in extractable evidence such as definitions, numbers and comparisons.
Structure earns absorption, and markup does not earn selection.
The Brand Layer Tests Strong
The correlation gap is the whole argument, and the direction of it is what makes it actionable rather than merely interesting.
- Retrieval is a prerequisite you buy once, and nothing in this set suggests buying it twice helps.
- Mention volume is ongoing work with a published correlation behind it.
- The two therefore do not compete for the same hour, because only one of them recurs.
Rand Fishkin’s case for why the website, and the brand behind it, still matters as answers get synthesised rather than clicked runs on the same logic. What an engine can repeat is a name people already use.
Your floor is probably fine. Your mention volume probably is not. A free diagnosis checks both in one pass: whether an engine can retrieve and read your pages at all, and how often your brand is actually named in the answers your buyers see. I run it against your own pages and prompts rather than a generic checklist. Book a free diagnosis
My Position: Fix the Floor Once, Then Buy Brand
King and Schwartz are both right, at different altitudes, and I am going to say which one to fund first rather than leave that as an exercise for the reader.
The technical and retrieval floor is real, cheap and worth building once. Both properties I run publish a hand-written llms.txt, and both files are public documents anyone can open:
- tdminsights.com/llms.txt is an 85-line content map, sectioned by topic.
- thegourmethost.com/llms.txt runs to 92 lines on the same pattern.
- Neither is an auto-generated stub, and neither was built on the expectation that it would earn a citation.
I shipped both knowing SE Ranking’s finding that the file has little to no influence on citation selection.
The reason is the trade rather than the evidence. The file costs an afternoon once, the downside is close to zero, and a prerequisite you cannot rule out is worth removing from the argument entirely. That is the correct shape for a floor item: cheap, done once, never chased quarterly.
Past the floor, the marginal hour buys more at the brand layer right now, and the gap is not marginal. A 0.664 correlation against 0.218 is the widest published separation between any two signals in this set, and no technical tactic tested here produced a comparable payoff in either direction.
So here is the position, stated plainly. Fix the floor as a checklist item, then spend nearly all ongoing effort on being mentioned and being asked for by name.
Where a Founder’s Next Ten Hours Should Go
If you have one afternoon a quarter for AEO and nothing more, the sequence matters more than the total, because floor work expires and brand work compounds.
- Hours one to three, the floor: confirm your key pages are crawlable and indexed, add clean schema for rich-result eligibility rather than for an AI-citation promise, and ship an llms.txt file if you do not have one.
- Hours four to five, measurement: decide how you will know anything changed. Tracking AI traffic that arrives with no referrer is a different job from rank tracking, and the tools that do it vary widely in what they sample.
- Hours six to ten, mentions: pitch journalists, answer source requests, get into comparison pieces, and give other people a reason to use your name when they describe your category.
None of the floor items moved a citation number in any test collected here, and none costs enough to argue about. The mention work is slower, harder to attribute, and the only line on that list with a published correlation attached.
For a wider map of what the category covers beyond these four claims, a landscape view of AI search strategy is a reasonable starting point.
Start with the claim founders act on first. Schema Markup and AI Citation: What a Test Found sets out the controlled test in full, with its sample, its method and the result nobody expected.
Frequently Asked Questions
What is AEO, in plain terms?
Answer engine optimisation is the work of getting your pages retrieved, quoted and named by AI answer systems such as Google AI Overviews, ChatGPT and Perplexity. It overlaps heavily with SEO. What differs is the outcome you optimise for: being named inside an answer rather than ranked beside one.
What is the Verdict Model this page runs on?
Four steps applied to every claim: state the claim as practitioners actually use it, name the specific dated test behind it with its sample size and method, state a verdict rather than a summary of both sides, and name what the test cannot tell you. Skipping that last step is how advice turns into overclaiming.
Which AEO claims currently hold up under testing?
One of the four holds. Unlinked brand mentions correlate with AI Overview visibility at 0.664 against 0.218 for backlink count across 75,000 brands. Schema markup and llms.txt both returned null results in published testing, and AI Overview prevalence figures vary with the keyword panel rather than with Google.
Is AEO an engineering problem or a brand problem?
Both, at different points in the work. The crawlability and structure floor is real, cheap and largely built out already, with no citation payoff left to claim on top of it. The evidence with the strongest published correlation to AI visibility sits at the brand layer, which is where ongoing effort belongs.
Where should a small business start with AEO?
Fix the technical floor once as a checklist item: confirm crawlability, add clean schema, ship an llms.txt file. Then move nearly all ongoing effort to earned mentions and digital PR, because that is where the published correlation to AI visibility is strongest and where the work compounds rather than expires.
How do I judge a new AEO claim on my own?
Ask four questions. What is the specific claim, what named and dated study sits behind it, what did that study actually measure rather than imply, and what does it explicitly fail to cover. A claim that cannot answer the second question is an opinion, however confidently it is written.
Continue Reading
More From This Series
- Schema Markup and AI Citation: What a Test Found
- llms.txt Has Adoption but No Evidence Behind It
- Every AI Overview Statistic Contradicts the Next
- Brand Mentions Beat Backlinks in Ahrefs’ Own Data
- Being Cited by AI Is Not the Same as Being Chosen
More from TDM Insights
- Answer Engine Optimization: What AEO Actually Is
- How AI Answer Engines Choose Their Sources
- Structured Data for SEO: The Founder’s 80/20
- GEO vs SEO vs AEO vs LLMO: What They Mean
- How to Measure AI Traffic With No Referrer
- SEO and AI Visibility Tracking Tools, Compared
Explore TDM Insights Topics