TDM Insights brand mark TDM Insights SEO & AI Advisory Book a diagnosis
AEO

Being Cited by AI Is Not the Same as Being Chosen

Being cited by AI vs being chosen are two outcomes, not one. See what the research measured and what to log so a citation count counts.

By David Jubé · Jul 2, 2026 · 13 min read
TDM Insights graphic showing a 0.57 correlation, emphasizing data analysis.

Read being cited by AI vs being chosen as two separate outcomes, because the published research says they only loosely track each other.

Visibility in AI answers is not one measurement. It is three: whether an engine reached your content, whether it named your brand, and whether it told the reader to choose you.

Only the third one is worth money, and it is the one almost nobody reports. An engine can use your page as a source and name a competitor as the pick inside the same answer.

Citations are a floor to hold. Recommendations are the outcome, and they need their own line in the report.

Key takeaways

  • Citation and recommendation are separate outcomes. Profound’s 224-task shopping study put the correlation between citation share and being chosen at 0.57.
  • Almost none of those decisions passed through a click, because 92.8 percent of the tasks ended with no meaningful click to the open web.
  • In 48.9 percent of those tasks the participant accepted the named options with no follow-up investigation at all.
  • Recommendation is settled by comparative evidence rather than citation volume. An arXiv study watched an incumbent brand’s recommendation lock break on a rating gap under 0.1 stars.
  • Log three states per prompt, per engine, plus a field for who won the ones you lost.

Being Cited by AI vs Being Chosen: Two Outcomes, One Report Line

No, a citation is not a recommendation.

An engine that cites you has confirmed it could reach, read and trust your page enough to build an answer with it. That is a precondition for being chosen, and nothing past it.

The gap that answer leaves is measurable.

Profound ran the cleanest test of that gap I have found. Fifty-six participants worked through 224 real shopping tasks inside ChatGPT, 221 of them codable, and the study matched the brands the engine cited during each session against what the participant went on to buy.

Brands that were chosen held 24 percent share of voice in those answers. Brands that were passed over held 11 percent, and the correlation between citation share and being chosen still landed at 0.57.

Double the share, half the predictive power you would want.

A citation count can climb for a quarter while the record of who got recommended sits still, and nothing inside that count will surface it.

Strip the metric back and a citation proves three narrow things:

  • Your page was reachable and indexed when the answer was assembled.
  • Its wording was extractable enough to quote or paraphrase.
  • The engine judged it relevant to one of the sub-queries behind the prompt.

Google’s own documentation on AI features describes a query fan-out, where one prompt becomes several related searches, each able to surface a different source. Your citation may have answered a sub-question the reader never asked, inside an answer that recommended somebody else.

None of the three proves a reader was told to choose you, which raises the question of why the citation count became the number everyone reports.

Why Counting Citations Feels Like Measuring Success

Citations became the default AI visibility metric because they are the cheapest thing in the stack to count. An engine either referenced a URL or it did not, and a tracker can log that across thousands of prompts without making a judgement call.

Whether the same answer told the reader to choose you is a harder extraction, so reporting settled on the countable thing instead of the decisive one.

Convenience picked the metric, not evidence.

That leaves three questions a citation count cannot answer:

  1. Whether your brand was named as the pick in the same answer.
  2. Whether a competitor was named instead, and which one.
  3. Whether the reader ever saw the citation, let alone followed it.

The habit carried over from classic search, where a citation and a visit were close to the same event. That link has broken at both ends.

SparkToro reports that in the first four months of 2026, 68.01 percent of US Google searches ended without a click, up from 60.45 percent across 2024. Profound’s shopping study finds the same behaviour inside the chat window, with 48.9 percent of tasks ending in the participant accepting the named options and investigating nothing further.

Readers are not withholding the click because the answer looks weak. Pew surveyed 5,153 US adults in August 2025 and found 20 percent rate AI summaries extremely or very useful and another 52 percent somewhat useful.

Answering those three questions means splitting one metric into three, which is the part almost every report skips.

Three Tiers Decide What an AI Answer Is Worth to You

Every AI answer files your brand into one of three states, and only the top one has a buyer attached. Averaging them into a single visibility number turns a win and a loss into a shrug.

Each figure below is sourced to its study: BrightEdge’s analysis of how differently each engine cites, Profound’s research on browsing after a brand mention, and Profound’s analysis of unsolicited editorial content.

TierWhat happenedWhat the evidence attaches to itSource
CitedYour page was used to build the answerGoogle AI Overviews sends 30% of transactional citations straight to a major retailer, ChatGPT 15%BrightEdge
MentionedYour brand was named inside the answerBrand site visits run 1.5x to 2.5x forecast baseline for 7 days, and 97%+ arrive with no UTMProfound, AI mention effect
RecommendedThe answer told the reader to pick youAbout 47% of response content is unsolicited editorial: comparisons, rankings, shopping guidanceProfound, parrot problem

Tier One: Cited, and Possibly Invisible

Being cited says the machine could reach you, and nothing about whether a person ever saw the reference. Profound’s 92.8 percent no-click figure puts a ceiling on how often anyone did. Selection is not yours to control either.

Google’s featured snippets documentation answers the question of how to make your page one with a flat “You can’t”, and the AI features guidance states there are no additional requirements or special optimisations for appearing in AI Overviews or AI Mode.

Tier Two: Mentioned, and Finally Measurable

A mention is the first tier with a demand signal behind it. Profound reviewed more than two million AI conversations between January and June 2026 and found that after an assistant names a brand, visits to that brand’s site run at 1.5 to 2.5 times the forecast baseline for seven days.

More than 97 percent of those visits carried no UTM. If your reporting counts only what arrives tagged, a mention looks like nothing happened, which is the measurement hole covered in how to measure AI traffic with no referrer.

Tier Three: Recommended, and Written by the Model

The recommendation is not lifted from your page. Profound’s analysis of 50,000 prompts across seven industries found roughly 47 percent of response content was unsolicited editorial: comparisons, rankings, caveats and shopping guidance nobody asked for.

That editorial layer is where the pick gets made, and the model writes it.

  • Cited: the engine used your page to build the answer.
  • Mentioned: the engine named your brand inside the answer.
  • Recommended: the engine told the reader to choose you.

Three states, three different jobs, and the job that moves a brand up a tier is not the one that earned the citation.

Moving a Brand From Cited to Recommended Takes Comparative Evidence

Xi Chu and Yupeng Hou tested three commercial models on skincare recommendations and published the result as Incumbent Advantage in June 2026. With product specifications held identical, the well-known brand was recommended 100 percent of the time.

That lock broke as soon as a competitor carried a rating advantage of less than 0.1 stars.

Read it as a mechanism rather than a skincare finding. Where nothing separates two options the model falls back on brand familiarity, and where something comparable separates them the comparable thing decides it.

A separate arXiv paper on brand bias in large language models found these systems disproportionately associate global brands with positive attributes, which is a headwind if yours is the smaller name in the category.

Which sources get pulled also varies by engine. Semrush’s AI Mode study analysed 5,000 keywords and over 150,000 unique citations, finding Perplexity’s citations overlapped Google’s organic domains by more than 91 percent while ChatGPT tracked Google’s rankings least closely.

The Levers That Move a Tier Are Comparative, Not Volumetric

So the levers that move a tier are comparative rather than volumetric:

  1. Publish the comparable attribute the model needs: ratings, counts, prices, coverage, anything that survives being lifted into a comparison table.
  2. Earn the third-party evidence, because the reviews and roundups an engine reads about you sit on other people’s domains.
  3. Structure the page so the comparison is liftable, which is the ground covered in comparison tables and formats LLMs love to cite.

None of that shows up in a citation count, which is why a property can look busy in AI answers and still lose the recommendation.

I see the same split in my own prompt checks on a food and hospitality property I run, offered here as an illustration of the pattern rather than as evidence for it.

Founder Insight: A Visibility Score Is a Floor Reading, Not a Verdict

A score built from mentions tells you the engines can find you, which is worth knowing and worth defending. It does not tell you whether any of those answers ended with your name as the pick, and that is the only one of the three a buyer acts on.

The Trap: Optimising the Metric You Can See

Chasing citation volume assumes the citation carries traffic. BrightEdge’s tracking through 2025 put AI search at under 1 percent of referral traffic while organic search kept delivering the majority of conversions, and reported AI referral traffic was not converting at the same rate.

Buying more of the cheapest outcome at the highest effort is a poor trade, and it is the trade a citation-only report recommends.

Visibility and search position have also come apart. Similarweb’s 2026 Generative AI Brand Visibility Index benchmarks leaders, fast growers and overachievers across six sectors, where an overachiever is a brand more visible in AI answers than in conventional search.

If those two can diverge that far, neither is a safe proxy for the other.

This is the last of six articles in this cluster, and the one where the argument stops being about which tactic to build and starts being about which number to report. The cluster’s pillar sets out how each claim was tested, and the brand mentions finding explains why mentions earned their place at all.

Three habits keep the trap open:

  • Reporting a single visibility score: one number cannot separate a citation that lost from a recommendation that won.
  • Treating a rising count as progress: volume rises while win rate falls, and the report still reads as a good quarter.
  • Waiting for a tool to solve it: the tracking tools worth comparing differ on what they count, so the definition has to be yours first.

Avoiding all three takes a reporting format, not a platform.

What Good Looks Like When Both Outcomes Are Tracked

Aleyda Solis has published the clearest version of this I have seen. Her three-layer measurement framework splits presence, readiness and business impact, and inside presence it carries a Recommendation Rate: appearances where the AI explicitly recommends the brand, divided by prompts where the brand appears, times 100. A Comparative Win Rate sits beside it for head-to-head prompts.

Two named metrics a founder can hold a tool to, and neither is a citation count.

The Four Fields to Log

The input layer matters as much as the KPI. A representative AI search prompt library fixes the prompts you test so the numbers stay comparable week to week, rather than moving because the question changed.

Run that fixed set on a repeat schedule and log four fields:

  1. Cited: did the engine use your page as a source.
  2. Mentioned: did the answer name your brand.
  3. Recommended: did the answer tell the reader to pick you.
  4. Winner: which brand was named instead when you were not.

The fourth field is the one people leave out, and it is the only one that separates a citation problem from a comparison problem. Those need different fixes, and a rising mention count cannot tell them apart.

Pair the log with referral measurement so the demand side stays visible. Semrush’s guidance on tracking AI referral traffic covers the mechanics, and the discipline of showing the control beside the result is what makes a claim hold up, as in the CTR fix and building the cluster rather than the post.

Fix the comparison layer first if you are cited and never named, and the retrieval layer first if you are neither, which is the sequence set out in how to be the source Claude recommends. If you would rather have that separation run for you against your own category prompts, start with a diagnosis and the split itself is the first deliverable.

Frequently Asked Questions

What is the difference between being cited, mentioned, and recommended by AI?

Being cited means an engine used your page as a source. Being mentioned means it named your brand in the answer. Being recommended means it told the reader to choose you. Only the third is a buying signal, and Profound’s shopping-task research found citation share predicts it at just 0.57.

Which check separates a recommendation from a citation?

Run a fixed set of category prompts across the engines you care about and read each answer twice: once for the source list, once for the body. Record whether your page was cited, whether your brand was named, and which brand was recommended. Repeat the same prompts on a schedule.

What moves a brand from being cited to being recommended?

Comparable evidence does. An arXiv study of three commercial models found that with identical product specifications the familiar brand was recommended every time, and that lock broke on a rating advantage under 0.1 stars. Publish the comparable attribute and earn the third-party reviews an engine reads.

Which prompts should I track for AI visibility?

Track the prompts a buyer would type at the point of decision, not the ones you already rank for. A representative library covers informational, comparison and evaluation phrasing across your category, and stays fixed between runs so week-on-week numbers move with performance rather than wording.

How much click-through should I expect from an AI citation?

Very little. In Profound’s shopping study 92.8 percent of tasks ended with no meaningful click to the open web, and SparkToro puts 68.01 percent of US Google searches in early 2026 in the same bracket. Treat a citation as influence you cannot see, not as traffic.

What should I fix first if I am cited but never recommended?

Fix the comparison layer. Being cited proves retrieval already works, so the gap sits in what a model can say about you against a rival: ratings, counts, coverage, prices, and the third-party reviews it reads. Publish the comparable attribute before publishing more pages.

Continue Reading

More From This Series

More from TDM Insights

Explore TDM Insights Topics