AI Recommendation Verification: Why Mentions Don't Sell

An AI assistant answers a buyer question with a shortlist. Ask for a plumber in a mid-sized city, and the response names three businesses. The buyer treats those names as options. The businesses absent from the list never enter consideration.
A mention in that answer looks like demand. It is visibility. The two diverge when the buyer starts verification. The buyer copies the business name into a search engine. They open review profiles. They check the website for a phone number and service area. They look for a second source that confirms the AI’s claim. If the second source is missing or contradicts the mention, the shortlist collapses.
This is the verification gap. A model names a business because its name appears in training data, a directory, or a retrieved page. The model also names a business with no current website, no recent reviews, and no consistent address. The mention still occurs. The sale does not follow.
aeod.app measures mentions and their position in AI answers. That measurement shows where a business appears. It does not show whether a buyer has a second source to verify the business after the answer ends. The next section defines the verification gap: what changes when a mention has a second source behind it, and how to measure the difference between being named and being confirmed.
What the verification gap measures: mention versus second source
The verification gap is the distance between an assistant naming a business and a buyer acting on that name. The answer ends with a shortlist. The buyer’s work starts after that. They open a new tab and search the business name for a second source. That second source lives on surfaces the assistant answer never shows: review pages, Reddit threads, maps, news, forum posts. The gap is the time and decision space between the name and the confirmation. A buyer who hears one name from an assistant treats that name as a lead. The lead requires a second source before it becomes a call, a booking, or a purchase.
Mention, citation, and position describe the answer layer. Mention counts whether the business appears. Citation records whether the assistant links to a source. Position records where the business sits in the shortlist. Those three metrics describe what the assistant said. They do not describe what the buyer does next. A citation points to a page the buyer never opens. A position of one fails if the buyer finds no independent confirmation. The answer layer is measurable from the assistant output alone. The verification layer is measurable only from the buyer’s next actions.
The two layers disagree inside the same account. A business shows a high mention rate across buyer questions and weak conversion from those same questions. That pattern means the second source did not confirm the first. The assistant named the business. The buyer checked. The check returned doubt or a competitor. The answer layer looked healthy. The verification layer failed. A dashboard shows a high mention count and no lift in qualified calls. The mention count is real. The conversion gap is also real. The pattern repeats across categories, as covered in why AI assistants recommend the same businesses.
Measuring the gap requires tracking both layers. The answer layer comes from model outputs. The verification layer comes from what buyers find after the answer: search results for the business name, review profiles, community threads. The two draw on different data, and one cannot stand in for the other. The next section puts numbers on that behavior: 98% check, 86% verify, half start on Reddit. Those numbers explain why a mention without a second source stays a mention.
The numbers: 98% check, 86% verify, half start on Reddit
The survey base is large enough to set expectations. Idea Grove’s 2026 study of 1,000 US consumers found that only 2% would buy from an unfamiliar brand on an AI recommendation alone. The other 98% took further steps: search, reviews, press, or the brand’s own site. Product.ai’s 2026 survey of 1,463 US shoppers found 86% verified a recommendation through another source. The Reddit Path to Purchase 2026 survey put roughly half of US shoppers fact-checking AI picks on Reddit.
These are self-reported figures. They are directional. They describe stated behavior, not observed purchase logs. The pattern still matters because it repeats across three independent samples. A buyer who receives an AI answer treats that answer as a starting point. The shortlist creates interest. It does not close the sale.
The counterweight comes from Adobe. By May 2026, Adobe found AI-referred retail traffic converting better than non-AI traffic. That finding keeps the verification story from turning into a rejection story. Buyers who check are not dismissing the AI recommendation. They are filtering it. The recommendation survives the check or it does not. The ones that survive arrive with more context and higher intent.
That distinction changes what a business should measure. A mention in an AI answer is an opening. The sale depends on what the buyer finds when they look up the business name. Search results, review profiles, press coverage, and community threads all enter the decision. Reddit carries particular weight because half of US shoppers in the Reddit survey used it for fact-checking. A thread from 2023 can outrank a homepage. A review profile with no recent activity can cancel the confidence the AI answer created.
An AI brand accuracy audit can show what the model says about a business. It cannot show what buyers find after the answer. The verification layer needs its own review, run against search results, review profiles, press pages, and community threads.
The next section follows the buyer after the answer to where buyers verify: the sources the answer never displays.
Where buyers verify: the sources the answer never displays
The answer ends the conversation with the assistant. It starts a second one in a browser tab. Buyers run the brand name through Google or Bing, open review platforms, read two or three community threads, scan recent news, and land on the company’s own pricing and location pages. Those surfaces carry the decision.
Assistants retrieve from the same pool. Retrieval draws on indexed pages, review corpora, forum threads, and news archives. The citation block shows a fraction of what was retrieved, often one or two URLs out of a dozen candidate documents. The rest stays invisible inside the answer.
That invisibility is the trap. A citation the buyer never clicks still shapes what they find when they search the brand. The cited page sits in the same index the buyer queries, so it feeds the results they see, the snippets they read, and the ordering of the page. Coverage that exists only inside the answer leaves no trace on the surfaces buyers actually use.
The gap widens when the answer compresses a claim. Pricing is the clearest case. An assistant summarizing AI answers to pricing questions states one starting figure, while the destination page lists tiers, add-ons, onboarding fees, and annual commitment terms. The buyer arrives expecting a number and finds a table. Location claims behave the same way. An answer says the business serves a metro area; the page behind AI visibility by location lists one address, a service radius, or three offices with different hours. The claim survives the summary and changes meaning at the source.
Review platforms add dates, reviewer history, and owner responses. Community threads add disagreement, including the objections the assistant omitted. News coverage adds context the model had no reason to surface. Each surface is searchable, and each one already contains the pages the assistant retrieved.
Comparing the answer against the sources it retrieved shows whether the claims that sent the buyer searching hold up on the pages they open.
Verification is where the answer and the record meet. Disagreement between them is where mentions die.
When verification fails: the answer and the record disagree
A buyer asks an assistant for a plumber. The assistant says the diagnostic fee is $89. The buyer opens the plumber’s page and sees $129. The assistant says the company serves Denver. The page lists Aurora and Lakewood. The assistant cites 4.9 stars from 812 reviews. The page shows 4.6 stars from 604 reviews. The answer and the record disagree. Verification fails at the moment of purchase. The buyer resolves the conflict in favor of the page they can open.
The conflict has a label. Each claim gets classified as supported, outdated, or fabricated. Supported means the claim matches the current first-party page and third-party sources. Outdated means the claim was true at an earlier date. A retired price is one case. A service area that shrank is another. A credential that lapsed or a review count that changed fits the same label. Fabricated means no source ever carried the claim. The label sets the action. Outdated claims need a source update and a crawl path check. Fabricated claims need a source trace and correction. Supported claims need no action.
Retrieval and recall correct on different clocks. Model recall comes from training data frozen at a cutoff. Retrieval comes from a live index built by crawlers. An assistant running GPT-4o with web search pulls a current page. Its parametric recall still holds the old price. If retrieval returns a directory listing from 14 months ago, the assistant repeats the old figure. The first-party page updated this morning. The index updated last week. The directory updated last year. The assistant answers from the middle of that timeline.
The crawler wrinkle makes the conflict worse. A corrected first-party page stays outside the retrieval path when AI crawlers are blocked. A robots.txt disallow or a WAF rule stops the crawl. Third-party pages keep the old figure. The assistant retrieves the stale third-party page. The buyer opens the live first-party page. The answer looks wrong even though the business fixed its own site. See blocking AI crawlers for the mechanism. Classifying a conflict means comparing the answer against the first-party record and the third-party sources that outrank it.
The resolution costs the mention its authority. The next failure is harder to see because the buyer never opens the page. Ghost citations: the mention the buyer never sees.
Ghost citations: the mention the buyer never sees
A citation is only useful to a buyer who can see it. Most are invisible.
Semrush and Kevin Indig analyzed 3,981 domain appearances across four AI engines. 61.7% of those citations were unattributed. The assistant retrieved the page and used its content. The answer it produced never named the source. The domain did the work and received nothing the buyer can act on. That is a ghost citation.
Verification requires a thread to pull. A buyer checks a claim by searching the brand or opening the URL. An unnamed citation supplies no thread. The buyer receives a synthesized answer and moves on. Verification traffic to the cited page from that session is zero, and it stays zero because the buyer has no reason to suspect the page exists.
The authority flows elsewhere. When an assistant states a fact without attribution, the reader credits the assistant. The page that supplied the evidence adds nothing to its own standing with that reader. Run that across thousands of sessions and the source becomes invisible infrastructure while the answer interface becomes the authority.
Then sources turn over. AI answers reshuffle as models update and indexes refresh. A page cited in March drops out by June. If the citation carried a name, the buyer already formed a memory and can return to the brand directly. If the citation was a ghost, the disappearance removes a path the buyer never knew existed. Source turnover has its own measurement problem, and persistence is tracked separately.
The gap between cited and named is a metric separate from mention rate and position. A business can appear in the retrieval layer of an answer and never appear in the sentence a buyer reads.
How much that costs depends on what the buyer asked. Verification shifts by prompt type.
Verification shifts by prompt type
A price prompt asks for a number. The assistant returns a range. The buyer then checks a pricing page or a quote calculator. That surface confirms the price. It does not confirm whether the vendor supports the buyer’s use case.
A pair prompt changes the destination. “X vs Y” produces a head-to-head answer. The buyer opens a comparison page to check the verdict. That behavior is the subject of AI comparison prompts: X vs Y verdict. The verification surface is a side-by-side table. The buyer wants the row that shows the difference.
A constrained follow-up sends the buyer somewhere else. The first answer names three vendors. The second turn adds a filter: “Which of those has a license in Texas?” The buyer then checks a license registry. If the filter is “Which integrates with HubSpot?” the buyer checks an integration directory. If the filter is “Which handles 500 concurrent users?” the buyer checks a capacity spec. That pattern appears in multi-turn AI visibility follow-up questions. The verification surface is attribute-specific evidence.
These surfaces produce different conversion paths. A price prompt ends on a pricing page. A pair prompt ends on a comparison page. A constrained follow-up ends on a documentation page. A single blended conversion figure sums all three and hides the split. The number looks stable. The underlying behavior changes by prompt class.
Position and citation rates can be read per prompt class, and the finding is simple: verification has to be read the same way. A price prompt needs price evidence. A pair prompt needs comparison evidence. A constrained follow-up needs attribute evidence. When the classes are merged, a drop in one class can be masked by a rise in another.
That split determines what to log. Each prompt class needs its own record of the answer and the verification surface it sends the buyer toward. The next section covers logging the gap: five fields per prompt.
Logging the gap: five fields per prompt
A useful prompt log has five fields. Record prompt wording verbatim, including any follow-up turn that changed the request. The follow-up turn narrows the shortlist and changes which business gets named. Record mention status with the exact wording the answer used. Record the cited source as the exact URL the assistant attached to the claim. Record position as the order in the shortlist, with first place logged as 1. Record the surface the buyer would check next, such as the linked page or local listing the answer points toward.
These fields create a trace. Prompt wording preserves the question. Mention status preserves the answer. Cited source preserves the evidence. Position preserves the ranking. Surface preserves the next click. Without the surface field, a mention looks like a win. With it, the mention becomes a claim that has to hold on a page.
Run each prompt five times as a floor. A single run samples one answer from a distribution. A model with temperature above zero returns different shortlists across runs. Even with temperature at zero, context and retrieval change. Five runs show whether the mention is stable, occasional, or absent. Re-run the same prompts on a schedule. The cited source has a half-life. The URL that supported a claim today decays as pages update, as described in AI citation half-life.
The unit of comparison is the claim in the answer against the page the buyer reaches. If the answer says “same-day service” and the page says “next available appointment,” the mention breaks at the second source. If the answer names a location and the map result shows a closed office, the mention breaks there. The log captures that gap per prompt, keeping the answer and the verification surface side by side.
Where the work starts: make the second source agree.
Where the work starts: make the second source agree
A single mention on a third-party list is a lead, not a sale. The buyer asks the assistant a follow-up: hours, license, service area, price range. If the assistant retrieves a second source and finds a different phone number or a closed location, the mention loses force. The shortlist entry survives only when the sources behind it agree.
Order the repairs. First, update the cited page in place. If an assistant cited a service page, fix that URL. Do not create a new page and leave the old one live. The old page remains retrievable and continues to contradict. Correct the name, address, phone, category, and hours on that page. Then reconcile those same fields across every profile an assistant can retrieve: Google Business Profile, Bing Places, Apple Business Connect, industry directories, chamber listings, review sites, and data aggregators. Each source is a candidate for retrieval.
Where third-party pages carry old figures, place a dated correction. A line such as “Updated March 2025: phone number changed to…” gives the retriever a timestamp and a reason to prefer the current value. Without a date, two pages with different numbers look equally valid. The assistant has no basis to choose the newer one.
These are the same inputs that decide shortlist presence. A mention and its verification stand or fall together. The assistant does not separate the recommendation from the supporting data. If the cited page says one thing and the profile says another, the recommendation becomes fragile. The buyer sees the inconsistency and leaves, or the assistant drops the business from the next answer.
The work is unglamorous. It is field matching and source hygiene. aeod.app assesses how assistants answer buyer questions by measuring mentions, citations, and position, then ties findings to a prioritized action list. That list points to the pages and profiles that disagree. The repair is manual. No tool can force an assistant to cite a page.
The limit is structural. The source pool gets rebuilt on every question, so a single edit cannot guarantee the next answer. The goal is to remove contradictions that make a business easy to skip. The next section takes up what to do with this: how to order the repairs when time and attention are finite.
What to do with this
The assistant answer and the purchase decision are separate events. The answer produces a shortlist. A buyer then opens a second source: the cited page or a review profile. That second-source check decides whether the shortlist becomes a call or a form fill. A mention opens a door. The buyer’s verification decides whether anyone walks through it.
The working conclusion is a measurement habit. Read the answer for a buyer question. Then read the page the answer points to. Compare the claims. If the assistant says the business offers emergency service and the page says “call during business hours,” the disagreement is the finding. If the assistant cites a service area that the page omits, that mismatch is the finding. If the assistant names a competitor and cites a comparison page that lists the business with stale hours, the stale source is the finding. The useful output is the contradiction a buyer would hit during verification.
Run that habit across a set of real questions: measure mentions, citations, and position for each, then map every mismatch to a prioritized action list. The list orders work by the gap that appears in the answer and on the cited page. A page that repeats the facts in the assistant’s cited source has less friction. A page that conflicts with that source creates a verification failure at the exact moment the buyer is checking.
The constraint remains. Retrieval is rebuilt per question. An assistant assembles the answer from sources available at that moment, using the query wording and session context. An edit changes one source in that pool. It does not guarantee the next answer. It does not guarantee a purchase. A corrected page can still lose to a stronger competitor page. A verified profile can still miss a question phrased in an unexpected way. The aim is to remove contradictions that make a business easy to reject during verification. A cleaner source pool raises the chance that the answer and the page agree. The buyer still decides.
Your own answers
See what AI says about your business.
One domain, one assessment. Mention and citation rates, competitors named instead of you, and a prioritized action list.
Get your report · $29 ↗More notes

31 Aug 2026 · 13 min read
Why AI Assistants Recommend the Same Three Businesses
Why AI assistants recommend the same businesses, explained by concentration data, retrieval sources, and entity consistency.

16 Aug 2026 · 14 min read
Multi-Turn AI Visibility: One Follow-Up Erases the List
Multi-turn AI visibility data: one added constraint removes 62% of named brands, and single-prompt tracking hides the drop.