AI Brand Accuracy Audit: Wrong Facts in AI Answers

An AI assistant asked for a plumber returns three names. If one of those names carries a wrong address or an expired license, the buyer sees a factual answer. The error does not announce itself. The shortlist still looks authoritative. The business either loses the call or spends the first minute correcting the record.
Wrong facts in AI answers come from retrieval and synthesis. A model retrieves web pages, directories, review sites, and cached snippets. It compresses that material into a sentence. When sources disagree, the model picks one version. When a source is stale, the model repeats it. The result is a confident answer with a specific phone number that no longer rings.
Most teams track traffic and rankings. Few track whether an assistant states the correct service area, hours, credentials, or price range. A wrong fact in an AI answer costs more than a missing mention because it disqualifies the business inside the buyer’s decision. The buyer acts on the answer before visiting the website.
aeod.app found this pattern by asking assistants buyer questions and checking the factual claims in their answers. The finding is narrow: accuracy is measurable at the claim level.
The next section defines what a brand accuracy audit measures: the specific claims and source types that shape an AI answer.
What a brand accuracy audit measures
A brand accuracy audit replays the same buyer prompts used for visibility tracking. The prompts stay fixed. The assistant answers get read line by line. For each answer, the audit extracts every verifiable statement about the business. Price. Plan tiers. Features. Service area. Credentials. Ownership. Hours. Refund policy. Any detail a buyer verifies. Each claim gets a label: supported, outdated, or fabricated.
Supported means the claim matches a current primary source. The business site, a regulatory filing, a licensed directory, a dated press release. Outdated means the claim was true at some point. The assistant repeats an old price, a discontinued tier, a former address. Fabricated means no source supports the claim. The assistant invents a service area, a certification, a partnership, a founder name. The distinction matters because the fix differs. An outdated claim needs a fresh source. A fabricated claim needs a correction at the source the assistant used.
The audit does not stop at the label. It pairs each claim with the source the assistant drew on. That source falls into known types: business website, review platform, directory listing, news article, forum thread, cached snippet. The pairing shows where the answer came from and which page needs attention. A wrong price on a review site creates a different action than a wrong price on the business’s own pricing page.
The output is a claim ledger: a list of specific statements, each with a status and a source. A business ranks well and still carries a fabricated credential. A business is absent from a shortlist and still has accurate claims. The audit separates those conditions. The same buyer prompts that track whether a business appears also track whether the answer gets the facts right. This is why claim-level review sits beside mention tracking. A business appears in the same shortlist repeatedly for reasons covered in why AI assistants recommend the same businesses. aeod.app runs this pass by replaying the prompts and extracting the claims. The result feeds a prioritized action list: fix the source, update the page, correct the directory, or leave a supported claim alone.
That ledger sets up the next layer: mentions, citations, position, and a fourth field: claims.
Mentions, citations, position, and a fourth field: claims
A business can hold position two in a three-name shortlist and still be described incorrectly. Mention tracking reads the name. Citation tracking reads the source attached to the name. Position tracking reads where the name sits in the answer. Claims tracking reads the sentence built around the name. Those are separate fields, and they can point in opposite directions without either one being wrong. A mention count of one and a position score of two say nothing about whether the answer states the business charges $95 for a diagnostic or serves only commercial clients. The name is present. The description is wrong. Both readings are true.
An audit separates those readings. In one answer, presence scores well: the business appears, the model cites its homepage, and the position is second. Accuracy scores poorly: the answer repeats a price the business retired 18 months ago. No contradiction exists between the two scores. The presence field answers “Is the name here?” The claims field answers “What does the answer assert about this business?” A report that stops at presence records a win while the buyer receives a false fact.
Pricing is the claim type that breaks most visibly. A page changes from $99 to $149. The old figure stays in cached copies, quoted passages, and retrieval indexes. The assistant repeats $99. The business keeps its shortlist rank. The claim attached to that rank is stale. The guide to AI answers to pricing questions walks through that mechanism. The citation half-life explains why a cited source can remain influential after the page updates.
Claims tracking records the sentence and its source, plus the date the audit found it. It treats pricing as one claim type among several. The report can show strong mention coverage and weak claim accuracy in the same answer. That split leads into two knowledge paths: recall versus retrieval.
Two knowledge paths: recall versus retrieval
An AI answer pulls a fact from one of two places. The first is model weights, the associations baked in during training. The second is a passage fetched at answer time, often from a web page or document. The paths fail at different rates, and the prompt decides which path carries the answer.
Ask “tell me about Acme” and the assistant has little to retrieve. It generates from recall. Benchmarks of open-ended prompts have measured hallucination rates across a wide range, from roughly a fifth of answers to nearly all of them, depending on the model. That range is wide because recall depends on training coverage, model size, tuning, and how often the entity appeared in the corpus. A company with sparse training data gets a confident sketch with wrong pricing, wrong leadership, wrong locations. The model has no passage to check against.
Ask “what does Acme’s pricing page say about the Pro plan” and the assistant retrieves. The answer is grounded in a passage. Error rates in page-grounded settings fall to a fraction of the open-ended range. Retrieval still fails. The page may be stale, the snippet may omit the relevant sentence, or the retriever may fetch a competitor’s page. The failure mode changes from fabrication to misreading. That difference explains why question wording changes accuracy. A prompt that names a source or asks for a page-grounded fact gives the model a target. A prompt that asks for a summary gives recall room to fill gaps.
For an accuracy audit, the split matters. An audit separates claims that appear in open-ended answers from claims that appear in page-grounded answers. An audit that tracks both paths will show different accuracy by prompt type. The same business can look accurate when the prompt names a source and inaccurate when the prompt asks for a general description. Tracking which path produced a claim also sets up the citation half-life question: a retrieved page can go stale, and a recalled fact can drift as models update. The next step is to sort claims by type. Which claims break first: pricing, features, location, people.
Which claims break first: pricing, features, location, people
The first wrong fact a buyer hits is usually a price. A model asked “what does this cost” pulls the most repeated number across its retrieved pages. If a 2023 rate sits on three directories and the current rate sits on one pricing page, the old number wins on weight. The buyer receives a number that is no longer offered. They anchor on it, ask for it, or leave. That is a blocked sale. The pricing error class has its own prompt pattern; see AI answers to pricing questions.
Discontinued features come next. A product page that stays live after a feature sunset becomes a source. The assistant answers “yes, it supports X” because the page says so. The buyer books a demo around a capability that no longer exists. The cost lands in wasted evaluation time and a trust drop when the demo team corrects the record. The action difference is direct: the buyer would have chosen a different vendor or plan if the feature fact were right.
Service areas rank high because the wrong answer stops contact. Location claims conflict across profiles. A business with a service radius, multiple offices, or recent expansion gets different boundaries from its website, its Google Business Profile, and third-party directories. A buyer in a covered ZIP is told the business does not serve that area. The buyer does not call to check. They move to the next name. Location conflicts also vary by prompt; AI visibility by location covers how the same business can appear in one market and vanish in another.
Credential and leadership details change years ago and persist. An old bio names a founder who left. A license number appears on a state board page but not on the company site. A buyer asking “who runs this” or “are they licensed” gets a stale fact. The commercial cost is lower than a wrong price. The buyer still calls, then asks for verification. For regulated categories, a wrong license blocks the sale. For most, it weakens trust.
Founding year and award years sit at the bottom. A wrong founding year changes no next step. A buyer does not choose a vendor because it started in 2009 instead of 2011. The cost is reputational only when the error is obvious.
Rank each class by buyer action. Wrong price blocks a sale. Wrong feature fit wastes a sales cycle. Wrong service area removes the business from consideration. Wrong credential adds friction. Wrong founding year does almost nothing. The common thread is source conflict. Most wrong facts trace to conflicting sources.
Most wrong facts trace to conflicting sources
Assistants do not invent facts from a void. They reconcile records. The mechanism is entity resolution: the model collects name, address, phone, category, hours, and license data from Google Business Profile, the company website, directories, review sites, and licensing boards. Then it decides whether those records describe one business or several. A phone number written as (555) 123-4567 in one source and 555-123-4567 in another is a minor mismatch. An address written as Suite 4 and #4 is a format conflict. A category label of plumber in one profile and contractor in another is a classification conflict. Each mismatch lowers confidence. Enough mismatches split the entity in two. The assistant then answers about one fragment and omits the other.
Name collisions produce worse errors. Two businesses share a name. One has a phone number. Another has an address. A third has a license. The model merges those facts into a single confident claim. The claim belongs to someone else. The buyer gets a wrong number or a wrong service area. In audits, these errors cluster around profile fields. The answer text is a symptom.
The practical consequence is direct. The repair target is the profile field. Editing the answer text changes nothing because the assistant rebuilds the answer from source records on the next retrieval. Fix the name, address, phone, category, and hours across every source that feeds the entity. The same address-format mismatch appears in AI visibility by location. Name collisions show up in AI comparison prompts, where merged facts decide a winner. Once the records agree, the remaining variable is time. A retrieval source updates in weeks. A fact inside model recall takes months. That gap sets the next measurement: correction clocks. Retrieval changes in weeks. Recall changes in months.
Correction clocks: weeks for retrieval, months for recall
The repair target splits into two clocks. An audit separates them per claim because the timelines do not match.
The retrieval clock runs on weeks. Live-search answers from Google AI Overviews and Bing Copilot assemble a response from indexed pages and entity records at query time. When a source record changes, the assistant reads the new value after the crawler revisits and the index refreshes. In practice that window lands between two and five weeks. A source edit on Monday does not appear in an AI Overview on Tuesday. The crawler revisits, the index rebuilds, and the answer regenerates. The mechanism is corrected source, crawl, index, answer. The source must remain reachable. Crawl budget sets the floor. Index refresh sets the ceiling.
The recall clock runs on months. ChatGPT and Gemini store factual patterns in model weights from pretraining. A source correction does not rewrite those weights. The next training cycle or fine-tune carries the update. Model providers release training updates on their own schedules. The gap between releases is measured in months, and a brand cannot book a slot. Major providers do not publish a brand correction portal or ticket queue for factual errors. OpenAI support handles account and policy issues. Google and Microsoft run comparable queues. No queue exists for a wrong phone number in a model. A wrong fact in recall survives after the web page is fixed.
Blocking AI crawlers removes the mechanism a correction would use to reach the index. If robots.txt disallows the crawler, the corrected page never enters the retrieval corpus. The retrieval clock stops. The recall clock loses the fresh signal that feeds a future training cycle. A corrected page behind a crawler block is invisible to the retrieval layer. The assistant keeps the old value from the last crawl. The source fix exists on the origin server and nowhere the assistant reads. This is the wrinkle in blocking AI crawlers: the block prevents the repair from traveling.
Two clocks produce two audit columns. A retrieval claim needs source verification and a crawl check. A recall claim needs repeated, consistent source signals across months. The next step is running the audit: prompt set, claim log, severity.
Running the audit: prompt set, claim log, severity
Start with ten to twenty buyer prompts written in the customer’s own wording. Cover service questions, price questions, availability questions, and trust questions. Pull phrases from sales calls, search queries, review language, and support tickets. Include multi-turn follow-up questions and comparison prompts such as X vs Y verdicts. Run the same set across ChatGPT, Claude, Gemini, and Perplexity. Repeat on separate days. Answers vary between runs because retrieval indexes change and model sampling shifts. A single run gives a snapshot. Three runs on three dates give a pattern. Five runs give a stronger pattern.
For every answer, log each factual claim. A claim is a sentence the assistant states as true: a price, service area, license number, owner name, years in business, warranty terms, or response time. Record the model, run date, cited source, and severity tag. The cited source matters. If the assistant cites a directory page from 2021, the wrong fact has a retrieval anchor. If no source appears, the claim came from model recall. Severity depends on buyer consequence. A wrong license number gets a high tag. A stale service area gets medium. A phrasing mismatch gets low. A missing fact gets a separate tag because absence changes the shortlist.
aeod.app records mentions, citations, and position per prompt. The claim log sits alongside that output. A wrong fact can be read next to the position it appeared in. If a high-severity error appears in the first named answer for a prompt, the risk is larger than the same error in a fifth-position mention. Position changes the cost of the error. A wrong price in first position reaches the buyer before the next name. A wrong price in fifth position reaches fewer buyers.
The audit output is a table with prompt, model, date, claim, source, severity, mention, citation, position. That table shows what is wrong and where it surfaced. The next question is where corrections actually take hold.
Where corrections take hold
The cited URL is the first place to fix. When an assistant cites a page, that page is the retrieval target. Editing it in place changes what the next crawl returns and preserves the URL’s history. Publishing a new page alongside it splits the signal. The old URL still carries the wrong fact, still gets retrieved, and now two pages disagree about the same address or phone number. Update the original, and confirm the corrected line sits in the first screen of content, where extraction reads.
Second, reconcile identifiers. Assistants assemble a business from many retrievable sources. Google Business Profile, Bing Places, LinkedIn, and the site’s own schema markup all get read. So do industry directories and review pages. An assistant that finds “Suite 400” on one profile and “Suite 410” on another has no procedure for deciding which is current. Audit the name, address, phone, hours, and category strings across every profile you control and make them byte-identical. Where a profile cannot be edited directly, claim it or submit a correction through its own process.
Third, a third-party page carrying the wrong fact needs a sourced correction. Directory listings and news pages are the pages a model quotes when it explains a claim. If the operator will not edit the page, leave a dated, sourced correction where that page’s crawler will read it: a comment, a press page, or a public statement with the correct figure and its source. That gives the next retrieval pass a competing claim with a citation attached.
Then re-run the exact prompts from the audit. Record four fields: whether the claim flipped, which models reflected the change, the date, and the days elapsed since the fix. The audit table supplies the baseline row; the re-run supplies the after row. The crawl question applies here too, since a fix on a page no assistant can retrieve changes nothing (blocking AI crawlers).
Segment the re-run by prompt type. A single-sentence category question updates first, because it draws on the broadest retrieval. A constrained follow-up such as “is that location open on Saturday” pulls narrower sources and holds the old fact for weeks after the general answer corrects. Log both rows separately. Averaging them hides the split.
The audit finding is that wrong facts are traceable. That does not make them easy to erase. Accuracy is maintenance because the sources behind an AI answer age at different rates.
A retired price on a page used by a retrieval platform is typically corrected after the next crawl and index refresh. That cycle runs in weeks. The same price in a model’s training data persists until the next training run or until the system retrieves a fresher page. Major model training cycles run in months. An answer that mixes both paths cites the new page while still repeating the old number in prose. The two clocks explain why one cleanup pass fails.
The working habit is a loop. Log the claim exactly as the assistant states it. Name the source the assistant cites or the source that supplies the fact. Fix that source at the point of error. Update the price. Correct the spec. Remove the retired service. Add the missing qualifier. Re-run the same prompt and record the new answer. The comparison separates a source problem from a model recall problem. If the cited page changes and the answer follows within a retrieval refresh cycle, the fix landed. If the answer holds the old fact after the source is current, the fact is sitting in model recall or in another source the first citation did not reveal.
The audit supports this loop by connecting each wrong claim to the cited sources and the prompt that produced it. The output is a maintenance queue for repeated checks.
The limit is simple. No page edit guarantees the next answer changes. A model drops the corrected source or cites a different page. It blends the old number with new context. The audit shows whether the answer moved and which source it used. It cannot force the next generation. Treat accuracy as a recurring check against a moving set of sources, and the wrong facts become manageable without becoming permanent.
Your own answers
See what AI says about your business.
One domain, one assessment. Mention and citation rates, competitors named instead of you, and a prioritized action list.
Get your report · $29 ↗More notes

31 Aug 2026 · 13 min read
Why AI Assistants Recommend the Same Three Businesses
Why AI assistants recommend the same businesses, explained by concentration data, retrieval sources, and entity consistency.

16 Aug 2026 · 14 min read
Multi-Turn AI Visibility: One Follow-Up Erases the List
Multi-turn AI visibility data: one added constraint removes 62% of named brands, and single-prompt tracking hides the drop.