AI Visibility Across Assistants: Why One Check Misleads

A buyer asks ChatGPT for a plumber in a mid-sized city. The reply names three companies and links two of them. The same question typed into Perplexity returns a different set: one overlap with the first list, two names that never appeared in the first answer. Google AI Overviews, drawing on a different index of pages, produces a third variation.
The owner who checked once, on one assistant, saw a favorable answer and concluded the work was done. The owner who checked once and saw nothing concluded the opposite. Both conclusions rest on a sample of one.
Assistant answers are assembled per query from whatever the system retrieves and how the prompt is worded. Recency signals shift the mix week to week. Different assistants pull from different corpora and cite different pages, so a single check captures one path through that machinery at one moment. It does not describe the range of answers the market actually receives.
This matters because a shortlist is where demand gets allocated. Absence from one assistant’s answer is a narrow finding. Absence from most of them is a pattern, and patterns are what a business can act on. aeod.app was built around that distinction. It queries the major assistants, records the mentions it finds along with the cited pages and the position of each mention, and reports the spread across systems instead of one result.
Reading that spread takes a defined method. The next section sets out what it means to measure AI visibility across assistants.
What it means to measure AI visibility across assistants
Measuring AI visibility across assistants starts with a fixed buyer prompt set. One list of questions. The same questions run against ChatGPT, Claude, Gemini, and Perplexity separately. Each returns a shortlist. Then the lists are compared. The comparison does not collapse into one brand score.
The unit of analysis is the assistant-and-prompt pair. For a prompt like “best payroll software for a 12-person restaurant,” ChatGPT returns one set of names. Gemini returns another. Perplexity cites live sources. Claude draws on its training data. A brand can appear third on one assistant and absent from another. Those are two separate observations. A brand-level score would average them into something like partial visibility. That average discards which assistant gave the answer and which prompt triggered it.
Why the pair matters: each pair is a distinct answer surface. The assistant’s retrieval mechanism and citation behavior shape the shortlist. Google’s Gemini can pull from Google Search results. Perplexity returns citations to web pages. ChatGPT’s answer can change across model versions. A brand that ranks in one surface does not automatically rank in another. The actionable question is specific: On which assistant, for which buyer prompt, does our name appear, and in what position?
That is the axis a blended score flattens. The same way a national score flattens a local one. A national average can hide a store that ranks first in Boise and nowhere in Miami. The local gap is the work. AI visibility by location covers that version of the problem. The assistant axis has the same structure. A single blended number hides that a business is named by Perplexity for procurement questions and missing from ChatGPT for the same questions. The team cannot fix the missing answer if the score reports only a midpoint.
aeod.app runs a fixed prompt set against each assistant separately. It reports mentions, citations, and position for each assistant-and-prompt pair. That output is a set of lists, not a single number. The lists show where the answers agree and where they diverge.
The divergence is the finding. The next section reports the measured disagreement: 46% agreement, 47% named by one.
The measured disagreement: 46% agreement, 47% named by one
A buyer question sent to four assistants produces four shortlists. aeod.app measured overlap across ChatGPT, Gemini, Claude, and Perplexity. Only 46% of the recommendations were consistent across those assistants. That means for every recommendation that appeared in common, another appeared in a different combination or in only one place. The shortlist depends on which assistant answered.
The numbers get sharper under controlled conditions. In a study where ChatGPT and Gemini received identical sources, they agreed on 64% of the products they named. The inputs were identical. The source set was the same. One third of the named products still diverged. The assistants did not see a different web. They processed the same material differently.
Across all three assistants in that study, 47% of products were named by exactly one assistant. Nearly half of the available shortlist existed as a single-assistant result. A business can appear in ChatGPT and go unmentioned in Gemini and Claude. If the buyer uses one of those assistants, the business is absent from consideration. This pattern shows up in prompt-level tests too. The same X vs Y prompt can yield different verdicts depending on the assistant and the sources it retrieves, as covered in ai comparison prompts X vs Y verdict.
The practical implication is measurement design. Checking a single assistant shows roughly 60% of the picture. The remaining 40% sits in assistants that were not queried. A one-assistant audit reports a complete-looking shortlist with a systematic blind spot. The blind spot is not random. It clusters around source pools and retrieval paths, which is why citation age and source churn matter over time. A source that supports a recommendation today can lose influence as assistants refresh their indexes, a dynamic described in AI citation half-life.
aeod.app produced these overlap figures by querying multiple assistants against the same buyer questions and comparing named entities. The output is a disagreement map: which businesses appear everywhere and which appear once, with the source patterns behind each appearance. That map shows where visibility is concentrated and where it is missing. It does not guarantee a mention. It shows the gap between one assistant’s answer and the market’s answer.
The measured disagreement is the baseline. It explains why a single check misleads. The next question is mechanical. Why do assistants with access to similar sources produce different shortlists? The answer sits in retrieval, memory, and source pools.
Why the assistants diverge: retrieval, memory, and source pools
The disagreement is mechanical. Each assistant runs a different retrieval stack. Perplexity behaves like a literature review. It issues live searches and cites sources inline. Its candidate pool is whatever the live web returns at query time. ChatGPT with browsing synthesizes across retrieved pages and its own weights. Citations appear, though less prominently than Perplexity’s. The answer reads as a summary. Claude weights model recall more heavily. It draws on training data and a narrower retrieval step. Gemini inherits Google’s index and AI Overviews behavior. It sees Google’s crawl and ranking systems. Four retrieval bets, four candidate pools.
Same brand, different bar. A brand with strong recent coverage clears Perplexity’s live retrieval. That same brand fails ChatGPT if the coverage lacks the comparison language the model uses to synthesize. Claude names a brand from older training data when current pages are sparse. Gemini surfaces a brand that ranks in Google’s index and appears in AI Overviews. The shortlist changes because the inputs change. aeod.app measures these differences across assistants and prompts; the finding is that a single check samples one retrieval path.
Crawler access adds asymmetry. If a site blocks AI crawlers, the effect lands unevenly. Perplexity’s live retrieval loses a source it would have consulted. ChatGPT browsing loses a page it would have synthesized. Claude’s model recall still names the brand from training data. Gemini’s Google index keeps the page from standard search crawling. The same block creates different deficits. That uneven retrieval is why crawler policy and AI visibility are linked, as covered in blocking AI crawlers.
The candidate pool is the mechanism. Assistants do not share one source pool. They build separate pools from live retrieval, model memory, and index inheritance. A brand that clears one assistant’s bar fails another because the bar sits on different evidence. Averaging those results produces a number that describes no assistant’s actual shortlist. The next section examines what a blended visibility score hides.
What a blended visibility score hides
Take four assistants a buyer set uses: ChatGPT, Gemini, Perplexity, and Copilot. Score a first-place mention as 100 and an absence as 0. A brand named first by ChatGPT and Gemini and absent from Perplexity and Copilot earns a mean of 50. That 50 appears in a dashboard as a middling result. No buyer sees a middling result. ChatGPT users see the brand at rank 1. Gemini users see it at rank 1. Perplexity and Copilot users see no brand. The underlying pattern is bimodal: two ceiling placements and two floor placements. Averaging draws a flat line through the gap.
The weighted version is worse. If a buyer works inside Microsoft 365, Copilot sits in Word and Teams. That buyer’s shortlist comes from Copilot. A first-place mention in Perplexity carries no weight for that buyer. The same logic applies to a developer who opens ChatGPT for code questions or a researcher who starts in Gemini. A blended score treats every assistant as interchangeable. Assistant choice follows the buyer’s workflow: the tool appears where the work happens. A mention in an assistant the buyer never opens reaches no one.
The useful question is coverage. A brand present in two of four assistants has 50 percent coverage. That number describes a portfolio. A mean of 50 also arises from a brand ranked second or third in all four assistants. Same score, different risk. The first brand is on two shortlists and invisible on two. The second brand is on four shortlists at lower positions. The arithmetic collapses both into one figure. aeod.app reports per-assistant mentions, citations, and position because the distribution carries the finding. The per-assistant record also exposes the verification gap (/blog/ai-recommendation-verification-gap): a citation in one assistant’s answer does not confirm that the same source entered another assistant’s retrieval pool.
Portfolio coverage maps each assistant to the buyer segments that depend on it and marks the gaps. A mean cannot answer that. The next question is why the pattern repeats: the source pool is shared, and unusually concentrated.
The source pool is shared, and unusually concentrated
Two findings look contradictory. Assistants return different brand lists for the same buyer question. Those same assistants pull from a narrow, overlapping set of sources. The 5W AI Platform Citation Source Index resolves the tension: the top 15 domains absorb roughly 68% of citation share, and Reddit is the top source across platforms. The pool is small and shared. The weighting inside the pool differs by assistant.
Each assistant runs its own retrieval and ranking stack. It assigns authority scores to domains and applies freshness rules. Prompt handling is assistant-specific. Reddit threads receive one weight in Assistant A and another in Assistant B. Review sites and vendor pages receive different weights again. The same 15 domains appear across outputs, yet their order and frequency shift. Identical inputs produce different shortlists because the ranking layer changes what the shared pool contributes.
This matters for any business that wants to appear in AI answers. If the business is absent from the concentrated sources, it is absent from the pool many assistants draw from. If it appears in those sources, each assistant still decides how much that appearance counts. A single assistant check reports one weighting. It does not report the shared source pool or the other assistants’ weights. The repeated recommendation pattern across assistants is covered in why AI assistants recommend the same businesses.
For measurement, aeod.app maps mentions, citations, and position per assistant, then ties the gaps to a prioritized action list. The output is a source map with different weights per assistant. A business is cited in one assistant and absent in another while both cite Reddit and the same top domains. That is the contradiction resolved: shared inputs, different weights.
The practical response is to read each assistant separately. The next section covers reading a per-assistant report: four fields worth logging.
Reading a per-assistant report: four fields worth logging
A per-assistant report is only useful if it separates the assistants. The first field is per-assistant mention rate: for each assistant, the share of prompts where the brand appears in the answer. ChatGPT and Gemini can return different mention rates on the same prompt set because they retrieve and rank sources differently. Those two numbers describe different systems and different retrieval paths.
The second field is cross-assistant agreement rate: the share of prompts where every assistant tested names the brand. If four assistants answer the same prompt and only one names the business, agreement is 25 percent. That number tells you whether the brand has a shared presence or a single-assistant artifact.
The third field is exclusive-naming share: the share of prompts where exactly one assistant names the brand. A high exclusive share points to an assistant-specific source or retrieval pattern. The fourth field is first-position share: for each assistant, the share of prompts where the brand is the first business named. Assistants return ordered lists, and the first name receives the first evaluation. A brand can have a high mention rate and a low first-position share, which means it is named after competitors.
aeod.app records mentions, citations, and position per assistant and per prompt, so agreement and exclusivity can be computed from the log rather than guessed. The citations matter for the same reason: if one assistant cites a directory page and another cites a review site, the split has a source-level explanation. For a related accuracy check, see /blog/ai-brand-accuracy-audit.
The fields stay comparable only if the prompt wording is held constant across assistants. Change “best CRM for small law firms” to “top CRM for law offices” and the assistant sees a different query and returns a different shortlist. Multi-turn sessions add another variable, since follow-up questions can shift the context. That behavior is covered in /blog/multi-turn-ai-visibility-follow-up-questions. Log the exact prompt, the assistant, the date, and the raw answer. Then compute the four fields. The next section turns to the case that makes the split most visible: when a brand appears on one assistant only.
When a brand appears on one assistant only
An exclusive mention is a diagnostic. A brand that appears in ChatGPT and nowhere else has a source map, not a victory. The pattern tells you which retrieval path one assistant follows and which paths the others ignore.
Exclusive names usually sit on a source type that one assistant weights heavily. ChatGPT retrieves Reddit threads and forum discussions at high frequency. Perplexity cites structured directory pages and comparison posts. Gemini pulls from Google Business Profile and indexed local pages. Claude tends to weight long-form documents and well-structured pages. If the only mention lives in a Reddit comment, ChatGPT can retrieve it while Perplexity and Gemini return an empty shortlist. If the only mention lives in a niche directory, Perplexity can surface it while ChatGPT returns nothing. The mechanism is retrieval weight. Brand quality does not determine retrieval.
Run the test with a fixed prompt across four assistants. Record the assistant and date with the raw answer. aeod.app does this across assistants and ties the result to a prioritized action list. When one assistant names the brand and three do not, the finding is source concentration.
The decision splits. Broaden the source base so other assistants can retrieve the same claim. That means placing the claim on source types each assistant already weights: a forum thread with real replies and a directory page with consistent categories. A pricing page with explicit numbers also gives assistants a structured claim to cite. The internal link /blog/ai-answers-to-pricing-questions shows how assistants answer price questions from different source types. Or accept a narrow position. If the buyer segment uses one assistant, the narrow position matches the audience. If the buyer segment spreads across four assistants, the narrow position leaves three shortlists unaddressed.
An exclusive mention can also carry an exclusive error. A single source with a wrong price or a stale address can feed one assistant and never get corrected by the others. The other assistants return nothing, so no contradicting source exists in their retrieval sets. That error then becomes the only available claim for that assistant. Check the cited source, not just the mention.
Setting a per-assistant cadence follows from this. Each assistant has its own retrieval behavior and its own update cycle. A monthly check on all four is the baseline. If one assistant changes its shortlist while three stay stable, the source concentration is moving. The next section covers how to schedule those checks without overreacting to a single snapshot.
Setting a per-assistant cadence
A single audit expires because AI assistants do not share one update clock. Live-retrieval assistants read the current index when they answer. Perplexity and ChatGPT with search enabled query fresh results at response time. Google AI Overviews draws from the current search index. Their shortlists move with the index. Weekly checks match that cycle. A source that appears on Monday drops by Friday when a new page earns the citation.
Assistants that lean on model recall operate on training cycles. Claude and some Gemini modes answer from weights updated at training cutoffs. Their shortlists change when the model refreshes, monthly or quarterly. A monthly check matches that cadence. The distinction matters because a weekly check on a recall-heavy assistant produces noise. A monthly check on a live-retrieval assistant misses the window when a competitor replaces your source.
aeod.app reports a 40 to 60 percent month-to-month shift in citation sources. That shift is the reason a one-time audit expires. The citation half-life is short enough that a source holding position in January drops by March. The replacement comes from a different domain type, such as a directory or a comparison post. The mechanism is retrieval. The assistant pulls what the index ranks for the query at that moment. When the index changes, the citation changes.
Log which assistant moved, which source replaced yours, and when. That log becomes the cadence record. If Perplexity shifts twice in a month while Claude holds steady, the live-retrieval layer is moving. If Claude shifts after a model update, the recall layer changed. The log separates a stable shortlist from a volatile one. It also shows whether the replacement source is a direct competitor or a directory. That pattern determines whether the response is content or source work.
The audit produces the log by assistant. Use the weekly run for live retrieval and the monthly run for model recall. Review the log after each cycle. A single snapshot is not a trend. Two cycles on the same assistant show direction.
The schedule tells you when to look. What to do with this depends on what moved. A source replacement in one assistant calls for one response. A replacement across three assistants calls for another. The next section covers how to turn the log into a prioritized action list without overreacting to one check.
What to do with this
Treat each assistant as its own market. ChatGPT and Perplexity build answers from different indexes and retrieval paths. A blended score treats them as a sample of one and answers a question no buyer asks. Buyers ask one assistant at one moment. The useful output is per-assistant coverage mapped to the buyer prompts where the brand is absent. If a plumbing company appears in ChatGPT for “emergency plumber near me” and is absent in Perplexity for the same prompt, that gap is the finding. aeod.app reports that coverage by assistant and prompt, which makes the next action specific: fix the source Perplexity retrieves, or add the page ChatGPT cites.
Track the disagreement rate over time. Count the share of prompts where two or more assistants return different shortlists. A disagreement rate that rises across monthly checks tells you the assistants are diverging. A single check would have hidden that movement. A stable rate tells you the market is holding.
Keep the limit honest. Assistants rebuild answers at retrieval time, so per-assistant coverage is a reading to repeat. It is not a fact to lock in. The same prompt can return a different shortlist next week because the index changed or a source dropped out. Schedule the next check, log what moved, and respond to the movement. That is how the log becomes a prioritized action list without overreacting to one check.
Your own answers
See what AI says about your business.
One domain, one assessment. Mention and citation rates, competitors named instead of you, and a prioritized action list.
Get your report · $29 ↗More notes

31 Aug 2026 · 13 min read
Why AI Assistants Recommend the Same Three Businesses
Why AI assistants recommend the same businesses, explained by concentration data, retrieval sources, and entity consistency.

16 Aug 2026 · 14 min read
Multi-Turn AI Visibility: One Follow-Up Erases the List
Multi-turn AI visibility data: one added constraint removes 62% of named brands, and single-prompt tracking hides the drop.