AI Visibility by Location: Why Brand Scores Mislead

A brand score compresses every market into one number. When a buyer asks ChatGPT or Gemini for a plumber in Austin, the assistant assembles an answer from local signals: the city in the prompt, the user’s approximate location, local business listings, and pages that mention that city alongside the service. Change the city to Denver and the shortlist changes. The brand stays the same. The visibility does not.
aeod.app measures AI answers by location, tracking mentions, citations, and position across prompts. In one multi-location review, a national brand showed healthy aggregate mention rates while individual metros had no appearances for high-intent queries. This pattern repeats whenever a company operates in several cities or has local competitors with stronger citations.
A single score also masks the shape of demand. A query for “emergency electrician near me” draws on different sources than “best electrician for panel upgrade.” Location adds another split: same service, different city, different retrieved pages, different competitors named. The aggregate score blends these into a mean. A mean cannot show which branch or city is missing.
That gap matters because AI shortlists are local. A buyer in one metro sees one set of names. A buyer in another metro sees another. The brand score says the company is visible. The location-level answers say where. Which locations carry the score, and which vanish inside it, is what the average cannot show.
A brand score averages away the branches that never appear
An enterprise AI visibility dashboard reports a single number for the brand. It gets there by sampling buyer questions, counting how often the brand appears in the answers, and rolling every market into one figure. The rolling is where the information goes. A metro where the business is named in nine of ten answers and a metro where it is named in none both feed that average, weighted by how many questions were run in each.
The arithmetic buries the second metro. Seven strong markets and two silent ones produce an aggregate that reads healthy, and that aggregate is what reaches the executive summary. The silent markets still hold buyers. They ask the same questions and receive shortlists that leave the business out. The brand score does not flag them, because it was computed at the wrong unit.
Brand is a reporting unit. Location is a buying unit. A buyer in Charlotte types a query and an assistant returns names it has retrieved from sources tied to Charlotte. The sources behind a Charlotte answer and a Seattle answer come from different retrieval sets. Local pages dominate each retrieval set, and the national sites that show up in both are the ones already ranking for the category. The recurring names in a metro come from the overlap in that retrieval set, which is why the same few businesses keep appearing while everyone else stays absent (why AI assistants recommend the same businesses).
That overlap also decays. Sources fall out of the retrieved set as new pages publish and older ones lose relevance, so a mention earned in one quarter does not hold at the same rate in the next (AI citation half-life). A brand score absorbs that decay exactly the way it absorbs an absent market. Gains in three metros offset losses in two, the top line holds steady, and the losses never surface as a line item.
Aggregates track a single business over time when nothing else changes. Every market moves at once in an aggregate. Markets move separately. One location gains a mention because a local page got indexed. Another loses one because a competitor published comparison content. Averaged together, the two events cancel and the dashboard reports stability.
Measure at the location. Run the buyer questions for each market and keep the answers separate. Then count how many locations put the business in the shortlist at all. That count is the figure an aggregate cannot produce.
What 16,000 location-level scans found
Birdeye’s July 2026 study ran more than 16,000 location-level ChatGPT scans across 1,500+ multi-location brands in 28 industries. Roughly 20% of locations were invisible. Fewer than 1% returned a fully correct profile across name, address, phone, website, and hours. Hours were the most frequent error. A brand with a strong aggregate score still has individual branches that fail the basic match between what the assistant says and what the location actually is.
SOCi’s 2026 Local Visibility Index shows the same selectivity from another angle. ChatGPT named about 1.2% of eligible locations. Google’s local results included roughly 35.9% of those locations. The gap is about an order of magnitude. One channel surfaces a third of eligible locations. The other surfaces a little over one in a hundred. For a multi-location brand, the average hides the unit that matters: the location.
These numbers line up with the category shortlist behavior already documented for AI assistants. When buyers ask for recommendations, assistants return a small set of names. The same pattern appears per market. A brand with 40 locations does not get 40 chances. It gets 40 separate tests, and most locations fail the first filter. The earlier analysis of how assistants recommend the same businesses across categories (/blog/why-ai-assistants-recommend-the-same-businesses) describes the shortlist mechanism. The location scans show what that mechanism produces when the question includes a city or a neighborhood.
aeod.app ran its own location-level scans one market at a time. The output is a count: how many locations put the business in the shortlist at all. That count separates a brand-level mention from a local presence. A brand gets cited in a national answer while staying absent in the suburb where the buyer is standing.
Why the same question names different branches in different cities comes down to how the assistant resolves location, how it matches the query to a branch profile, and how local signals vary across markets.
Why the same question names different branches in different cities
When a buyer asks an AI assistant for a plumber near a specific city, the assistant resolves location before it builds any shortlist. It reads the city name, neighborhood, postal code, device location, or prior conversation context. That resolved location becomes a filter on the record set. The assistant then retrieves local entity records from Places-style APIs, directories, and review platforms. Each record is bound to an address, service area, category, and hours. The candidate pool is local by construction.
Retrieval ranks those candidates with nearby evidence and freshness. Nearby evidence includes distance to the resolved location, category match, service area overlap, and local review volume. Freshness includes recent reviews, updated hours, recent photos, and active listings. A branch with strong national brand recognition but thin local records drops out. A single-location competitor with dense local signals enters. The same prompt in Denver and Miami produces two different pools because the records pulled are different. Identical wording does not force identical candidates.
The entity and retrieval-layer logic matches what governs national category questions. The assistant resolves entities and matches categories before scoring evidence. Location adds a hard boundary to the record set, which changes which branch profiles are available to score.
A national brand score averages across markets and hides branch-level differences. A brand appears in one city prompt and is absent in the next. The average looks stable. The local reality is uneven. Running prompts with city context and recording branch-level mentions and citations is how the location effect becomes visible.
Building a city-level prompt set
A city-level prompt set starts with the buyer’s category question in their own words. If a buyer asks “Who is the best plumber in Austin?”, the market is Austin. The same question resolves to Denver, Seattle, and every other market by changing only the place name. Wording stays fixed. The only variable is location. That control is what makes the comparison readable. If the wording changes between cities, the result measures the prompt, not the market.
Run two prompt forms for each market. The first names the city explicitly: “Who is the best plumber in Austin?” The second uses “near me” with location set in the assistant or device: “Who is the best plumber near me?” The explicit prompt tests the city as a string. The “near me” prompt tests the assistant’s location resolution. Both belong in the set because assistants handle them through different retrieval paths. A city string pulls a directory page. A location signal pulls a map pack or a local index.
Run each prompt across several assistants. ChatGPT, Gemini, Claude, and Perplexity each assemble answers from different sources and ranking logic. A single run on one assistant returns one sample from a distribution. Repeat each prompt several times per assistant to see whether the same business appears across runs. Repetition separates a stable recommendation from a one-off generation.
Where buyers narrow by constraint, extend the thread. A first prompt identifies the category and city. A follow-up asks “Which of those handle emergency calls?” or “Which of those are open on Sunday?” That follow-up keeps the prior shortlist in context. Restarting the prompt as a new question discards the context and changes the test. The method for that is covered in multi-turn follow-up questions. When the buyer asks a head-to-head question, such as “Plumber A vs Plumber B in Austin,” treat it as a separate prompt family. The verdict pattern is covered in AI comparison prompts.
Run this set as a matrix: prompt by market by assistant by repetition. The output records mention, citation, and position for each cell. A brand-level score collapses that matrix into one number. The city-level set keeps the rows separate.
Once the matrix is populated, each row has to be read on its own terms. A mention in Austin differs from a citation in Denver. A recommendation for a branch in Seattle is a third finding.
Reading per-location results: mention, citation, and the wrong branch
Reading a row starts with four fields per market. The first field is whether the correct branch is named in the answer. The second is whether a sibling branch is named in its place. The third is the rank position of the named business among the shortlist. The fourth is the domains cited for that market. These fields turn a city row into a diagnostic record.
The split works because mention and citation status are reported with a rank position per question. Split the prompt set by city, and each question maps to those fields. A prompt like “who is the best dentist in Austin” returns a named business at a rank with cited domains. The same question with Denver substituted returns a separate record. The matrix row for Austin does not inherit Denver’s citations.
The wrong-branch result is the field a brand score cannot show. A brand score aggregates every city into one number. A national brand with mentions in many markets shows a healthy aggregate while the Seattle row names the Portland branch. The brand is present. The correct branch is absent. That result points at entity resolution. The assistant has a page or profile for the brand, and it has pages for individual locations. When the location entity is ambiguous, the model substitutes a sibling it has seen more often. Content volume on the brand site does not fix that substitution. Clean location entities and consistent branch names in structured data do.
The pattern connects to why AI assistants recommend the same businesses: retrieval favors entities the model can resolve. Follow-up questions add constraints, so a single prompt misses the branch substitution that appears on turn two.
That distinction points to the mechanism behind the substitution: why one branch gets named instead of another.
Why one branch gets named instead of another
A brand-level page that names the company but never resolves to a street address creates a gap the assistant fills from other sources. If the site says “Serving Chicago” while the Google Business Profile lists a Naperville address, the model has two candidate locations. Entity resolution depends on field agreement. Name, address, phone, hours, and category are the join keys. When one field conflicts, the record splits. One node inherits the profile reviews. The other node inherits the website copy. The assistant names the node with more corroboration, and the other branch disappears from the shortlist. This is the mechanism behind many wrong-branch answers.
Hours and phone numbers are common splitters. A directory lists an old phone number. The location page lists a new one. The business profile lists holiday hours that differ from the site. Each mismatch creates a separate entity candidate. The model does not reconcile them by brand name alone. It resolves them by repeated, consistent location identifiers across independent sources. Review volume then acts as a tiebreaker. If review volume sits mostly on one market’s profile, that market wins the entity match. Other markets have no comparable evidence, so they remain absent even when the brand has a physical presence there.
The repair order matters. First, reconcile per-location identifiers. Standardize the business name, street address, phone, hours, and category on the site, the business profile, and the major directories. Fix the old phone number and the mismatched hours before adding new content. Second, add market-level evidence. Local reviews and local coverage give the assistant market-specific corroboration. A Chicago review from a Chicago customer attaches to the Chicago entity. A regional news mention does the same. Third, build genuine per-market pages. A page with a unique address, staff, service area, and local proof is an entity asset. A templated city duplicate with swapped city names is a near-duplicate and adds no resolution signal.
The concentration of citations follows the concentration of corroborating sources. A citation to one location and a missing mention for another is a different finding from a low brand score. Once the per-location records resolve, the next problem is timing.
Set re-measurement cadence per market
Once the per-location records resolve, the cadence follows the retrieval path. A brand score averages different markets into one number. That number changes for reasons that belong to one market. Weekly runs on live-retrieval prompts for priority markets catch source changes within seven days. Live-retrieval prompts pull current web results, so a new review, a closed location, or a local news citation lands in the answer set. Run the same prompt set across each model in your measurement set. Perplexity and Google AI Overviews retrieve live pages; ChatGPT with browsing enabled does the same. A monthly pass covers prompts answered from model knowledge, where weights change on a slower release cycle.
An immediate re-run follows any per-location data change. Update a phone number, address, hours, service area, or review profile and the assistant’s retrieval path shifts. The next run shows whether the cited source changed. Without that immediate re-run, the log mixes the old path with the new one.
Log each run with the market, prompt, model, cited source, and the source’s publication date. That record shows which markets hold their position across cycles and which lose citations as sources rotate. A brand-level score reports one number for all of them. The per-market log reports what each location earns.
Your own answers
See what AI says about your business.
One domain, one assessment. Mention and citation rates, competitors named instead of you, and a prioritized action list.
Get your report · $29 ↗More notes

31 Aug 2026 · 13 min read
Why AI Assistants Recommend the Same Three Businesses
Why AI assistants recommend the same businesses, explained by concentration data, retrieval sources, and entity consistency.

16 Aug 2026 · 14 min read
Multi-Turn AI Visibility: One Follow-Up Erases the List
Multi-turn AI visibility data: one added constraint removes 62% of named brands, and single-prompt tracking hides the drop.