The Disambiguation Threshold | SearchShopAI
Back to Blog Methodology

The Disambiguation Threshold: measuring whether AI actually knows which brand you are

August 25, 20269 min read

When two brands share a name, "the AI got it right" is not a coin flip and it is not one number. It's a metric that depends on how much of the entity context the question already carried. This is the metric we use, and how we score it — anchors, denominators, per-mention verdicts, published.

The short version

  • Every branded prompt is classified into one of four anchoring strata — the amount of entity context the prompt itself hands the model.
  • The Disambiguation Threshold is the least-context stratum at which the identity ladder resolves answers to your brand at 90 percent or better. It is a stratum label, not a rate.
  • Entity work — Wikidata, schema, an entity home, naming consistency — is judged by whether the threshold moves down one stratum at a time: product first, category next, then bare-name.
  • We publish the anchors, denominators and per-mention verdicts openly, per edition. The metric moving is the deliverable; the number alone is not.

The 94 percent problem

"Accurate when named" — the model produced the brand's name in the answer — reads like an entity-accuracy metric. It isn't one. It's a name-presence test in an entity-accuracy costume.

Here is how the two come apart. A small brand shares its name with a large, unrelated one and with a homonym in an adjacent category. The bare-name prompt "is X actually good or just hype" gets a coherent answer, mentions the string that is the brand's name — and the reader would need to already know the brand to notice the answer is about the wrong one. Name-presence: passed. Identity-verified: failed. Averaged together they read as one high number and hide the real signal.

The fix isn't a better model. It's a different denominator. Score answers by how much entity context the prompt supplied, then look at the identity resolution rate inside each context bucket. The averaged number stops being a metric and becomes a mix.

Averaged together, they read as one high number and hide the real signal.

The four anchoring strata

Every branded prompt is classified by how much entity context the prompt itself carries before the model even starts answering. Four buckets, from least to most context:

The four anchoring strata
StratumWhat the prompt carriesExample
Bare Name alone.The hardest bucket. Nothing scopes the entity — the model resolves on prior alone. "Is [Brand] actually good or just hype"
Category Name + category context.Implicit anchors count too — "dermatologists" scopes to skincare. "Do dermatologists recommend [Brand]"
Product Name + one of the brand's own product names.Product names come from the brand's actual catalog, not a hand list. "Is the [Brand] [Product Name] good"
Domain Name + the brand's domain.The easiest bucket. The URL is a coordinate; the model doesn't have to resolve anything. "Is [brand-domain].com legit"

Classifying prompts this way makes an averaged number legible. A bare-name miss and a domain-anchored miss are the same one-percent-of-answers loss in the top-line rate, but they mean completely different things about entity work.

The threshold, defined

The Disambiguation Threshold is the least-context stratum at which the identity ladder resolves answers to your brand at 90 percent or better. Threshold = the label of that stratum. It is not a percentage; it is a name — bare, category, product, or domain.

The 90 percent boundary is deliberately high because the threshold is a client-facing reliability claim: give the AI at least this much context and it reliably finds you. Setting it lower would let a coin-flip stratum masquerade as the threshold.

An illustrative measurement, from one small brand's most recent 210 branded answers:

One brand, one nightly, 210 branded answers

45%
bare
57% of prompts (119 / 210)
64%
category
n = 14 (wide CI)
99%
product
threshold today
100%
domain
by construction

Identity ladder resolution rate per stratum, one anchored brand, most recent nightly run. Numbers illustrative of the methodology, not a claim about any specific brand.

Threshold: product. The bare-name stratum sits at 45 percent, but 57 percent of the current prompt portfolio is bare-name questions. That mix is why an ungrouped pooled rate reads as roughly two-thirds — an average over a mix that has no single meaning. In the bare stratum specifically, the failure is mostly honesty: unverified answers plus explicit clarification requests dominate; the model saying "I am not sure which X you mean" is a different failure mode from the model confidently answering about the wrong X.

What the metric is for

Entity work — Wikidata, on-site schema, an entity home page, name consistency across the platforms AI reads — is judged by whether it moves the threshold down one stratum at a time. Category next, then bare. That turns a shared-name embarrassment into a progress metric with a legible sequence.

Two properties of the metric matter more than the numbers themselves:

  • The bare-name stratum stays low indefinitely for generic names, and that is not a defect to fix this sprint. A popularity prior alone scores roughly two-thirds correct on public entity-linking benchmarks: absent context, resolution defaults to the famous namesake about two thirds of the time. Small brands with common-word names inherit this headwind. The threshold makes the headwind visible instead of averaging it away, and it colour-codes movement between editions rather than banding the number.
  • "No confident entity" is a first-class output. Academic entity linking treats it that way — either by thresholding the top candidate or by ranking a "no match" option alongside real entities. The shrinking share of unverified answers is a health metric on its own. It never gets folded into verified, never gets folded into wrong.

The remediation stack, and its success metric

The metric would be nothing without the levers it scores. Four, in the order they compound:

  • Wikidata item. The load-bearing statement is the official website property — that alone kills the lookalike-domain confusion. Notability is cleared with a handful of independent editorial references, not a Wikipedia-grade bar. This is what AI systems that rely on Google Knowledge Graph or Wikidata-adjacent embeddings actually resolve against.
  • schema.org disambiguatingDescription. A property purpose-built for entities that share a name, sitting inside your Organization schema. One line that says which entity this is; readable by every schema consumer.
  • Entity home. One canonical URL — usually /pages/about or /about — that acts as the entity's reconciliation point, corroborated by every off-site source. Called the "entity home" doctrine in the Kalicube tradition; the same idea as a Wikipedia article's canonical URL, but on your own domain.
  • Naming consistency. Brand name reads exactly the same on LinkedIn, Ulta, Amazon listings, the merchant feed, wherever an AI reads about you. Inconsistent names erode entity confidence even when everything else is correct.

The success metric for that stack is the threshold moving down a stratum. Not "we added schema." Not "we filed with Wikidata." The threshold moved from product to category.

The reporting rules

The metric is honest only if the reporting rules are honest. Three, per edition:

  • Never present name-presence and identity-verified as one measure. They are different tests on different denominators. Error and empty responses carry no identity signal and are excluded from the identity denominator, while name-presence counts them as misses — so identity-verified can legitimately exceed the name-presence rate on the same run. Any surface that shows both carries the denominator note.
  • Colour encodes movement, not band. A generic name's bare stratum will sit low indefinitely; red-banding it every edition communicates nothing. Colour marks the threshold moving down or up a stratum, and per-stratum rates improving or regressing between editions.
  • Thresholds are published per edition. Each report states the threshold stratum, per-stratum n and verified rates, and the prompt mix that produced them. Any prompt-portfolio change carries a methodology note in the same edition.

Why this goes public

Buyers of AI-visibility work have started listing "entity disambiguation accuracy" as a criterion, and none of the vendors publish how they measure it. The score has become the deliverable, without the anchors and denominators that let a buyer check the score.

So publishing the anchors, denominators, and per-mention verdicts is the sales lever. Anyone can now audit our threshold claim: replay a public prompt set through the identity ladder, classify the strata, average per bucket, and reach the same number. If the audit disagrees, the disagreement is legible — a specific stratum, a specific denominator, a specific prompt.

That is the metric we run against internally, and the one we publish externally, and they are the same metric. If you want us to score how well AI resolves your brand's identity today — and give you a stratum-by-stratum plan for moving the threshold down — start with a free audit. The threshold and the per-stratum breakdown are in every report we ship.

Get Your Free Audit

Share this note

See how AI resolves your brand's identity today.

Get Your Free Brand Audit