Skip to content

How to Check If AI Mentions Your Brand Reliably in Practice

Learn how to check if AI mentions your brand with real buyer questions, repeated runs, captured sources and a defensible before-and-after method.

Five-stage text-free workflow on a deep axis, with question cards splitting into repeated runs, captured answer and source cards, a comparison grid, a recheck loop, and an olive folder at the final station.
By AI Priority Map Editorial

A single prompt is the tempting method because it is quick, vivid and almost impossible to compare. Learning how to check if AI mentions your brand requires a different approach: predefine real buyer questions, repeat them under recorded conditions, preserve the answers and score the same fields every time. The protocol turns an anecdote into bounded evidence.

Quick Answer. A reliable method for how to check if AI mentions your brand uses real buyer questions, repeated runs and a fixed capture sheet. Record exact answers, mentions, recommendation context, competitors and sources. Hold wording and buyer context constant, then make one evidence change and rerun the same protocol for comparison.

Last updated: 27 August 2026

One prompt gives you a snapshot

One query can answer a narrow question: what did this assistant say in this run? It cannot show whether the result recurs across buyers, prompts, assistants or repetitions. That limitation applies whether the brand appeared or did not appear; a pleasing answer is no more stable merely because it is pleasing.

The distinction between snapshot and measurement is operational. A snapshot may be worth saving, especially when it describes the business wrongly or names an unexpected competitor. A measurement defines the question set, conditions, repetitions, capture fields and scoring rules in advance. Someone else should be able to read the protocol and understand what was sampled.

For the decision that follows the measurement, the guide to whether ChatGPT recommends a business explains how to separate a mention, recommendation, absence and unusable answer. The method here concentrates on producing those observations consistently.

An owner can complete a small manual baseline in a spreadsheet. The hard part is not the interface but resisting changes halfway through: adding easy questions after poor results, removing unusable answers, switching personas between runs, or counting a name mention as a recommendation because the distinction was never written down.

Define questions around buyer decisions

The question set should represent demand, not brand recognition. “What is Acme Studio?” tests whether the assistant can say something about a supplied name. “Which accessibility consultancy suits a regional retailer redesigning its checkout?” tests whether the business appears in a buyer's consideration set for a specific job.

A useful set covers several stages:

  • Category discovery: options for a broad but real need.
  • Fit: providers suitable for a company size, industry or constraint.
  • Location: services available in a named market or service area.
  • Specialism: a narrow capability that differentiates the business.
  • Alternative: options compared with a known incumbent or approach.
  • Risk or constraint: providers able to work under a practical limitation.

Each question needs an identifier, exact wording and reason for inclusion. Ten vague variants of the same vanity prompt do not create ten independent buyer needs. Conversely, one broad question cannot represent every stage of a purchase.

Twenty questions is the internal full-audit product protocol, not a universal research law, and a manual baseline may use fewer if the business genuinely has a narrow offer. The defensible rule is that each included question maps to a plausible buyer decision, and each excluded question has not been quietly removed because of its answer.

Questions should also avoid leading language. “Why is Acme the best?” presupposes the conclusion. “Which firms can help a ten-person clinic move booking data?” gives the assistant a buyer, a task and a constraint while leaving the recommendation open.

Separate persona, location and language

Buyer context can change the recommendation set, so it belongs in the design. A 2026 preprint audited 2,000 runs across ten personas, eight prompts and three model configurations with repeated samples. Persona prefixes reduced set similarity by 0.12–0.20 Jaccard points against the study's same-persona baseline 2. Those values belong to that audit, not every assistant.

The practical implication is grouping. If the business serves both solo operators and enterprise procurement teams, the same category question can be run as two labelled persona conditions. Results from those groups should not be pooled into one score and called “brand visibility”; the buyer contexts are doing different work.

Location should be equally explicit because “near me” depends on context the recorder may not be able to inspect. A named city, region or country is easier to preserve; if the business is remote, the question can say so, while a protocol for which geography does not matter should state that rather than leaving it implicit.

Language variants are separate conditions, not translations casually swapped between repeats. A 2025 preprint comparing AI-search services reported differences in cross-language stability and sensitivity to query phrasing within its controlled experiments 1. A multilingual business can therefore build a matched set by buyer intent, while still analysing each language independently.

The rule is simple: one row group changes one planned variable. Persona comparison changes persona, language comparison changes language, and a held-constant repetition changes neither. When several variables move at once, the output may be interesting, but the cause of the difference remains opaque.

Freeze the run protocol before asking

A protocol prevents the researcher from negotiating with the results. It specifies what happens before the first query, how an unusable answer is handled and when the run is complete. A short version can fit at the top of the capture sheet.

The five-stage sequence is shown below.

In prose, the exact sequence is 1 Define questions, 2 Repeat fixed runs, 3 Capture answers and sources, 4 Compare the pattern, then 5 Change and recheck. It ends with Same protocol, comparable evidence. The order matters: a change made before the baseline is complete removes the baseline, while comparison before capture invites memory to replace evidence.

The protocol header should name:

  1. the assistants and product surfaces being tested;
  2. the date range and language;
  3. the buyer persona and location treatment;
  4. the exact questions and number of repetitions;
  5. the account or session handling where relevant;
  6. the definition of usable, mentioned, recommended and accurate;
  7. the treatment of refusals, cut-offs and ambiguous identities.

No universal repetition count creates certainty; more runs reduce dependence on one generation, but they still describe a bounded sample. The supplied full protocol uses five runs for each researched question across three assistants, whereas the free check asks one question on one assistant twice and describes itself as a snapshot. Those are product facts and scope choices, not statistical guarantees.

A pilot is sensible before full collection. One question can reveal that the sheet lacks a “wrong namesake” state or that answers are being truncated by the capture method. The pilot repairs the protocol, then restarts cleanly; its rows should not be mixed into the final sample if the coding rules changed.

Capture answers, sources and competitors

Verbatim evidence is the centre of the method. Screenshots can preserve appearance, but structured text is easier to compare and search. The complete answer or an explicitly marked retained excerpt should sit beside the coding fields, never be replaced by a summary written after the run.

Capture fieldAllowed values or formatPurpose
QuestionID plus exact textKeeps variants separate
ConditionsAssistant, language, persona, location, timestampDefines the run
UsabilityUsable, refusal, blank, cut-off, errorStops missing output becoming brand absence
Brand outcomeRecommended, mentioned only, absentSeparates endorsement from appearance
AccuracyCorrect, partial, wrong, ambiguousReveals harmful visibility
CompetitorsNames plus recommendation contextBuilds the comparison set
SourcesURLs, titles and relevant passagesShows evidence paths
NotesMinimal coding explanationPreserves edge-case reasoning

Source capture requires more than copying links: a 2023 human audit of Bing Chat, NeevaAI, Perplexity and YouChat found that 51.5% of generated sentences were fully supported by citations, while 74.5% of citations supported the sentence associated with them 3. The figures describe those four systems in the 2023 study and justify checking coverage and support separately, not distrusting every current citation by default.

For each citation relevant to the brand or competitor, note the claim it appears to support, whether the page covers that claim, and whether the passage supports it. A correct URL can still be irrelevant. A relevant page can still be stale. A source can also describe another business with a similar name.

Competitors are recorded as evidence, not a fixed rank. Preserve any order the answer explicitly gives, but do not invent one when the response groups options by use case. A competitor count without context loses whether the brand was an incumbent, specialist, budget option or merely an example.

Score mention, accuracy, stability and evidence

The analysis should retain several dimensions instead of collapsing everything into one impressive number. Mention rate answers how often the correctly identified brand appeared in usable runs. Recommendation rate answers how often it was presented as suitable. Accuracy rate answers whether the statements about it were correct. Stability describes recurrence within a defined question and condition.

MeasureCalculation inside the sampleImportant boundary
Mention rateRuns with a correct brand mention / usable runsA mention is not necessarily favourable
Recommendation rateRuns recommending the brand / usable runsContext must fit the buyer need
Accuracy rateCorrect brand descriptions / runs naming the brandAbsences do not become inaccuracies
Question coverageQuestions with at least one mention / questions testedDoes not show repetition stability alone
Competitor recurrenceRuns naming each competitor / usable runsNot a universal market share
Source recurrenceRuns citing each domain / runs with citationsDoes not prove source quality

An unusable answer stays visible as an unusable answer. Whether it belongs in the denominator depends on the named measure, and the choice should be stated. The cleanest approach uses usable runs for brand mention rates while reporting the unusable count separately. That prevents system failures from being treated as negative brand judgements.

Variation can be described without pretending to know its cause: a brand named in every repeat for one question is stable within that cell, while a brand alternating between present and absent is unstable within that cell. Neither statement travels automatically to another persona, assistant or week.

The seven-day audit of 48 authentic queries across four topics found query- and topic-sensitive sentiment plus commercial and geographic source bias in ChatGPT, Bing Chat and Perplexity 4. That scoped evidence supports retaining question, topic and geography in analysis. It does not supply a benchmark against which a business should score itself.

A concise report can show the matrix first, then quote representative evidence. A long stream of screenshots forces the reader to perform the analysis again. The report should always preserve a route back to the underlying captured answer.

Diagnose the gap without guessing causation

A measurement tells you where to look, not automatically why the output occurred. An absence for a specialist question may coincide with no clear service page. A wrong location may match stale profiles. Competitor citations may reveal better independent corroboration. These are evidence-backed hypotheses, not proof that one factor caused a model response.

The plain GEO guide describes the supported visibility levers: clear passages, relevant scope, verifiable claims and earned evidence. The competitor recommendation explainer adds prominence, source coverage and recommendation-set variation. Together they provide a diagnostic vocabulary without promising inclusion.

Four questions keep diagnosis grounded:

  • Does an owned page directly answer the buyer's need?
  • Are the business name, location and service facts consistent?
  • Is there independent evidence for the claimed expertise?
  • Do cited or likely sources remain current and actually supportive?

Sometimes the answer is “the sample is too thin”, which is a valid diagnosis. Increasing repetitions or adding a missing buyer question is better than rewriting content to explain noise. Sometimes the answer is “the assistant described another entity”, in which case identity disambiguation comes before visibility work.

Make one bounded change and rerun

A before-and-after comparison works best when the intervention is narrow. Rewriting the entire site, changing listings, commissioning coverage and modifying the prompt set at once may improve the business's evidence, but it prevents a useful interpretation of the recheck.

The change should match the measured gap. A missing specialist page earns a focused page with scope, exclusions, examples and evidence. Inconsistent addresses earn corrections at authoritative profiles and owned pages. Weak corroboration earns genuine third-party evidence, not copied claims. Unsupported assertions earn removal or primary support.

The record needs a change log: what changed, where, when and which finding justified it. After the new evidence is available, the original protocol is rerun without altering questions, personas or scoring rules. New exploratory questions can be added as a separate wave, not smuggled into the comparison.

A changed output is an observation after an intervention, not causal proof, because models, retrieval indexes and source availability may have changed too. The report can say “mentions increased within the rerun” or “the old location no longer appeared”; it should not say the edit guaranteed or permanently secured a recommendation.

If the result does not move, the intervention may have been irrelevant, unavailable to the tested system, insufficiently corroborated or simply overwhelmed by output variation. The next decision comes from the evidence, not from repeating the same edit at greater volume.

Keep manual checks and measured audits distinct

A manual check is valuable for learning the method, validating question wording and collecting a small baseline. It becomes misleading when its limitations disappear from the report. One question on one assistant twice is still two observations of one condition, even when both answers agree.

A fuller audit expands three axes: buyer-question coverage, assistant coverage and repetition. Its defensible advantage is broader coverage across those dimensions, not discovery of an eternal truth or any guaranteed analysis beyond the sampled answers.

The report should name its boundary in the headline or methods box. “April sample: twelve questions, two assistants, four runs” is informative. “AI visibility score” without the underlying protocol invites a reader to assume universality the data does not have.

Manual observation and product measurement can therefore coexist. The first helps a team understand its questions and edge cases. The second samples more buyer questions, assistants and repetitions. Neither should be marketed as a durable rank, and neither makes optimisation outcomes certain.

Choose the fuller measurement when the decision is costly

A broader protocol is warranted when a decision justifies sampling more than one question on one assistant twice. Its defensible advantage is broader buyer-question, assistant and repetition coverage; the scope does not guarantee any particular analysis or conclusion.

The full protocol uses twenty researched buyer questions, three assistants and five runs per question. The free check asks one question on one assistant twice. The difference is buyer-question, assistant and repetition coverage, not a claim that uncertainty disappears.

When that scope fits the decision, use the 20-question, three-assistant, five-run AI Answer Audit measurement. That protocol defines a broader sample while preserving the observation boundary.

Frequently Asked Questions

How do I check if AI mentions my brand?
Start with a bounded set of questions that real buyers ask, then run each question repeatedly under recorded conditions. Save the complete answers, mark whether the brand was recommended, merely mentioned, absent or described wrongly, and capture competitors and sources. The result is a sample defined by the protocol, not a permanent AI ranking.
How many times should I repeat each AI question?
There is no universal repetition count that guarantees certainty. The full product measurement uses five runs for each of twenty researched questions across three assistants; the free check uses two runs of one question on one assistant and calls itself a snapshot. Whatever count you choose, define it before starting and apply it consistently.
Should I use exactly the same prompt on every assistant?
Use the same buyer intent and preserve a master wording when comparing assistants, but record any syntax change required by a product surface. A fair comparison holds persona, location, language and constraints constant. If the research question is phrasing sensitivity, create separate labelled variants instead of quietly changing wording between runs.
What counts as an AI brand mention?
A mention means the correctly identified brand name appears in a usable answer. A recommendation is a narrower event in which the answer presents that brand as suitable for the buyer's need. Code those states separately, and add an accuracy field, because a recommendation that assigns the wrong service or location is visible but misleading.
Do AI citations prove a brand claim is accurate?
No. A citation can fail to cover a claim, and a cited page may not support the associated sentence. A 2023 human audit of four named generative search systems found substantial gaps in both citation coverage and support. Inspect each relevant source directly, and record unsupported, stale or mismatched evidence rather than trusting fluent presentation.
How often should I rerun an AI brand visibility check?
Rerun after a defined evidence change or on a scheduled review cycle that suits the business, using the same question set and protocol. Frequent unstructured checking creates noise and encourages reaction to single outputs. A stable baseline, a recorded intervention and a comparable recheck provide more information than daily prompts with changing wording and conditions.

Sources

  1. 1.Generative Engine Optimization: How to Dominate AI SearcharXiv · 2025
  2. 2.Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider AuditarXiv · 2026
  3. 3.Evaluating Verifiability in Generative Search EnginesarXiv · 2023
  4. 4.Generative AI Search Engines as Arbiters of Public Knowledge: An Audit of Bias and AuthorityarXiv · 2024

Want this run on your business?

AI Foundation Audit — a structured assessment of your AI footprint: integration risks, governance gaps, ROI opportunities. Delivered as a comprehensive report you can act on.

Start your audit

You receive your AI Opportunity Report and Implementation Brief — tailored to your business and delivered immediately.