Skip to main content
日本語

How to Measure AI Visibility

Tomohiro Iida · Published July 7, 2026 · Updated July 7, 2026

Noticing that ChatGPT mentioned a company’s name, or that AI Overviews cited it, has become an everyday occurrence as more companies engage with generative-AI search. That raises the question of how to track AI visibility, how visible a company is inside AI, as a number. But reporting a single-shot figure, such as an AI visibility score of 85, from one observation has a pitfall that should not be ignored. This article starts from the non-deterministic nature of generative-AI answers to work out why a single-shot score cannot be trusted, and what an honest way to measure it looks like.

Key takeaways

  • Generative-AI answers are non-deterministic: the same question can produce a different output each time it is run. A single observation is just one sample drawn from a population, and a single-shot score has no backing for reproducibility.
  • The right approach is to observe repeatedly, across multiple conditions, and treat the result as a distribution. Rather than rounding to a single overall score, observe with separate metrics at each stage: mention rate, citation selection rate, citation absorption rate.
  • Attach a range to any proportion, using something like a Wilson confidence interval, and break results down by condition. Because causation for a given tactic is hard to establish, it is more realistic to read the numbers as a trend than to state them as fact.

Generative-AI answers are non-deterministic

ChatGPT, Gemini, Google AI Overviews, and other generative-AI tools usually give a different output each time the same question is asked, run to run. The cast of citations, whether a company gets mentioned, even fine details of the wording, none of it is guaranteed to line up twice. This comes down to factors such as sampling, model updates, and whether web search is used; it is closer to how generative AI is built to behave than to a bug. Fixing temperature and seed through an API can improve reproducibility in some cases, but ordinary search use runs with this variation as the default. In other words, a single observation is just one sample drawn from a population that varies. The whole picture cannot reasonably be inferred from one sample.

Why a single-shot score cannot be trusted

Round data with a sample size of one into a phrase such as an AI visibility score of a given number, and that number is heavily pulled by whatever happened to come up that one time. Ask the same question again tomorrow and there might be no mention at all, or a different competitor pushed to the front. A single-shot score looks precise, but it has no actual backing for reproducibility; how a number looks and how trustworthy it is are two different things. Treating a single-shot value as grounds for a decision risks reading a change into something that is not really there.

Covers the structural reasons behind not being cited by AI search.

Seven Reasons AI Search Does Not Cite You

Observe repeatedly, as a distribution

The right approach is to observe repeatedly, across multiple conditions, and look at the result as a distribution. Conditions here means things such as provider (ChatGPT versus Gemini), model, observation date, variation in how the prompt is phrased, region setting, and whether web search is toggled on. Varying these and observing repeatedly reveals not a single number but a sense of what share of the time, and with what spread, a company shows up. The starting point is treating this as a picture with a range, not a single point.

Covers the practical side of running this kind of repeated observation.

Tracking AI Citations

Break the metric into stages

Rounding AI visibility into one overall score hides what is actually happening. Instead, observe with several metrics broken into stages. The main ones are as follows.

Mention rate
The share of AI answers in which a company or brand appears.
Citation selection rate
The share of the time a company is chosen as a cited source.
Citation absorption rate
The share of the time a citation is actually reflected in the content of the answer itself.
Supported claim rate
The share of claims that are verifiable and backed by a source.
Answer entity consistency
How consistent a company’s information is within a given answer.
Competitor inclusion
How often a competitor is mentioned alongside the company.

Citation selection rate and citation absorption rate, in particular, are not official, industry-standard metrics defined by a search engine or AI vendor. They are operational categories used in this article for diagnostic purposes; treat them with the understanding that definitions and measurement can vary by whoever is running them.

Being chosen as a source and having that content actually reflected in the answer’s body text are two different phenomena. Breaking the metric into stages makes it possible to isolate which layer things are getting stuck at.

Rather than a single-shot score, SIGNAL looks at current standing through repeated observation and staged metrics. The free scan and diagnostic give a reference AI-visibility score and the top three priority issues. It does not guarantee search ranking or inclusion in AI answers.

See What SIGNAL Covers

Attach a confidence interval to any proportion

A proportion from repeated observation is still only an estimate from a finite number of trials. Wilson confidence intervals or bootstrapping can make explicit the range a proportion could actually fall in. If a mention rate came out to 3 out of 10, for instance, the caveat is needed that the true rate could be scattered across a wider range. The average or range from a small number of trials is a useful reference, but it is a weak sample size and cannot ground a firm claim. Showing a range is not an admission of weakness, it is how numbers are handled honestly. A wide interval signals a stage where trials are still being accumulated. One note: a Wilson interval is a guide for a binomial proportion within a single, consistent condition. Mixing provider, model, day, and region together into one interval breaks the interpretation, so the premise is to break results down by condition first.

Causation for a given tactic is hard to establish

Chasing a single-shot perfect score tends to make people read too much into every up and down in a number that naturally varies, and misread the causation behind a tactic. Just because an AI’s answer changed does not mean it can be attributed to a company’s own initiative. Run-to-run variation, confounding from other factors, and unpublicized changes in model behavior all pile up. Even where a change is observed, attributing the cause to a single tactic calls for caution. It is more realistic to read this as a trend than to state it as causation.

Building this thinking into a tool

Netsujo’s own tool, SIGNAL, is designed around the thinking described in this article. Three things anchor it: repeated observation instead of a single-shot score, staged metrics instead of one overall figure, and a confidence interval showing a range instead of a single point. That said, SIGNAL does not guarantee search ranking or inclusion in AI answers. It is a tool for continuing to observe and tracking how the distribution changes as a reference value. Measurement is not magic, it is steady repetition and honest caveats, accumulated.

Frequently asked questions

Why can AI visibility not be measured with one observation?
Generative-AI answers are non-deterministic: even the same question can change, run to run, in whether a source is cited or a company is mentioned. A single observation is just one sample drawn from a population, and with too small a sample size, a single-shot value has no guaranteed reproducibility.
About how many times should this be observed?
There is no one-size-fits-all answer. A useful gauge is whether the confidence interval on a proportion has narrowed enough to support a decision. While trials are still few, the interval stays wide, which signals a stage where firm claims should be avoided. Keep accumulating trials while varying conditions such as provider, model, day, and phrasing.
Is it a problem to manage this with a single overall score?
Rounding into an overall score hides which layer a change actually happened at, mention rate, citation selection rate, or citation absorption rate. Observing with metrics broken into stages is more useful for isolating an issue and designing a response.
Does using SIGNAL guarantee citation by AI?
No. SIGNAL is a measurement tool for grasping the distribution of visibility through repeated observation, staged metrics, and confidence intervals; it does not guarantee search ranking or inclusion in AI answers. The free scan and diagnostic are free, and current pricing for the paid tiers beyond that is on the pricing page.

The free scan and diagnostic give a reference AI-visibility score and the top three priority issues; from there, repeated observation and staged metrics identify what to improve next. It does not guarantee search ranking or inclusion in AI answers.

Check Your Current AI Visibility, Free