How to Measure AI Visibility
Tomohiro Iida · Published July 7, 2026 · Updated August 6, 2026
Asking an AI once and judging "AI visibility" purely on whether a company appeared has limits. A generative AI answer is not necessarily the same twice for the same question. The wording of the question, whether search was used, model updates, and the point in time all change which companies are recommended and which URLs are cited. AI visibility is therefore read not as a single score but as the distribution of repeated observations under fixed conditions.
This article is for people at B2B companies who want to manage AI visibility as a number. It works through what to call visibility, how to fix the observation conditions, how to design questions by stage, how to decide the number of repetitions, how to attach a confidence interval, and the limits of a composite score. What it sets out is a discipline for observation, not a method that guarantees inclusion or citation in AI answers.
Key takeaways
- Generative AI answers are non-deterministic: the same question produces different output from one run to the next. A single observation is one sample drawn from a population, and a single score has nothing behind it in terms of reproducibility.
- Record and fix the observation conditions — question text, engine, model, whether search was used, run date, number of runs — and observe repeatedly across several conditions. Rather than rounding to one overall score, observe with metrics separated by stage: mention rate, citation selection rate, citation absorption rate.
- Attach a Wilson confidence interval to an observed rate, and break results down by condition. Because causation for a measure cannot be claimed easily, reading the numbers as a trend rather than stating them as fact is the realistic approach.
AI answers are non-deterministic
ChatGPT, Gemini, Google AI Overviews, and other generative AI can produce different output each time the same question is asked. The set of sources cited, whether a company is mentioned, even fine details of the wording, are not guaranteed to line up twice. This comes down to factors such as sampling, model updates, and whether web search is used; it is closer to how generative AI is built to behave than to a bug. Fixing temperature and seed through an API can improve reproducibility in some cases, but ordinary search use runs with this variation as the premise. In other words, a single observation is one sample drawn from a population that varies. Describing the whole picture from one sample is statistically unreasonable.
Why one score cannot support a decision
Round data with a sample size of one into "AI visibility: N points," and that number is heavily pulled by whatever happened to come up that one time. Ask the same question tomorrow and the company may not be mentioned at all, or a different competitor may be pushed to the front. A single score looks precise but has nothing behind it in terms of reproducibility; how a number looks and how trustworthy it is are two different things. Treating a single value as grounds for a decision risks reading a change into something that is not there. The structural reasons behind not being cited by AI are covered separately.
Fix the observation conditions and repeat
Measurement starts with recording and fixing the conditions. When comparing, keep at least the following aligned and on record.
- The question text and its purpose
- Put into words what the question is trying to find out, and keep it.
- Branded or non-branded
- Whether the question contains the company name, or searches by category.
- Target region and target industry
- The candidate set changes, so always record these separately.
- The AI service used
- Which service, and through which route, the observation was made.
- Model or mode
- Output differs within the same service, so distinguish them.
- Whether search was enabled
- Whether web search was used changes the sources cited.
- Run date and number of runs
- Keep when, and how many times, as the denominator.
For example, "companies in Kyoto to consult about AI development" and "companies to consult about deploying RAG in manufacturing" have different candidate sets. Averaging questions together makes it impossible to see which market a company is being found in. Observing repeatedly while varying conditions reveals not a single number but the rate at which, and the spread with which, a company appears. The practical side of tracking is covered separately.
Design questions by stage
Split the questions you observe along the customer’s decision stages. The pages and the answer content each stage needs are different.
| Decision stage | Example questions |
|---|---|
| Problem awareness | What is stopping inquiries from increasing / How can internal documents be searched with generative AI |
| Solution | What is the difference between RAG and an AI agent / How should a B2B site be improved |
| Comparison | Companies in Kyoto that can build AI agents / Companies that can support a Web3 PoC from concept through implementation |
| Purchase | The cost of an AI agent PoC / Companies that can take on everything from SEO analysis to code implementation |
Not being found at the problem-awareness stage and being dropped from the candidate set at the comparison stage call for fixes in different places. Mixing the stages into a single number erases that difference.
Decide what you call visibility
The phrase "AI visibility" covers several different states. Mixing them into one score makes it impossible to judge what to improve.
- Mention
- The company or service name appears in the body of the answer.
- Recommendation
- It is picked as a candidate, with an explanation of why it fits the user’s conditions.
- Citation
- The company’s URL is shown as a source.
- Reflection of content
- Features, pricing, or track record written on the company’s pages are used in the answer.
- Accuracy
- The company name, services, pricing, and target customers are described correctly.
- Action
- A user who saw the answer moves to the site and gets in touch.
The metrics below are those states turned into something observable. In Netsujo SIGNAL, the analysis side that aggregates repeated observations treats these six as the unit of observation. What a scan alone judges mechanically is the part corresponding to mention and citation of the official site.
- Mention rate
- The observed rate at which the company or brand appears in an AI answer.
- Citation selection rate
- The observed rate at which it is chosen as a source.
- Citation absorption rate
- The observed rate at which a citation is actually reflected in the body of the answer.
- Supported claim rate
- The observed rate of claims that are verifiable and backed by a source.
- Answer entity consistency
- How consistent the company’s information is within an answer.
- Competitor inclusion
- The degree to which competitors are mentioned at the same time.
Citation selection rate and citation absorption rate are not official or industry-standard metrics defined by a search engine or an AI vendor. They are operational categories used in this article for diagnostic purposes; treat them on the premise that definitions and measurement can differ between whoever is running them.
Being chosen as a source and having that content reflected in the body of the answer are two different phenomena. Separating the stages isolates the layer where things are getting stuck.
Decide the number of repetitions
Running a question once tells you nothing about how often it appears. Observe the same conditions 20 times and see 6 mentions, and the observed rate is 30 percent — but that does not establish 30 percent as the true level. The smaller the sample, the greater the uncertainty.
There is no single correct answer, so in practice we decide the number of runs on one of the following. These are operating guidelines, not universal standards.
- A budget ceiling
- Decide the monthly number of runs first, from the cost of a single observation.
- The size of the difference you want to detect
- Work backwards from how large a month-on-month difference you want to distinguish.
- The width of the confidence interval
- Accumulate trials until the interval is narrow enough to support a decision.
- The importance of the question
- Increase repetitions for important comparison questions; start exploratory questions with fewer runs.
Rather than a single score, we look at the current state through repeated observation and staged metrics. The free diagnosis checks public AI-search visibility and returns likely blocker hypotheses and the first repair instruction. It does not guarantee search rankings or inclusion in AI answers.
Check your current AI visibility, freeAttach a confidence interval to an observed rate
Show only the observed rate, and the gap between 30 percent and 32 percent looks like a meaningful change. With a small sample it may be within the margin of error. An observed rate from repeated observation is a value estimated from a finite number of trials, so we attach a binomial confidence interval as a reference range showing the uncertainty in the estimate. In SIGNAL, the analysis side that aggregates repeated observations uses a Wilson confidence interval (95 percent level by default), because with few trials it does not run outside 0 to 1, or become excessively narrow, the way a normal approximation can.
A confidence interval does not express "the probability that the true value falls within this range." It is a guide to how often the interval would contain the true value if the same procedure were repeated. It assumes observations that are close to independent.
For reporting to management, a form like the following is easier to work with.
- Mentions: 6 of 20 runs
- Observed rate: 30 percent
- Estimated range: shown alongside as a reference value
- Previous month: 5 of 20 runs
- Judgment: cannot be called a clear improvement
The aim is to give information a decision can be made from, rather than to make numbers look precise. A wide interval reads as a stage where trials still need to be accumulated. A Wilson interval is a guide for a binomial proportion within one consistent condition; mixing engine, model, run date, and region into a single interval breaks the interpretation, so the premise is to break results down by condition.
Constraints when using a composite score
If several metrics are combined into one figure, publish the following at the same time. Publishing only the score, without these, leaves the reader unable to check what it means.
- The component metrics and their weights
- The target questions and target engines
- The number of runs
- How missing data is handled
- The update date
- What the score does not represent
A composite score is usable as an aid for reviewing month-on-month change. For choosing what to improve, though, we use the staged metrics. A high citation rate with incorrect pricing calls for an accuracy fix. A high citation rate with no inquiries calls for a review of the path to an inquiry and the service description.
Do not state a change as the result of a measure
If mention rate rises the month after an article is published, that article is not necessarily the cause. The following factors move at the same time.
- Model updates
- Changes in the search index
- Competitors’ published information
- Social media and events
- Differences in the wording of the question
- Seasonality
- Service news
When evaluating a measure, fix the periods before and after publication, and record the pages changed, the question sets, and other activity. Where possible, also compare against question sets that were not changed. How to approach generative AI search is covered in our articles on generative AI search being SEO and on GEO. Reading the numbers as a trend rather than as causation is the realistic stance.
A minimum monthly report
A monthly AI visibility report needs only the following.
- The period covered
- The question sets and their purpose
- Engines and number of runs
- Mention rate, citation rate, accuracy rate
- The content of incorrect answers
- The difference from the previous month
- A confidence interval, or the uncertainty
- The changes published
- The next pages to fix
- Items where judgment is held
Knowing "which question, which stage changed, and what to fix next" leads to implementation more reliably than a report saying the score went up.
How Netsujo SIGNAL approaches this
Netsujo SIGNAL does not settle AI visibility with a single judgment. We split questions along the customer’s decision stages, observe several times, and check mention, citation, and accuracy separately. The conditions under which a scan of an individual site is run are published on the methodology page as follows.
| Item | Conditions for an individual scan |
|---|---|
| Number of questions | Generated per target across up to four lines — branded, service, problem, and latest — according to the information supplied |
| Repetitions | Each question is run twice by default, and aggregated as an observed rate |
| Engine | OpenAI Web Search. Perplexity Sonar is implemented but not enabled at present |
| Scoring | Out of 100: AI presence 25, information accuracy 30, official-site citation rate 20, share within AI answers 15, information freshness 10 |
| Judgment | Mention and official-site citation are judged mechanically. Accuracy and freshness are provisionally judged automatically, then confirmed or rejected by a person |
Attaching a confidence interval is a step on the analysis side, which aggregates repeated observations on an ongoing basis. A scan on its own presents the observed rate and a reference score; it does not display the interval itself. The engine, model, whether search was used, the region, and the run date and time are recorded in the individual report.
When the question set, judgment criteria, number of runs, or engine change, we update the version and keep the history. The detailed conditions are published on the SIGNAL Lab methodology page.
The figures shown in the free diagnosis are reference values under limited observation conditions. They do not represent an absolute level across the market, and they do not guarantee rankings or inclusion in AI answers. What matters is not a high score but identifying the stage where things stop and converting it into an improvement that can be implemented. The record of our own practice is collected in the practice log.
Frequently asked questions
- Why can AI visibility not be measured with one observation?
- Generative AI answers are non-deterministic: even for the same question, the sources cited and whether the company is mentioned change from run to run. A single observation is one sample drawn from a population, and with too small a sample the reproducibility of a single value cannot be guaranteed.
- About how many times should we observe?
- There is no single correct answer. The gauge is whether the confidence interval on the observed rate has narrowed enough to support a decision. While trials are few the interval stays wide, which reads as a stage where firm claims should be avoided. The number also changes with the budget ceiling, how large a month-on-month difference you want to detect, and the importance of the question. Accumulate observations while varying conditions such as provider, model, day, and phrasing.
- Is it wrong to manage this with a single composite score?
- Rounding into a composite score hides which layer a change happened at — mention rate, citation selection rate, or citation absorption rate. It is usable as an aid for reviewing month-on-month change, but staged metrics are what you use to choose improvements. If you do show a composite score, publish the component metrics and weights, the target questions, the target engines, the number of runs, how missing data is handled, the update date, and what the score does not represent.
- If the AI visibility score rises, can we call it the result of our work?
- Not conclusively. Model updates, changes in the search index, competitors’ published information, differences in the wording of the question, and seasonality all move at the same time as your own work. Fixing the periods before and after publication, recording the pages and question sets you changed, and comparing against question sets you did not change reduces misreadings.
- Does using SIGNAL mean AI will definitely cite us?
- No. SIGNAL is a measurement tool for grasping the distribution of visibility through repeated observation, staged metrics, and confidence intervals; it does not guarantee rankings or inclusion in AI answers. The free diagnosis is free, and current pricing for the paid tiers beyond it is on the pricing page.
The free diagnosis checks how you appear in AI search and returns the first repair instruction; from there, repeated observation and staged metrics identify the priority improvements. It does not guarantee rankings or inclusion in AI answers.
Check your current AI visibility, free