AI Visibility Research Methodology
This page publishes the procedure we actually run for AI-search-visibility observation in SIGNAL Lab and the free web sales-infrastructure diagnosis. It describes only the method currently in use — no target values or unimplemented techniques. When we change the method, we update the version and keep the history.
01 / Definition — AIO and AI visibility
In this Lab, AIO means AI Optimization (AI search optimization): work that lets AI search — ChatGPT, Gemini, Perplexity, and similar — correctly understand a company or service and cite it accurately. It is unrelated to security-industry abbreviations or to "all-in-one" product naming.
AI visibility means whether a brand is mentioned in an AI answer, whether the official site is cited, and whether the business is described accurately.
02 / Engines measured
- OpenAI Web Search (gpt-4o) — uses the Responses API's web-search tool; captures the answer text and cited URLs.
- Perplexity Sonar — implemented but not currently enabled. As of this page's publication, every measurement uses OpenAI Web Search only. If we enable it, we will state the engine used per measurement and update the version.
Both are API-based observations and are not guaranteed identical to what a person sees in the ChatGPT or Perplexity consumer apps.
03 / Fixed question set
Fixed-point baseline measurement runs a fixed set of 20 questions under identical conditions, across 4 categories: company recommendation (e.g. "who can I commission for AIO measures?") × 5, problem-solving (e.g. "why doesn't my company show up on ChatGPT?") × 5, purchase decision (e.g. "what does AIO work typically cost?") × 5, and target-company questions (e.g. "what does AIO work look like for a BtoB technical-services company?") × 5.
Individual-site AI-search-visibility scans generate prompts across 4 tracks — named, service, problem, and latest-news — for the specific target, and run each prompt multiple times (2 by default).
04 / Judging criteria
- Brand mention — a machine check of whether the target company's name (including spelling variants) appears in the answer text.
- Official-site citation — a machine check of whether the target's official domain appears among the answer's cited URLs. Mention and citation are tracked as separate metrics.
- Accuracy and freshness — an automatic provisional judgement (cross-checked against the self-reported business description and the answer excerpt) is confirmed or dismissed by a human. We do not call something "misinformation" on the automatic judgement alone; we keep a review state (automatic / human-verified / dismissed).
- Competitor identification — requires human review of the answer excerpt; not asserted automatically.
05 / Scoring
An individual scan is scored out of 100:
- AI presence rate — 25 points (mention rate across all prompts)
- Information accuracy — 30 points (based on the judgement; where unjudged, official citation is used as a proxy and flagged "needs confirmation")
- Official-site citation rate — 20 points
- AI answer share — 15 points (mention rate on category/problem questions)
- Information freshness — 10 points (whether the current core business is reflected; a proxy is used where unjudged)
Status labels (AI absent / misrepresentation risk / citation shortfall / competitor dominant / recognised, and similar) follow mechanically from the score and observation rate; only the misrepresentation-risk label is confirmed by a human before it is finalised.
06 / Handling variability
- AI answers vary run to run and over time, so we do not treat a single answer as a final result.
- Individual scans run multiple times per prompt and are evaluated as a rate.
- Fixed-point observation compares month to month, under the same fixed questions, engine, and conditions.
- We do not casually assert causation between an observed change and a specific action (experiment records keep a flag for whether a causal claim is allowed).
07 / Limits of this research
- Fixed-point observation (v1) runs each question once; more runs per question are planned for a future version.
- API-based observation does not reflect personalisation from login state, region, or history.
- Questions are in Japanese only.
- The lag before a web change is reflected in an AI answer (indexing lag) cannot be controlled.
- Accuracy judging is based on matching self-reported information against the answer excerpt, so anything outside what was self-reported cannot be judged.
08 / Methodology versions and history
| Version | Contents |
|---|---|
| 2026-07-02.v1 | Initial version. Fixed 20 questions (4 categories × 5), OpenAI Web Search (gpt-4o), one run each, machine judgement of mention and official citation. |
When the question set, judging criteria, run count, or engine changes, we update the version and add a history row here. We do not compare figures across versions directly.
09 / Disclaimer
This methodology and its observations do not guarantee search ranking, inclusion in an AI answer, citation, recommendation, or enquiry volume. An observation reflects "the result at that point, under those conditions"; future results may differ as AI services change how they behave.