Who Does AI Search Actually Cite?
Netsujo Inc. · Published July 4, 2026 · Updated July 4, 2026
We ran a fixed set of 30 BtoB purchasing-and-procurement questions through OpenAI Web Search and classified every citation in the answers by hand. Under a realistic setting, 29 of the 30 questions returned zero citations; when we forced a web search on every question, 94% of the 137 citations were to companies' own official domains. Research date: 2026-07-04. Engine: OpenAI Web Search (gpt-4o). Version: 2026-07-04.r1 / 2026-07-04.r2.
01 / Background and question
In BtoB purchasing and procurement, the entry point for finding a supplier is shifting from search engines to AI search. Whether a company is cited in an AI's answer may become a new form of exposure, comparable to search ranking today.
A common assumption around AI-search citation is that "comparison sites and roundup media get cited" — but there is very little published data measuring the actual distribution of who gets cited, for Japanese-language BtoB queries. If the assumption is wrong, priorities for action are wrong too.
So we narrowed this study to one question: for BtoB procurement-style questions, who does AI search actually cite? We fixed the question set, the run conditions, and the classification rules, and had a human classify every cited URL.
02 / Research design
- Population — a fixed set of 30 questions modelled on BtoB procurement and purchasing decisions (6 categories × 5 questions: systems development, SaaS selection, AI adoption, web/SEO, manufacturing procurement, professional services). The questions, not a list of companies, are the population — no specific company is evaluated by name.
- Execution — each question run once, under two conditions: (1) a realistic setting where the AI decides whether to search (tool_choice=auto, "r1"), and (2) a setting that forces a web search on every question ("r2").
- Recording — full citation URLs, evidence of whether a search call happened, and each answer's opening text, all saved. Every cited URL's page title was fetched by HTTP and classified by a human, one by one, into operator type (official / media / unknown) and page type (service page / article / broken).
03 / Main findings
- Under the realistic setting (tool_choice=auto), 29 of 30 questions returned zero citations. The answers read like a list of well-known companies and services — general answers that look like they come from trained knowledge rather than a live search.
- The only question that triggered a search was the geographically specific one, "which company in Kyoto can I commission for systems development?" — all 5 citations for that question were to companies' own official sites.
- When we forced a search on every question (r2), 94% of the 137 citations (129) were to the official domain of a company or service provider; third-party comparison, roundup, or media sites accounted for 5.
- Of the 129 official-domain citations, 100 were service/company pages, 27 were explainer or comparison articles on the company's own domain, and 2 were pages now returning 404.
- The 137 citations spread across 133 unique domains; no domain appeared more than twice — we did not observe a small number of sites dominating.
04 / Breakdown by category
| Category | Citations | Official | Third-party media | Unclassifiable | Of which article-type |
|---|---|---|---|---|---|
| Systems development | 25 | 25 | 0 | 0 | 8 |
| SaaS selection | 15 | 11 | 4 | 0 | 7 |
| AI adoption | 23 | 21 | 1 | 1 | 5 |
| Web/SEO | 24 | 23 | 0 | 1 | 9 |
| Manufacturing procurement | 26 | 26 | 0 | 0 | 0 |
| Professional services | 24 | 23 | 0 | 1 | 3 |
| Total | 137 | 129 | 5 | 3 | 32 |
"Article-type" means explainer, comparison, or cost-guide articles (27 on official domains plus 5 on third-party media = 32). Figures are aggregated from a human-verified classification file.
05 / Data
| Citation operator type (r2, 137 citations) | Count |
|---|---|
| Company/service-provider official domain | 129 (94%) |
| Third-party comparison, roundup, or media site | 5 (4%) |
| Could not be classified (fetch failed) | 3 (2%) |
| Breakdown of the 129 official-domain citations | Count |
|---|---|
| Service/company page | 100 |
| Explainer or comparison article on the company's own domain | 27 |
| Page now returning 404 (broken link) | 2 |
This study observes the distribution of citation sources; it is not an evaluation of which listed domains are better or worse. Full data, including cited URLs and page titles, is kept in classification.json in the repository; the Japanese version of this page includes an expandable table of all 137 rows.
06 / Analysis
The main actor in citations was not comparison media, but companies' own official domains. Under the forced-search condition (r2), 94% of the 137 citations were to a company or service provider's own official domain. The assumption that "AI-search measures = getting listed on comparison sites and roundup articles" does not match what we observed across these 30 questions — what AI drew on to construct its answers was, in the main, the product or service's own primary source material. Third-party comparison and media citations numbered only 5, and 4 of those were concentrated in the SaaS-selection category, so where comparison-media exposure seems to matter may be limited to SaaS selection rather than category-wide.
Which type of page gets cited differs by category. Article-type citations (explainers, comparisons, cost guides) numbered 32 overall, concentrated in web/SEO (9), systems development (8), and SaaS selection (7) — in these areas, explainer content on the company's own domain functions as an entry point for citation. By contrast, all 26 citations in manufacturing procurement were official service or company pages, with zero article-type citations; in this category AI appears to reference primary pages showing what a company can make or sell, rather than explainer articles.
Citations were not concentrated — they spread across 133 domains, with no domain appearing more than twice. We did not observe the kind of concentration among a handful of top sites that is typical of page-one search results, so the opportunity to be cited looks broadly distributed. We did not classify the size or brand recognition of cited companies in this study, so any pattern by company size is a question for future research.
If no search happens, no citation happens. Under the realistic setting (r1, where the AI decides whether to search), 29 of 30 questions returned zero citations — as long as AI answers from trained knowledge, there is no room for a new company or page to surface. The only question that triggered a search was the geographically specific one about Kyoto. We cannot generalise from a single observation, but it suggests a hypothesis — that more specific, localised questions are more likely to trigger a search — which we plan to test in future studies.
07 / Limitations of this study
- Each question was run once, so a single answer is affected by run-to-run variance.
- The only engine measured was OpenAI Web Search (gpt-4o); behaviour may differ from the ChatGPT product, Perplexity, or Google's AI features.
- The question set was designed by us and carries selection bias (we have published all 30 questions in full as a research brief).
- The realistic-setting (r1) data has no record of whether a search call happened, so we do not claim more than "no citation was attached."
- This study does not guarantee the effect of any specific action, or inclusion/citation in an AI answer.
08 / Practical implications
- The primary battlefield for winning AI-search citations is your own official domain, not placement in external media — service and explainer pages that answer questions directly on your own site are the starting point.
- We only observed comparison-media placement translating into citation within SaaS selection. Outside SaaS, prioritising your official page's citability (structure, direct answers, clearly stated facts) over media placement is a reasonable basis for decisions.
- In areas like manufacturing procurement, primary pages showing product and manufacturing capability were referenced more than article content — which page type to strengthen depends on the business area.
- Two of the 137 citations were to pages now returning 404. Since past URLs can end up cited, keeping redirects in place when a URL changes is worth the effort from a citation-exposure standpoint too.
- Citations spread across 133 domains, suggesting the opportunity may be distributed toward any company with a page that directly answers a specific question.
09 / Reproduction
- Question set — a fixed 30 questions across 6 categories × 5 (the full set is published as a research brief); no question names a specific company.
- Run conditions — OpenAI Web Search (gpt-4o, ja-JP). r1 used tool_choice=auto (the AI decides whether to search); r2 forced a web search. Each question run once.
- Recording — all cited URLs, evidence of the search call, and the start of each answer were saved.
- Classification — fetched the page title of every cited URL by HTTP, cross-checked against URL structure, and had a human classify each one, one by one, into operator type (official/media/unknown) and page type (service/article/broken).
- Aggregation — figures aggregated from the classified data by category and type; unmeasured or unclassifiable items are stated as such, not shown as zero.
7 reasons AI search doesn't cite you(日本語)
What are AIO, GEO, and LLMO? A guide based on Google's official guidance(日本語)