Antony Deroshan · Figshare 2026 · 2026
DOI: 10.6084/m9.figshare.33994620.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Search-enabled language models such as ChatGPT may call a web-search tool or answer from parametric knowledge alone. When the tool is called, the model writes a first rewritten web query—the fanout query, q1—before any observed result is returned, and that fanout already inserts brand names. Prior measurement finds brands named in the fanout are cited 68.9% of the time, against 2.1% for names only fetched later: those brands have already won visibility before retrieval happens. This study asks three questions. Can ChatGPT know a brand and still omit it from the fanout query? What makes a brand enter the fanout rather than fail to make the cut? Does category-relevant web exposure track where a brand sits in the underlying ranking? The results have three parts. First, the model can know a brand and still omit it: asked with search off, it names Campaign Monitor on 49 of 50 prompts, but in live fanout queries issued via the OpenAI Responses API with web_search it writes Campaign Monitor into only 1 of 50; seven other category-native brands show the same gap. Second, the omission is decided when the query is written and is a budget effect. Using a fanout query simulator validated against real OpenAI fanout data (Spearman rho = 0.965), we define an entry threshold K50—the smallest name budget K at which a brand appears in at least half of prompts—that measures how deep each brand sits. Third, category-conditioned unique-host exposure in Common Crawl tracks that depth (rho = +0.76 on 29 brands, +0.85 on the 25 with defined K50, and +0.986 on a held-out e-signature category). Exposure must be category-conditioned, and quality resampling or fuzzy dedup does not strengthen the association. This is a measurement of fanout brand selection on one model family, not a causal account: it does not show that web exposure trains the ranking, and K50 is an instrument defined on the authors' own rewriter prompt rather than a production parameter.
No comments yet — start the discussion below.