Author's Note
Key Takeaways
2.8X
more fan-out queries per run from Opus relative to ChatGPT and Sonnet.
70%
shift in final citation mix between Claude Sonnet and Opus.
93%
shift in final citation mix between ChatGPT and Claude Sonnet.
Glossary
Generative Engine: The system and infrastructure which orchestrates generative AI models (such as LLMs or diffusion models) to synthesize original responses to prompts.
Fan-out Query: Sequential sets of search strings issued by the engine to retrieve sources.
Lead Query: First fan-out query, if multiple fan-out queries exist.
Follow-up Query: Fan-out queries issued after the lead query, if multiple fan-out queries are issued.
Interim Citations: URL paths and domains returned by fan-out queries.
Final Citations: URL paths and domains included as supporting evidence to accompany the engine’s final response.
Key Exhibits

This chart shows how often five engines name each of the top 10 brands in four categories, as a share of answers that recommend at least one brand. Only six of the 40 brands appear in nearly every answer on every engine, while 26 swing by at least 50 points depending on the engine. For example, Sunbasket appears in 99.7% of Sonnet's meal kit answers and 0.2% of Google AI Overviews'.
High visibility on one engine does not carry over to another. Bank of America appears in 98.2% of ChatGPT's secured card answers and 3.4% of AI Overviews'. Even two Google products split: Discover appears in 99.6% of AI Overviews' answers but only 20.3% of Gemini's.
Experiment Setup
We ran a fixed prompt template across five engines.
4 prompts across four industries: secured credit cards, meal kit delivery, cryptocurrency exchanges, and live sports streaming, all written from a single template (“Recommend {X}.”).
5 engines on their consumer product surfaces: ChatGPT (GPT-5.5 Instant), Claude Sonnet 5, Claude Opus 5, Gemini 3.5 Flash-Lite, and Google AI Overviews. ChatGPT, Sonnet, and Opus expose fan-out queries and interim citations; Gemini and AI Overviews expose final citations only.
500 target runs per prompt per engine, for roughly 10,000 total runs collected from a United States vantage point between July 27 and August 2, 2026.
Each run is traced from search to answer: the fan-out queries issued, the pages returned, the pages cited, and the brands named in the response. We run the following analyses:
A tagging of every fan-out query by position (lead or follow-up) and content (branded, regulatory, both, or general).
A retention and leak analysis to measure how much of what an engine retrieves it ultimately cites.
A permutation test to establish whether cross-engine citation distances exceed each engine's run-to-run noise, including a comparison under a shared lead query.
A study of brand visibility across engines, and of how visibility shifts when a run issues a branded or regulatory fan-out query.
Outline
Section I: Introduction — Tracing the path from search to recommendation
Generative engines are a growing channel for product recommendations, and the same prompt can lead different engines to search different evidence and recommend different brands. We treat each engine as an observable pipeline and compare it stage by stage: the searches it issues, the pages it retrieves, the sources it cites, and the brands it names.
Section II: Experiment setup — One prompt template, five engines
We run four recommendation prompts from a single template (“Recommend {X}.”) on five engines: ChatGPT, Claude Sonnet, Claude Opus, Gemini, and Google AI Overviews, targeting 500 runs per prompt and engine. The prompts cover secured credit cards, meal kit delivery, cryptocurrency exchanges, and live sports streaming. Engines expose different parts of the pipeline. ChatGPT, Sonnet, and Opus log fan-out queries and interim citations, while Gemini and AI Overviews expose final citations only. Opus was collected on July 27, 2026, and the other engines on August 2.
Section III: Methodology — Four observable artifacts
We measure four artifacts per run. Fan-out queries are tagged by position (lead or follow-up) and, by keyword, as branded or regulatory. Interim citations are the pages those queries return, and final citations are the pages cited in the response. A brand's visibility is the share of an engine's answering runs that name the brand or its products. Citation profiles are compared with TVD against permutation noise floors, and brand visibility is compared through cross-engine ranges.
Section IV: Results — Engines diverge at every stage
Opus averages 2.77 fan-out queries per answer and routinely follows up with branded or regulatory searches; ChatGPT averages 1.02 and Sonnet 0.82. Engines also retrieve and cite different pages. Final path TVD ranges from 0.698 (Opus vs Sonnet) to 0.927 (ChatGPT vs Sonnet), against within-engine noise floors of 0.078 to 0.151, and the gaps are already present at the interim stage.
Brand visibility diverges more sharply. Six of the 40 top brands exceed 99% visibility on every engine, while 26 differ by at least 50 percentage points. Branded and regulatory fan-out queries coincide with large shifts in which brands are named. When ChatGPT's lead query names Coinbase, Kraken, and Binance, OKX visibility falls from 39.8% to 2.7% and Robinhood rises from 4.9% to 27.0%. The design cannot distinguish whether the query causes the shift or records a plan the engine had already formed.
Section V: Discussion — Engines are not interchangeable
Because engines agree on only a handful of brands, visibility on one engine is a poor guide to visibility on another, and similar citations do not imply similar recommendations. Programs should measure visibility by engine and prompt. These results are observational and come from four industries, one prompt format, and non-simultaneous collection dates. Controlled experiments that vary queries independently of retrieved evidence are the next step.
Section VI: Conclusion — Measuring the pipeline, not just the answer
Generative engines leave observable traces at each stage of the recommendation process, and those traces differ substantially across engines. Tracking fan-out queries, interim citations, final citations, and brand visibility separately, across multiple engines, shows where a brand gains or loses visibility and gives future experiments concrete conditions to test.
Abstract
Generative engines (GEs) are decision-making pipelines that search the web, select evidence, and use that evidence to construct recommendations for a user. In this campaign we ask whether different engines follow different search strategies and whether those strategies shape the brands they recommend. We use one prompt template (“Recommend {X}.”), four prompts, and a single United States collection window. We examine four types of artifacts from the search-to-recommendation process: fan-out queries, interim citations, final citations, and brand visibility. We compare ChatGPT 5.5 Instant, Claude Sonnet 5, Claude Opus 5, Gemini 3.5 Flash-Lite, and Google AI Overviews with the Gemini 3 model family on the same four prompts. In this campaign, the five engines showed substantial differences in search and brand visibility, and we observed four main patterns in the data. First, the number of fan-out queries issued by each engine differed: Claude Opus averaged 2.77 fan-out queries per run and frequently issued follow-up searches, whereas ChatGPT and Claude Sonnet typically issued only one search. Second, interim and final citation overlap was low. Total variation distance between two engines’ interim citation profiles ranged from 0.743 to 0.887, and between their final citation profiles from 0.698 to 0.927. Third, most brands showed wide cross-engine visibility ranges. Out of the top 10 brands in each prompt, 26 of the 40 most frequently mentioned brands differed by at least 50 percentage points across engines. Fourth, branded and regulatory fan-out queries often coincided with large shifts in the brands named in the response. These findings, sourced from a single sample across four industries, suggest that AI visibility programs should evaluate engines independently rather than assume interchangeable behavior.

