Research Report

How Access Surface Shapes LLM Search and Citation Behavior

Sami Akkawi

,

CEO, Co-Founder

April 2026

Download the white paper

Author's Note

Most brands track their AI visibility and performance using one of the many new-age AEO software vendors (including ourselves). For ChatGPT specifically, most of these tools gather their data from logged-out sessions. Some tools gather data directly from the API. But, none are able to gather data from logged-in, paid accounts, at scale.

This poses a major question. Does your AI visibility tool actually reflect what your customers are seeing in the real world? We conducted this research, the first of its kind, to test whether the same model, given the same prompt, produces systematically different search behavior, citations, and brand recommendations across access surfaces.

Sami Akkawi, CEO, Co-Founder

Most brands track their AI visibility and performance using one of the many new-age AEO software vendors (including ourselves). For ChatGPT specifically, most of these tools gather their data from logged-out sessions. Some tools gather data directly from the API. But, none are able to gather data from logged-in, paid accounts, at scale.

This poses a major question. Does your AI visibility tool actually reflect what your customers are seeing in the real world? We conducted this research, the first of its kind, to test whether the same model, given the same prompt, produces systematically different search behavior, citations, and brand recommendations across access surfaces.

Sami Akkawi, CEO, Co-Founder

Most brands track their AI visibility and performance using one of the many new-age AEO software vendors (including ourselves). For ChatGPT specifically, most of these tools gather their data from logged-out sessions. Some tools gather data directly from the API. But, none are able to gather data from logged-in, paid accounts, at scale.

This poses a major question. Does your AI visibility tool actually reflect what your customers are seeing in the real world? We conducted this research, the first of its kind, to test whether the same model, given the same prompt, produces systematically different search behavior, citations, and brand recommendations across access surfaces.

Sami Akkawi, CEO, Co-Founder

Glossary

Access surface — the specific mode through which a user reaches a model (paid logged-in chat, free logged-out chat, or the programmatic API), all resolving to the same underlying model.

Fan-out queries — the intermediate web searches a model issues during its research phase, before it composes a final answer.

Interim citations — the sources a model retrieves and consults during its research phase.

Final citations — the subset of sources the model explicitly cites in its final, user-facing response, recorded at both the domain and full-path level.

Visibility — the percentage of trials in which a given brand appeared in the model's recommendations.

Entity rank — the average ordinal position of a brand when it appeared in the final response; a rank of 1.0 means it was always listed first.

Key Takeaways


32PP

Largest swing in one brand’s visibility between surfaces

70%

API searches that repeat verbatim, vs under 10% on chat

95%

Interim sources that API keeps in its final answer, vs 16% on Logged-In

Key Exhibits


Brand visibility by access surface, for the six most-recommended ski brands. Each bar is the share of 300 trials in which GPT-5.2 recommended that brand; the same brand can rise or fall by more than 30 points based only on which surface was queried.

  • The two chat surfaces track each other more closely than either tracks the API, yet no surface is a stand-in for another. Even between logged-in and logged-out results, 4 of 6 brands swing by more than 15% visibility.

  • Blizzard peaks on the API (62%) while Line and Völkl peak on chat and fall to ~18% on the API.

Experiment Setup

We submitted a single fixed product-research prompt to three access surfaces of OpenAI's GPT-5.2, 300 times each, and captured the responses exactly as an end user would receive them. All trials ran on one day to hold the search index, model version, and web content constant.

  • 900 total trials — 300 runs of the same prompt on each of the three surfaces, all completed on February 26, 2026, ruling out temporal variation in search indices, model versions, and web content.

  • One multidimensional prompt — a request for premium all-mountain twin-tip skis encoding several competing constraints.

  • Three surfaces of the same model — Logged-In paid ChatGPT and Logged-Out free ChatGPT (each trial in a fresh incognito window, no history or custom instructions) via the "Auto" selector, plus the Responses API (gpt-5.2-chat-latest, temperature 0.7, no system prompt), with web search enabled and forced on every trial.

Outline

Section I: Introduction — An open question in the literature

Prior work compared different ChatGPT models and documented that LLMs are non-deterministic, but not whether the same model behaves differently across the surface it is queried through. This paper closes that gap for what matters commercially: the queries a model issues, the sources it cites, and the brands it recommends. 

Section II: Experiment Design — One prompt, three surfaces, 900 trials

We submitted a single fixed all-mountain ski prompt 300 times to each of three GPT-5.2 surfaces (Logged-In paid chat, Logged-Out free chat, and the Responses API), all on one day with web search forced. From each response we captured fan-out queries, interim and final citations, and per-brand visibility and rank.

Section III: Results — Fan-out queries

The API issued the most searches per trial, but recycled the same few templates. The chat surfaces wrote longer, natural-language queries that reused fragments of the user's own prompt. Even the two chat surfaces diverged, with Logged-In retaining nuance words like "fun" and chasing different products.

Section III: Results — Citations and source selection

Chat surfaces consulted far more sources during research, Logged-In most of all, but discarded the bulk of them. API kept nearly everything it read, which left it citing twice as many sources. API’s final mix was almost half marketplace and social, against two-thirds independent editorial for Logged-In.

Section III: Results — Brand visibility and rank

The two chat surfaces tracked each other far more closely than either tracked the API, yet even between them, brands swung by as much as 23pp. Against the API the gaps topped 32pp: Blizzard doubled on the API while Line and Völkl collapsed there. One brand, Renoun, showed up across chat trials but never once on the API.

Sections IV–V: Conclusion

Practical guidance for anyone whose visibility depends on AI discovery: no single surface is a proxy for the model, so a brand's presence and its supporting citations can hinge on the access surface queried, and single-surface studies may not generalize.

Abstract

Large language models with integrated web search are becoming a primary channel for consumer product research, and major providers expose the same underlying model through multiple access surfaces—paid authenticated chat, free unauthenticated chat, and programmatic APIs—that may differ in undocumented ways. Whether the same model produces systematically different outputs across these surfaces given an identical prompt has not been investigated empirically. The answer has direct consequences for research reproducibility and for any entity whose commercial visibility depends on AI-mediated discovery. Across 900 trials, we tested this question for OpenAI’s GPT-5.2 by submitting the same multidimensional product research prompt 300 times each to three surfaces (Logged-In ChatGPT, Logged-Out ChatGPT, and the Responses API), all conducted on a single day, and analyzing the resulting search queries, cited sources, and recommended brands. The three surfaces diverged on every dimension measured. The two chat surfaces behaved more like each other than either behaved like the API (brand-visibility correlation of r = 0.891 between chat surfaces vs. 0.711–0.815 against the API), though they still differed meaningfully in which prompt-derived nuance words they retained and which specific products they drilled into. The API issued more search queries per trial (4.08 vs. 2.75–2.95) but reused the exact same query wording across trials 70% of the time (vs. 5.6–9.5% for the chat surfaces), drew from rigid templated constructions rather than varied natural language, and retained nearly every source it consulted (5% filtered vs. 73–84% for the chat surfaces). Its brand recommendations over- or under-represented specific brands by as much as 32 percentage points relative to the chat surfaces, and one brand present in 15–18% of chat-surface trials was absent entirely from the API. These findings suggest that a single access surface cannot serve as a general proxy for model behavior, and that commercial visibility in AI-generated recommendations depends materially on which surface a user queries.

Let’s turn AI search into your next growth channel

Let’s turn AI search into your next growth channel

Let’s turn AI search into your next growth channel