Author's Note
Key Takeaways
18pp
Gap in citation movement between CPG and consumer fintech
85%
Shift in the citation mix under a compliance-bound persona
97%
Of frequently cited sources in a single category are covered by just 8 of 15 prompt types
Glossary
Citation space: The set and distribution of sources a generative engine cites across possible ways of asking a related question.
Seed prompt: The baseline query from which controlled prompt variants are created.
Perturbation: A deliberate, typed edit to a seed prompt, such as a surface rephrasing, a contextual reframing, a persona change, a competitor reference, or an intent shift.
Prompt radius: The measured distance between a seed prompt's citation profile and that of a perturbation. Larger radius means a larger citation-space shift.
Weighted Jaccard Distance (WJD): A 0 to 1 measure of citation-profile difference that accounts for how often each source is cited.
Total Variation Distance (TVD): The fraction of citation attention that would need to move between sources to transform one profile into another.
Permutation-based noise floor: The amount of apparent movement expected from rerunning the same prompt distributionally. A pair exceeds its floor when its observed movement is unlikely to be explained by run-to-run stochasticity alone.
Upper-head citation space: Frequently cited paths retained at the 0.1% pooled-share threshold; the observable, reliably measurable portion used for coverage analysis.
Key Exhibits

This chart shows the change in cited share when the same question is edited in 14 different ways across four industries. Even the smallest tweak (a typo) shifted about 20% of citations, while rewrites that add the most context to the original query reshuffle citations 80% to 90%.
Every perturbation type sits above the permutation noise floor (dashed line, 95th percentile at roughly 7% TVD). Even the mildest edit moves more citation share than rerunning the identical prompt does.
Experiment Setup
We hand-wrote a controlled prompt panel and applied one common set of edits to every seed, which estimates each perturbation type across multiple topics rather than from a single query.
16 seed prompts across four industries: B2B SaaS, consumer packaged goods (CPG), consumer fintech, and healthcare, with four seeds per industry.
14 perturbations applied to every seed: five surface-form changes, eight contextual reframings that preserve commercial intent, and one shift to instructional intent. The crossed design yields 224 seed–perturbation pairs and 240 unique prompts, including the 16 unperturbed seeds.
150 runs per unique prompt, for 36,000 total runs collected over a six-day window on one engine and one product interface.
Each prompt is summarized as a citation-frequency profile, so the object of analysis is a distribution rather than one answer. We run the following analyses:
A permutation test to establish whether each difference surpassed the engine's run-to-run noise.
A bootstrap test to quantify the magnitude and confidence interval.
An investigation of domain turnover from seed to perturbation.
A study of which seeds and perturbations most effectively cover the total observable citation space.
Outline
Section I: Introduction — Measuring citation space
Generative engines are stochastic, and prior work shows that content-side edits can change what they cite. We vary the query instead, using controlled edits, and measure how far prompt changes move the citation distribution, whether that movement exceeds the noise floor, and how efficiently the observed space can be sampled.
Section II: Prompt construction — A shared perturbation panel
We write 16 seed queries across four industries and apply one common panel of 14 typed edits to every seed: five surface-form changes, eight contextual reframings that preserve commercial intent, and one shift to instructional intent. This crossed design yields 224 seed–perturbation pairs and estimates each perturbation type across multiple topics rather than from a single query.
Section III: Methodology — Prompts as distributional estimators
Each prompt is run 150 times and summarized as a citation-frequency profile. WJD compares pooled citation counts, while TVD measures bulk-tier citation-share reallocation. Pair-level permutation tests yield the noise floor, and BCa bootstrap intervals characterize magnitude and confidence. URL- and domain-level turnover analyses distinguish changes in pages from changes in publishers. A separate exact subset search measures how efficiently prompt types cover industry-level ground sets.
Section IV: Results — Citation movement is real, graded, and directional
At the pair level, 223 of 224 comparisons reject the WJD permutation null at α = 0.05. Surface edits produce less movement on average than contextual reframings (mean WJD 0.54 versus 0.80; mean TVD 0.39 versus 0.66). The how-to intent shift produces the largest mean movement (WJD 0.97; TVD 0.90). CPG moves the most on average (WJD 0.77; TVD 0.65), closely followed by B2B SaaS, while fintech moves the least (WJD 0.61; TVD 0.47).
Large-radius perturbations introduce and remove many URLs, whereas short-radius edits retain more of the original set. The same structure enables efficient in-sample measurement. Out of 15 prompt types, 5 cover 85.1% of the upper ground set, 7 cover 93.9%, and 8 cover 96.7%. Long-radius contextual and intent probes tend to add more exclusive paths, while several short-radius surface variants add none. The relationship is strong but not absolute.
Section V: Discussion — Building efficient prompt maps
The structure of prompt radius yields a campaign-specific way to construct a prompt map: begin with a broad baseline, then add prompt types whose exclusive paths justify their cost. The optimized subsets reported here are in-sample results from one engine, one product interface, a six-day collection window, four industries, and the studied prompt panel. Their exact ordering and budgets should be re-estimated for other engines, categories, campaigns, or time periods.
Section VI: Conclusion — A statistical language for citation space
Controlled prompt edits move citations well beyond rerun noise, yet that movement is structured around stable anchor domains, so a small optimized panel recovers most of the observed upper-head support. More broadly, prompt radius provides a common statistical framework for comparing industries, tracking citation drift, evaluating engines, and assessing future GEO interventions.
Abstract
Choosing which prompts to track determines what an AI visibility program can observe and take action on. Real buyers do not phrase a question one way: they vary wording and voice, and they carry context about who they are and what they need. Generative engines (GEs) respond to that context, so a single prompt may miss most of the answers a category's buyers encounter in the real world. Tracking every permutation is infeasible, which makes the operative question how efficient a prompt map is: the smallest set of prompts that covers the largest share of a category's consequential citations. Answering it requires separating prompt-induced citation movement from the run-to-run variability GEs exhibit under repeated identical queries. We introduce prompt radius, a statistical framework for characterizing citation patterns through controlled prompt variations, or perturbations. We apply a panel of 14 perturbations to 16 seed prompts across four industries, sample each prompt 150 times, and measure citation movement using Weighted Jaccard Distance, Total Variation Distance (TVD), and resampling-based inference. Movement is graded and almost always exceeds the noise floor: 223 of 224 seed–perturbation pairs reject the permutation null, with mean TVD ranging from 0.18 under a one-character typo to 0.90 under a shift in user intent. Surface edits move citations less than context edits (mean Weighted Jaccard Distance 0.54 versus 0.80), and the same panel produces systematically different movement across industries (0.61 in Fintech versus 0.77 in CPG). Reallocation occurs around a stable core: the top-20 domains hold 45% to 63% of all citations in every industry, and the highest-share domains are retained under nearly every framing. An optimized prompt map of 5 of 15 prompt types covers 85.1% of the citation paths above a 0.1% industry-wide share floor, and 8 cover 96.7%


