Key takeaways
ChatGPT kept citing seven pages for four days after they began returning 404s, then dropped them to zero and kept them there.
Citation volume did not register the removal at all for two days: 221 the day before, 220 on the day itself, 229 the day after. It fell to 185 on day 2 and 64 on day 3.
Answer inclusion moved first. The share of citations reaching a final answer fell from about 81% before removal to 58% on day 0, while retrieval kept surfacing the same URLs at close to the old rate.
The same lag applies to other site changes, so a measurement window shorter than it will read a real change as no change.
What happens after you take a page down
Brands remove pages all the time. Sites get migrated, products get sunset, thin pages get pruned, and duplicate pages get consolidated.
In traditional search you can at least watch what happens next. Search Console lists which of your URLs are indexed and flags the ones now returning 404s, and inspecting any single URL tells you the last time Google crawled it. You also get a lever: the Removals tool pulls a page out of results within about a day. Though this removal lasts only for roughly six months, permanent removal occurs when Google recrawls the page and sees that it's gone.
LLMs don't provide any of that visibility. There's no index status report for ChatGPT and no lever to pull a URL out of the pool. Nothing tells you when the model last looked.
How ChatGPT interacts with your site
This is further complicated by how many separate ways a model can reach the web. OpenAI documents four crawlers, and three of them matter here: GPTBot gathers training data, OAI-SearchBot builds the index that surfaces sites in ChatGPT's search results, and ChatGPT-User fetches a page live when a question calls for it.
The index is doing two jobs at once. It is the catalog of pages eligible for a live fetch, and it also holds enough of what those pages said to inform an answer on its own. Which of those two jobs is doing the work decides how quickly a removal shows up in responses.
There isn't clear data on how long that window is, and it's a valuable figure to know for anyone managing their AI search visibility. Without it, you can't tell whether removing a page costs you anything, because you aren't sure if the changes will have occurred within the time horizon you're testing.
The timing is also informative in its own right. If the citations stopped straight away, live retrieval would be catching the 404 and gating what reaches the answer. If they ran on for days, the index would be the thing driving responses. And if the two layers disagreed, we would see it as a gap between how often a page still got retrieved and how often it survived into an answer.
We measured that window. We had daily citation data running on a set of pages that were about to be taken down, which let us watch the removal land in real time instead of reconstructing it afterwards.
What we tracked
We tracked seven pages on a brand's website that were removed and began returning 404s. All seven were being cited by ChatGPT beforehand, several of them heavily, and together they accounted for roughly a sixth of everything ChatGPT cited about that brand. These pages were tracked before and after they were removed from the site. Throughout what follows, day 0 is the first day the pages returned a 404, and days before that are counted back from it.
Two things make the before and after directly comparable. Collection was already running well before the pages came down, so none of this is reconstructed after the fact. And run volume was held constant across the whole window against a fixed prompt set, so any change in citations came from the model.
Pages elsewhere on the same site that stayed up are the control. If the removed pages go quiet while comparable pages keep getting cited, the silence belongs to the removals.
The setup also rules out the impact of training data because none of these pages was more than four months old, which puts every one of them after the February 2026 knowledge cutoff of ChatGPT's most recent models.
The citations kept coming for four days
The seven pages drew 221 citations on day -1, the day before they came down. On day 0, with every one of them returning a 404, they drew 220. On day 1 after the removal, they drew 229. The removal did not register in the citation volume at all for two days.
Day 2 brought the first real decline, to 185. Day 3 fell to 64. On day 4 the pages drew nothing, and they have drawn nothing every day since.

Figure 1: Every one of these pages was returning a 404 from day 0. Citation volume didn't move until day 2, and reached zero on day 4.
Comparable pages that stayed up kept getting cited straight through day 4 and after, so the zeros belong to the removals. Brand visibility fell over the same period, which tells us the citations these pages had been earning did not simply rotate to other pages on the site.
Answer inclusion fell before retrieval did
Petra separates a citation the model retrieved while working from a citation that survived into the answer a person reads. The removal showed up in that split before it showed up in volume.
Over the four days before removal, 81% of the citations to these pages made it into final answers. On day 0 that fell to 58%. It held at 58% on day 1 and recovered only to 63% on day 2. Retrieval kept surfacing the pages at close to the old rate, but what changed was how often the model was converting that retrieval into final citations.

Figure 2: Retrieval held steady while the model grew less willing to use what it found. The fall on day 0 is the live fetch registering the 404, at a point when the index was still nominating the same URLs.
For those four days the pages were still landing in answers people saw. Day 0 alone produced 127 final-answer citations from pages that were already returning 404s.
What the lag means for measurement
The presence of live retrieval bots doesn't mean LLMs respond to page changes immediately. It's tempting to read a live fetch as proof that the model sees your site as it stands today, and to treat crawler traffic on a URL as evidence the model is current on it.
But there is a lag between your current site and the version of it the model answers from. While this is relevant to page takedowns here, it also applies to site updates. Updated messaging and product catalog changes both take time to reach ChatGPT's answers.
Measurement windows on the impact of site updates must account for a lag. Too short a window returns a clean result that looks identical to a real one, so the change reads as inert and the next decision gets made on that footing. The lag also has to be accounted for in attribution.
A change that only shows up in visibility days later gets tied to whatever else happened on that later day. Run volume matters for the same reason, and we've measured how much of a short window is just sampling noise.
Methodology notes
Daily ChatGPT simulations at constant run volume across a fixed prompt set, collected through Petra's platform in August 2026. Day 0 is the first day the removed pages returned a 404. Interim citations (retrieved during the model's working process) and final citations (present in the delivered answer) are counted separately where noted. Page removal was confirmed by HTTP status. None of the pages was more than four months old.

