Traffic is down, AI Overviews are the reason, and the advice arriving in every inbox is the same: audit the archive, delete the underperformers, refresh everything else. Content pruning has become the default response to a decline nobody fully understands.
It is the wrong response, and the evidence against it is not subtle. Google states the case explicitly in its own documentation, and the citation data points the same way. This piece is about what the numbers actually support, because the reflex is expensive and it removes the asset you need.
The reflex, and where it comes from
The logic feels sound. If AI answers are taking clicks, the reasoning goes, then only your strongest pages will survive, so cut the weak ones so the site looks leaner and fresher. Thin content dilutes quality signals. Old content signals neglect. Prune, refresh, republish.
Every step of that chain is an inference. None of it comes from a stated policy, and two of the steps are contradicted by the source everyone is inferring about.
We are not going to re-argue the size of the traffic decline here. Our pieces on ranking and AI citation decoupling and what publisher traffic collapse looks like at small scale cover the numbers and the measurement problems with them. Take the decline as given. The question is what to do about it.
Google says this plainly, in writing
Google’s helpful content documentation includes a self-assessment question that addresses the tactic by name:
"Are you adding a lot of new content or removing a lot of older content primarily because you believe it will help your search rankings overall by somehow making your site seem ‘fresh?’ (No, it won’t)"
The parenthetical is Google’s, not ours. That is about as direct as the documentation gets on any subject.
The same page warns against the cosmetic half of the practice: "Are you changing the date of pages to make them seem fresh when the content has not substantially changed?" Bumping a modified date without substantive change is listed alongside the tactics Google treats as manipulation.
Separately, John Mueller addressed the other half of the reflex: "keep in mind that removing content doesn’t make the rest rank higher." Deletion does not redistribute equity to survivors. There is no pool being freed up.
The number that undercuts the freshness argument
If AI Overviews rewarded recency, mass refreshing would at least be aimed at the right target. They do not.
A study of 1,000 AI Overviews collected in April 2026 found the median cited page was 14 months old. Not 14 days. The set of pages Google’s own AI chooses to cite skews substantially older than the content calendar most sites are being told to maintain.
A larger analysis, drawing on Ahrefs data across roughly 17 million cited URLs on seven AI platforms, adds nuance that is worth taking seriously rather than ignoring. It found AI-cited content averaging about 1,064 days old, close to 2.9 years, against organic top-ten results at about 1,432 days, close to 3.9 years. On that measure cited content is roughly 26 percent fresher than what ranks organically.
Both things are true and they point the same direction. AI citation does mildly favor newer material relative to organic results, and "newer" in this context means three years old instead of four. Neither figure supports deleting pages at twelve months or running a quarterly refresh cycle across an archive. A page old enough to be on most pruning lists is sitting comfortably inside the age distribution of cited sources.
That same analysis is also worth citing for what it debunks. A widely circulated claim that AI-cited content is "4.3 times fresher" than organic results traces to no primary study. If you have seen that figure used to justify a refresh program, the figure is unsourced.
What the citation data actually rewards
The same 1,000-overview study reports which page-level attributes correlate with being cited, and the pattern is consistent.
| Attribute | Reported effect |
|---|---|
| Pages over 2,500 words | Cited 1.6x more than pages under 800 words |
| At least one named source cited in the body | Cited 2.1x more often |
| Schema markup present | 2.3x more citations than unstructured equivalents |
| Domain authority proxy | +0.61 correlation with per-query citation rate |
Depth, sourcing, structure, and accumulated domain strength. Every one of those is built over time, and three of the four are damaged by aggressive pruning. Delete a hundred pages and you have not made the survivors deeper, better sourced, or more structured. You have made the domain smaller.
The concentration figure sharpens the point further. The top one percent of domains take 47 percent of all citations, the next nine percent take 31 percent, and the remaining ninety percent of domains share 22 percent. Citation follows domain-level standing, which is exactly the asset that takes years to build and that a mass deletion erodes.
Our piece on writing for AI citation covers the page-level craft this implies. The archive-level implication is the subject here, and it is the opposite of pruning.
How much to trust these numbers
Both studies are commercial analyses published by companies selling services in this space, not peer-reviewed research, and we should say so before anyone builds a strategy on them.
The 1,000-overview study describes its method in reasonable detail: 100 queries across each of ten intent classes, roughly thirty verticals, collected between April 8 and 22, 2026 on US-English desktop, private sessions with rotated IPs, 4,243 cited URLs crawled against about 50,000 control URLs. That is more transparency than most vendor research offers. It also has a real gap: it does not report whether cited pages ranked in the organic top ten, which would tell you how much of the citation effect is simply ranking under another name.
Treat the direction as well supported and the precise multipliers as soft. This is the same posture we recommended in vendor benchmarks versus independent leaderboards: a single vendor study establishes a hypothesis, not a fact. Crucially, the argument here does not rest on those multipliers. It rests on Google’s own written guidance, which says the tactic does not work, and on the age distribution, which two independent datasets agree is measured in years.
What to do with your archive instead
Content pruning is not always wrong. It is wrong as a response to AI Overviews. The distinction is what the action is for.
Consolidate rather than delete. Three thin pages on overlapping topics should become one substantial page with the others redirected. That preserves any accumulated links and moves you toward the depth the citation data rewards. Deleting all three moves you away from it.
Fix pages instead of refreshing dates. If a page is genuinely out of date, correct it, and say what changed. If nothing has substantively changed, changing the date is the tactic Google names as manipulation.
Add outbound citations to what you already have. Naming and linking a primary source is the single cheapest change available, it is reported as one of the strongest correlates of citation, and it can be applied to existing pages this week without rewriting them.
Deepen the pages that already have standing. A page that ranks and does not get cited is a better investment than a new page with neither. Extend it, source it, structure it.
Where content pruning still makes sense
None of this makes deletion always wrong, and it would be a poor reading of the evidence to conclude that nothing should ever come off a site.
Remove pages that are factually wrong and not worth correcting, pages describing products or processes that no longer exist where no reader is served by the record, duplicate or near-duplicate pages that should have been one page, and pages built purely to capture a search term that never delivered anything to the person who arrived. Those removals are defensible on their own terms and were defensible before AI Overviews existed.
What separates that from the reflex is the reason. Content pruning driven by a page-quality judgment is maintenance. Content pruning driven by a traffic chart is a guess about a mechanism Google has said does not work. The same action, taken for the second reason, is how sites end up smaller without becoming better.
The practical test is whether you can state what is wrong with a page without referring to its traffic or its age. If you can, remove it. If the only complaint is that it is old or that nobody visits, the evidence says leave it and improve it.
The uncomfortable summary is that there is no fast move here. The attributes that earn citations accumulate, and the reflex to prune is attractive largely because it feels like action. Our GEO versus SEO piece covers where the two disciplines diverge, and this is one of the places they do not: both reward depth built over time, and neither rewards a cleared-out archive.
Frequently Asked Questions
Does content pruning help with AI Overviews?
There is no evidence that it does, and Google’s helpful content documentation states directly that adding new content or removing older content primarily to make a site seem fresh will not help rankings. The parenthetical “No, it won’t” is Google’s own.
How old is the typical page cited in an AI Overview?
One study of 1,000 AI Overviews collected in April 2026 found the median cited page was 14 months old. A larger analysis across roughly 17 million cited URLs put the average at about 2.9 years, compared with 3.9 years for organic top-ten results.
So does freshness matter at all?
Mildly, and less than the advice suggests. AI-cited content is roughly 26 percent fresher than organic top-ten results on the larger dataset, but both populations are measured in years. That supports keeping content accurate, not deleting it at twelve months or refreshing on a calendar.
Does deleting weak pages help my strong pages rank?
John Mueller of Google has said that removing content does not make the rest rank higher. There is no equity pool that deletion redistributes. Removing genuinely unhelpful content is defensible as a quality action, but not as a way to promote the survivors.
What actually correlates with getting cited?
In the 1,000-overview study: length over 2,500 words cited 1.6 times more than under 800, at least one named source in the body cited 2.1 times more often, schema markup 2.3 times more, and a domain authority proxy correlating at +0.61. Depth, sourcing, structure, and domain standing.
Is the “AI cites content 4.3 times fresher” statistic real?
No primary study supports it. The analysis that examined the claim traced it to no source and reported a verified figure of about 26 percent instead. If a refresh program has been justified to you using the 4.3 times number, ask where it came from.
Should I ever delete content?
Yes, when it is wrong, obsolete, or serves no reader. That is a quality judgment and it is still correct. What does not work is deleting content because it is old, or because traffic fell, or to make a site appear fresher.
How reliable is the citation data in this piece?
Both studies are commercial analyses rather than peer-reviewed research, and one does not report whether cited pages also ranked organically, which is a meaningful gap. Treat the direction as well supported and the exact multipliers as provisional. The core argument rests on Google’s own written guidance rather than on either study.