How to Run a Content Audit That Ends in Decisions
- Set your thresholds before you look at the data, or you'll rationalise every page
- Every URL gets one of four outcomes, and no page is allowed to stay unlabelled
- Consolidation is the highest-value action and the one teams skip
- Measure against a holdout group, because before-and-after can't separate your work from seasonality
A content audit has exactly one job, which is to decide what happens to every page you've published. Keep it, update it, merge it into something else, or take it down. If the exercise ends with a tab of 800 URLs and a column of notes, it wasn't an audit. It was an inventory, and inventories don't change anything.
The gap between those two outcomes isn't effort. Most audits collect more than enough data. The gap is that the decision rules get invented while the data is already on screen, which is exactly when they bend around whatever you're looking at. A page you spent a fortnight writing gets a generous reading. A page a departed colleague wrote gets a harsh one.
So the method below front-loads the judgement calls and leaves the data collection until they're already written down.
Decide the four outcomes first
Every URL in the audit ends up in one of four buckets, and the definitions get agreed before anyone opens analytics.
Keep: the page works and needs nothing. This should be a minority, and if it isn't, your thresholds are too soft.
Update: the page targets something people still want, but the content has aged out, the examples are stale, or it never fully answered the question.
Consolidate: several pages cover the same ground thinly. One of them becomes the real page and the others redirect into it.
Remove: the page has no traffic, no links, and no reason to exist. It goes, with a redirect if there's anywhere sensible to send people.
There's no fifth bucket, and specifically there's no bucket called "review later". That's where pages go to be audited again next year by someone who also doesn't decide.
Write the thresholds down before you look
Now set the numbers that sort pages into those buckets, while you're still ignorant of which pages they'll catch.
Pick a traffic floor over a twelve-month window, since anything shorter mistakes seasonality for decay. Pick a staleness date, meaning the publish or update date beyond which a page needs a fresh look regardless of performance. Pick a conversion rule, so a page with modest traffic that generates real pipeline is protected from the traffic floor.
Reasonable starting points are a floor of around 100 organic sessions a year, a staleness date of two years, and an exemption for any page with a meaningful conversion contribution. Those are starting points to argue about, not standards. The value is that you argued about them in the abstract.
Pull the data
Three sources cover almost everything.
- A crawl of the site, for the complete URL list plus titles, meta descriptions, word counts, status codes and canonical tags
- Analytics, for twelve months of sessions, engagement and conversions by landing page
- Search Console, for clicks, impressions, average position and the queries each URL actually gets seen for
Join those on URL into one sheet. The join is where you'll find your first surprise, which is usually the number of pages that appear in the crawl and nowhere else. Pages with no impressions at all across a year are invisible, and invisible pages are the easiest decisions you'll make.
Backlink data is a useful fourth source but not a blocker. Its job is narrow, telling you which dead pages deserve a redirect rather than removal.
Sort mechanically before you judge
Run the thresholds as formulas and let the sheet assign a first-pass bucket to every row. Nobody reads anything yet.
This step matters more than it looks. A mechanical first pass means the default outcome for a page is set by a rule you wrote earlier, and a human has to actively override it. Overriding is fine and often correct. Having to type a reason in the override column is what keeps it honest.
| Condition | First-pass outcome | What to do next |
| Above traffic floor, updated recently, converting | Keep | Nothing. Recheck at the next audit |
| Above traffic floor, older than the staleness date | Update | Refresh facts and examples, keep the URL |
| Below floor, but another page covers the same topic | Consolidate | Merge the useful parts, redirect to the survivor |
| Below floor, no links, no conversions, no strategic role | Remove | Redirect to the closest relevant page, or let it 404 |
| Below floor, has backlinks | Consolidate or keep | Never delete outright. Redirect at minimum |
| Zero impressions across twelve months | Remove | Check it isn't blocked or orphaned before removing |
The two checks the sheet won't do for you
Two problems don't show up in per-page metrics and both cause quiet damage.
The first is orphan pages, meaning pages with no internal links pointing at them. They're reachable in theory and invisible in practice, and a page in that state may be underperforming because nothing points at it rather than because it's bad. Fixing the links is a much cheaper action than a rewrite, and it needs testing before you conclude the content failed.
The second is cannibalisation, where several of your pages compete for the same query. You'll see it in Search Console as a term where the ranking URL keeps swapping. It's overdiagnosed in general, but where it's genuinely happening it's a strong argument for consolidation. Two pages half-answering a question rarely beat one page answering it.
Consolidation is where the value hides
Of the four outcomes, consolidation produces the biggest gains and gets used least, because it's the only one that requires real editorial work.
Pick the survivor by existing authority rather than by which draft you prefer, so usually the URL with the links and the ranking history. Move the genuinely useful material from the others into it, rather than stapling the text together, since a merged page that reads like three articles in a coat performs like three articles in a coat. Then redirect every merged URL to the survivor.
The redirect isn't optional and it's the step that gets forgotten in a rush. Deleting a page that had links throws away the one asset it had.
A caution worth holding. Consolidating pages that serve genuinely different intents makes both worse. A comparison page and a how-to page about the same product are not duplicates, even when the topic label matches.
Removal, done carefully
Removal has three flavours and picking the wrong one costs you.
Redirect the page when there's a relevant destination and the page has any links or residual traffic. That preserves whatever value it accumulated. Let it return a 404 or 410 when there's genuinely nowhere sensible to send someone, because redirecting everything to the homepage is worse than a clean 404 and mostly annoys people. Keep the page and noindex it when it needs to exist for users but shouldn't compete in search, which covers things like thin location pages and internal utility pages.
Do removals in batches with a record of what went where. When something breaks two months later, and occasionally it will, the record is what lets you reverse it.
Sequence the work by effort and payoff
The output is a queue, and the order matters because audits lose momentum.
Start with the cheap high-yield items, which are usually pages ranking just below the fold for terms they nearly win, and orphan pages that need links rather than rewrites. Those produce visible movement in weeks and buy patience for everything else.
Do consolidations next, since they're the biggest wins and the most work. Leave straightforward refreshes for the steady middle of the quarter. Do removals last and in bulk, because they're the least urgent and the easiest to get wrong when rushed.
Cap the queue at what the team can genuinely do in a quarter. An audit that produces 300 actions produces zero actions.
Measure against a holdout, not against last month
The instinct is to compare each updated page to its own past performance. That comparison can't tell you much, because the same window contains seasonality, algorithm updates, competitor changes and your own work, and it attributes all of it to you.
Set aside a group of pages that met the update criteria and deliberately leave them alone. Compare the updated group against that holdout over the same period. Both groups absorb the same external noise, so the difference between them is a far better estimate of what your edits did.
This takes discipline, since leaving pages untouched on purpose feels like negligence. It's the only version of this measurement that survives a sceptical question, and it's the same logic that makes a controlled test worth more than a chart with an arrow on it.
What an audit can't fix
An audit works on pages that already exist, which means it can't fix a strategy problem. If the library is full of content aimed at people who were never going to buy, the audit will tidy that content and the pages will keep not working. Pruning and refreshing raises the quality of what you have. It doesn't change what you decided to make.
Nor does it fix a production problem. Teams whose content underperforms for structural reasons will regenerate the same mess between audits, and running the audit more often just makes the cleanup more frequent. If the same categories keep filling with thin pages, that's a brief and commissioning issue, and no amount of auditing touches it.
A reasonable cadence
Once a year suits most libraries under a few hundred pages. Twice a year suits larger or faster-moving ones. Quarterly is almost always too often, because the changes need two or three months before their effect is legible, and auditing inside that window means reacting to noise.
Between audits, the useful habit is much smaller. When a page gets updated, record why. When one is removed, record where it went. That running log turns the next audit from an archaeology project into an afternoon, which is the difference between a process that survives and one that gets done once by someone who then leaves.
Frequently asked questions
Q: How often should you run a content audit?
A: Once a year for a library under a few hundred pages, and twice for anything larger or faster moving. More often than that and you're measuring noise, because the changes you make need two or three months to show an effect. Less often and the backlog of decayed pages gets big enough that nobody wants to start.
Q: Should I delete old content or just update it?
A: Deleting is right when a page has no traffic, no links, and no strategic reason to exist, and even then a redirect to the closest relevant page beats a 404. Updating is right when the page still targets something people want and has simply gone stale. The action people skip is consolidation, where several thin pages on one topic become a single good one.
Q: What data do I need to start?
A: A full list of URLs from a crawl, twelve months of sessions and conversions per page from analytics, and clicks and impressions per page from Search Console. That combination is enough to make every decision in this piece. Backlink data is useful for deciding what to redirect rather than remove, but you can run a first audit without it.
Q: How do I know whether the audit worked?
A: Compare the pages you changed against a holdout group of similar pages you deliberately left alone, over the same window. A straight before-and-after comparison can't separate your work from seasonality or an algorithm update, because both groups experience those and only one experienced your edits.
Q: What is keyword cannibalisation and does it matter?
A: It's when several of your own pages compete for the same query, which you can spot in Search Console when the URL ranking for a term keeps changing. It matters less than it's often claimed to, but where it's real it's a strong signal that those pages should be one page. Consolidating them usually helps more than optimising each one separately.