RelayMag
Analytics

Why Marketing Benchmarks Are Mostly Useless

Key takeaways
  • Benchmarks average metrics that different companies define differently
  • The sample is usually one vendor's customer base, not an industry
  • Even an accurate benchmark can't tell you what to do about the gap
  • They work as a smoke alarm for measurement bugs, and fail as targets

Someone brings a slide to the quarterly review. The industry average email click rate is 2.6%, and yours is 1.9%. The room agrees this is a problem, someone is assigned to fix it, and the meeting moves on feeling productive. Nobody asks where 2.6% came from, which companies are in it, or what a click meant to whoever counted it.

That exchange happens constantly, and it's worth taking apart, because almost nothing about it holds up. The problem isn't that benchmark numbers are made up. Most are collected carefully by people doing honest work. The problem is that the thing they produce can't support the weight put on it.

The same word, three different metrics

Start with what's being averaged. A conversion rate needs a numerator and a denominator, and both are choices.

The denominator might be sessions, or users, or only sessions that reached a product page. Sessions and users diverge by a wide margin on any site with returning visitors, so the same underlying performance produces materially different rates depending on which one a company picked. Nobody's cheating. They just answered a setup question differently three years ago.

The numerator has the same problem with more room in it. One company counts any form submission. Another counts only leads that passed qualification. A third counts a demo request but not a content download. Averaging those together produces a number that describes no company's actual definition.

Two benchmark reports can disagree by a multiple on the same metric in the same industry, and both can be honest. That fact alone should change how much authority a single figure gets in a meeting.

The sample is a customer list

Benchmark data comes from somewhere, and it almost always comes from a vendor's own platform.

That means the sample is that vendor's customers. An enterprise marketing platform's benchmark describes companies large enough to buy an enterprise marketing platform. A self-serve email tool's benchmark describes small senders. Neither describes an industry, and both get published with the industry's name attached.

There's a second filter on top. The companies in the dataset are the ones still paying, which quietly removes everyone whose programme failed badly enough that they churned. The average survives upward for reasons that have nothing to do with performance.

None of this makes vendor data worthless. It's often the only data that exists at that scale, and collecting it is genuinely useful work. It does mean the honest label is "our customers last year" rather than "the industry", and the honest label is rarely the one on the chart.

Averages describe nobody

Marketing metrics tend to be skewed rather than symmetrical. A small number of very high performers pull the mean up, and the bulk of companies sit well below it.

In a distribution shaped like that, the average is a poor description of a typical company. Reporting a median helps, and better reports do. But a median tells you about the company sitting at the middle of a sample you have no other connection to, which isn't obviously more relevant to you than the mean was.

The variance is the part that gets dropped, and it's the part that mattered. A benchmark reported as a range with quartiles is a useful object, because it shows you how wide the spread is and therefore how much a single position within it means. A benchmark reported as one number has thrown that away before you saw it, and one number is what fits on a slide.

The gap doesn't tell you what to do

Suppose every objection above is handled. Same definitions, representative sample, full distribution. You're at the 30th percentile.

You still don't know anything actionable. The gap could be your pricing, your market, your product's fit, the mix of channels driving your traffic, or the fact that you sell something with a nine-month consideration cycle to people who visit eleven times before buying. Every one of those produces the same below-average number and every one implies a different response, or no response at all.

Benchmarks are diagnostic-shaped without being diagnostic. They tell you a gap exists and stay silent on cause, which is the only part that would have helped. That's why the meeting assigns someone to fix it. The number created an obligation without creating any information about what to do, and something has to fill that space.

Being above the benchmark is the more expensive failure mode. It reads as permission to stop, and the company that stops optimising a channel at the 70th percentile because a slide said it was fine has been actively harmed by the number.

What they get used for

The uncomfortable observation is that benchmarks mostly aren't used to learn anything. They're used to win arguments.

Someone who wants more budget brings the benchmark showing the company behind. Someone defending a channel brings the one showing it ahead. Both are available, because the definitional variance above means there's usually a report supporting either reading, and the person choosing which to present has generally decided the conclusion first.

That's not dishonesty so much as how evidence gets used in organisations. It's still worth naming, because it explains why benchmark slides are so persuasive and so rarely followed by anyone checking what happened next. The slide's job finished when the decision was made. The same dynamic makes certain metrics surface at certain moments for reasons unrelated to what they measure.

The one job they do well

Benchmarks have a real use, and it's narrower and less flattering than the way they're sold.

They're an excellent smoke alarm for measurement bugs. If a published range for your metric sits around 2% and you're reporting 40%, you have a tracking problem, and you've found it in seconds. Duplicate event firing, a bot filter that isn't on, a conversion counted at two stages of the same funnel. Order-of-magnitude comparison catches these fast and it's the fastest check available.

They also help calibrate people outside the team. A board member who believes 20% email click rates are normal is going to make bad decisions, and an outside number is more persuasive than an internal explanation. That's a communication use rather than an analytical one, and it's legitimate.

And with no data of your own, a benchmark is a reasonable input for deciding whether a channel is worth testing at all. Rough is fine when the decision is whether to spend two weeks finding out.

Notice what those three uses have in common. All of them are about orders of magnitude and none is about a few percentage points of difference. That's the resolution these numbers actually support.

What to compare against instead

Your own history is a better benchmark than any published one, for the reasons the published ones fail. The definitions are consistent because you chose them once. The sample is the business you're actually running. The comparison controls for market, product and audience automatically, because those are held constant.

Twelve months of your own conversion rate, with its seasonality visible, tells you more than any industry figure. It shows what's normal for you, how much it moves on its own, and therefore how big a change has to be before it means anything. That last part is what almost every benchmark conversation is missing.

Where you need to know whether something worked, the answer is a controlled comparison rather than any benchmark. A holdout group, the same window, everything else equal. That's the only structure that separates your work from everything else happening at the same time, and it's a different tool entirely from a number that describes other companies.

Building a comparison you can actually trust

If you want an outside number badly enough, the way to get a good one is to build it, and it's less work than it sounds.

Pick four or five companies genuinely comparable to yours, meaning similar motion, similar price point, similar sales cycle, and not competitors you'd mind talking to. Agree the metric definition in writing before anyone shares a figure, since that single step removes the largest source of error in every published benchmark. Then swap numbers directly.

People assume this conversation is impossible and it mostly isn't. Operators at non-competing companies trade this kind of thing readily, because the exchange is symmetrical and everyone involved has the same problem. A peer group of five with an agreed definition beats a sample of ten thousand with a vague one, because you know exactly what's in it and you can ask follow-up questions when a number looks strange.

The follow-up question is the real value. A published benchmark can't tell you that the company beating you does it by spending four times as much on the channel. A person can, in one sentence, and that sentence is the thing you were looking for when you opened the report.

Reading a benchmark report properly

Published benchmarks are worth reading. They're worth reading the way you'd read a survey, with the method in view.

Find out where the sample came from and whose customers those are. Find the definition of the metric and check it against yours before comparing anything. Look for the distribution rather than the headline, and treat a report that gives you quartiles as substantially more useful than one that gives you a number. Notice who published it and what they sell, which isn't cynicism so much as reading a source.

Our own piece on conversion rates by industry is worth holding to the same standard. It's a reasonable orientation for someone who has no idea whether their number is plausible, and it can't tell you why yours is what it is. That's a limit of the format rather than of that particular effort, and it applies to every version of it, including the ones with better data than ours.

The failure isn't publishing these numbers. It's the quiet promise that comparing yourself to one is the same as understanding your own performance. It isn't close, and the comparison is most seductive exactly when you have the least idea what's driving your actual numbers.

Frequently asked questions

Q: Are marketing benchmarks ever worth using?

A: Yes, for one job they do well. If your number is an order of magnitude away from a published figure, you probably have a measurement problem rather than a performance problem, and that's worth knowing within minutes. Benchmarks work as a smoke alarm. They fail as a target.

Q: Why do two benchmark reports give completely different numbers for the same metric?

A: Because they're measuring different things under the same word. Conversion rate can use sessions or users as the denominator, and can count any form submission or only qualified leads in the numerator. Those choices move the result by multiples, so two honest reports can disagree wildly without either being wrong.

Q: What should I compare my performance against instead?

A: Your own history, measured the same way over at least twelve months. The definitions are consistent, the sample is the business you actually run, and the comparison controls for everything a published average can't. Where you need a real answer about cause, a controlled test against a holdout beats any comparison to an outside number.

Q: Isn't being below the industry average still a useful signal?

A: Only if you already know why, and if you know why then the benchmark added nothing. The gap between you and an average could be your pricing, your market, your definition of a conversion, or the composition of the sample the average came from. None of those are distinguishable from the number itself.

R
RelayMag is an independent publication on marketing, search, and how companies get found.