Recap
- A blanket 'good reply rate' number is mostly noise. Results swing by ICP, offer, list quality, and company stage.
- Open rate is broken for measuring interest. Apple Mail Privacy Protection counts opens that no human made.
- The benchmark that matters is your own baseline. Compare against your past self with the same offer and list.
- Read published reports for direction, not as a grade. They average across things you cannot see.
Most cold email benchmarks do not apply to you. A reply rate is the product of who you sent to, what you offered, how clean the list was, and how known your company is, and those four things differ for every sender. A seed-stage founder emailing 200 hand-picked accounts and a 40-person team blasting 50,000 contacts cannot share a benchmark. So stop chasing someone else's number and build your own baseline.
Why does one benchmark number not work?
Because a reply rate is a property of your specific situation, not of cold email. The same copy that pulls eight percent from a tight, in-pain list pulls a fraction of one percent from a scraped industry dump. When a report says the average reply rate is X, it has averaged across senders who have nothing in common with each other or with you.
Four variables move the number more than your copy ever will.
| Variable | Pushes reply rate up | Pushes it down |
|---|---|---|
| ICP fit | Narrow, in-pain segment | Whole industry, no filter |
| Offer strength | Clear, urgent, easy yes | Vague 'quick chat' ask |
| List quality | Verified, recently sourced | Scraped, stale, role accounts |
| Company stage | Known brand, warm category | Unknown name, cold category |
Change any one of these and the benchmark you were chasing becomes meaningless.
What is wrong with open rate as a benchmark?
Open rate stopped measuring interest in 2021, when Apple Mail Privacy Protection began pre-loading the tracking pixel inside every message it fetched. The open gets logged whether or not a person ever looked. A large share of inboxes now run through Apple Mail, so a chunk of your open rate is machines, not humans.
That does not make the metric useless. A sudden drop in opens can flag a deliverability problem before replies fall off. But treat it as a smoke alarm, not a scoreboard. The teams that optimize copy to lift open rate are tuning a number that no longer maps to attention.
What should you actually measure against?
Yourself, last month, with the same offer and list quality. That is the only comparison that controls for the variables nobody else's report can. Pick a small set of metrics that tie to revenue, segment them by ICP, and watch the trend rather than a single snapshot. The goal is to know whether this week beat last week, and why.
- Positive reply rate. Replies that show real interest, not 'unsubscribe' or 'wrong person'. This is the metric closest to pipeline.
- Meetings booked per thousand sent. Cuts through everything upstream. If this climbs, the machine is working.
- Bounce rate. Your list-quality early warning. Rising bounces mean stale data and a deliverability risk, full stop.
- Reply rate by segment. One blended number hides the truth. One ICP might pull ten percent while another pulls nothing.
Track these per ICP and per offer. A blended average across segments is how you stay convinced an experiment worked when only one slice carried it.
How do you read a published benchmark report without getting fooled?
Read it the way you would read a stranger's restaurant review. Useful for direction, useless as a verdict on your kitchen. Before you trust any figure, ask what it actually counted and what it quietly left out. Most reports answer none of these questions.
- What counts as a reply? All replies, or only positive ones? Auto-responders and 'not interested' often get bundled in.
- Whose data is this? A report from a tool sells the best outcomes its best users got, not a typical result.
- What ICPs and offers are inside? A blend of SaaS, agencies, and recruiting tells you nothing about your vertical.
- What volume and time window? Averages over millions of sends hide the long tail where most senders actually live.
If a report cannot answer those, its headline number is a vibe, not a benchmark. Borrow the structure and the questions it raises. Throw away the average.
How do you set a baseline when you are just starting?
You will not have a baseline on day one, and that is fine. Send to one tight, well-defined ICP with one offer, and keep the variables steady long enough to read a signal. A few hundred sends to a consistent segment is usually enough to see a trend. Below that, every result is an anecdote, so resist the urge to rewrite everything after twenty sends.
Once you have a baseline, change one thing at a time. New subject line, same list. New segment, same copy. If you swap the list and the copy and the timing all at once and replies double, you have learned nothing you can repeat. This discipline is also where an autonomous operator like LaunchSurface earns its keep. It holds the variables steady, runs the comparison against your own history, and shifts capacity toward the segments that actually reply.
What does this mean in practice?
Stop screenshotting benchmark charts and stop asking whether your reply rate is 'normal'. Normal does not exist across senders this different. Define your ICP, pick three or four metrics that map to revenue, segment them, and beat last month. The only benchmark you can trust is the one you measured yourself, under conditions you controlled. Everything else is someone else's average wearing a confident headline.
