To split test sales outreach you need enough volume to read the result. At a 5% reply rate, detecting a 50% lift (5% to 7.5%) takes 1,467 sends per variant — 2,934 in total. At 500 sends a week that is 6 weeks. At 100 a week it is 29 weeks, so the test is over before it can conclude.
At a glance: how much volume do I need to A/B test?
- The metric: reply rate = replies ÷ messages delivered, counted per unique contact, over a fixed send window.
- The formula: n per arm = (1.96 + 0.84)² × [p₁(1−p₁) + p₂(1−p₂)] ÷ (p₁−p₂)², at 95% confidence and 80% power.
- The shortcut: about 75 replies in the control arm before a 50% improvement is readable. About 400 before a 20% improvement is.
- Test first: who you send to, then the offer, then the channel. Subject lines last — they move the number least and so need the most volume.
- Under ~250 sends a week: stop running head-to-head tests. Make sequential changes big enough to see with the naked eye, and fix volume first.
How it works
How to run a split test that can actually conclude
Size the test first
Pick the smallest lift worth acting on. At a 5% reply rate, reading a 50% lift needs 1,467 sends per arm.
Check it against volume
Divide total sends needed by your weekly volume. If the answer runs past eight weeks, test something bigger instead.
Randomise per contact
Assign each unique contact to one arm and hold that assignment across every touch in the sequence. Per-message randomisation contaminates both arms.
Read it once, at n
Wait until both arms reach the number, then read it. Watch deliverability and opt-outs daily, never the metric under test.
MAKE MORE SALES.
Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.
How is reply rate actually measured?
Reply rate is unique contacts who replied ÷ unique contacts who received a delivered message, over a stated window. Three details decide whether the test means anything, and most teams get at least one wrong.
Delivered, not sent. Bounces and filtered mail are not exposure. If one variant lands in spam more often you are measuring deliverability, not copy — a different problem with a different fix, covered in email and SMS deliverability for outbound at scale.
Unique contacts, not messages. A six-touch sequence sent to 500 people is 500 data points, not 3,000. Counting touches inflates your sample by the length of your cadence and makes every test look significant.
One window, both arms. If variant A ran in March and variant B in April, the difference includes March and April. That is a before/after comparison, not a split test.
Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.
The sample size calculation, worked end to end
This is the standard two-proportion sample size formula. Nothing here is proprietary — it is in every biostatistics text, and you can check it in five minutes.
n per arm = (Zα/2 + Zβ)² × [p₁(1−p₁) + p₂(1−p₂)] ÷ (p₁ − p₂)²
At the conventional settings — 95% confidence two-sided, 80% power — Zα/2 = 1.96 and Zβ = 0.84, so the leading term is (1.96 + 0.84)² = 7.84.
Say your current reply rate is 5% and you want to know whether a new opening message gets you to 7.5%:
- p₁(1−p₁) = 0.05 × 0.95 = 0.0475
- p₂(1−p₂) = 0.075 × 0.925 = 0.069375
- Sum = 0.116875
- (p₁ − p₂)² = 0.025² = 0.000625
- n = 7.84 × 0.116875 ÷ 0.000625 = 1,466, round up to 1,467 per arm
Two arms, so 2,934 sends before the test can conclude. Substitute your own reply rate and your own target and the arithmetic takes a minute. One caution: this normal approximation degrades when the expected count in any cell drops below about five, which the NIST/SEMATECH e-Handbook of Statistical Methods states as the min{Np, N(1−p)} ≥ 5 restriction. Below that, use Fisher’s exact test.
How long does a split test take at my weekly send volume?
This is the table almost nobody publishes, because it says most outreach tests cannot work. Total sends needed: 864 for a doubling, 2,934 for a 50% lift, 16,292 for a 20% lift, all from a 5% baseline at 95%/80%. Divide by your weekly volume.
| Sends per week (both arms) | Read a doubling 5% → 10% |
Read a +50% lift 5% → 7.5% |
Read a +20% lift 5% → 6% |
|---|---|---|---|
| 100 | 9 weeks | 29 weeks | 163 weeks |
| 250 | 3 weeks | 12 weeks | 65 weeks |
| 500 | 12 days | 6 weeks | 33 weeks |
| 1,000 | 6 days | 3 weeks | 16 weeks |
| 2,500 | 2 days | 8 days | 7 weeks |
| 5,000 | 1 day | 4 days | 3 weeks |
| 10,000 | 1 day | 2 days | 11 days |
Read the bottom-right against the top-right. The same question a 10,000-a-week sender answers in eleven days takes a 100-a-week sender three years. The constraint on testing is not discipline or tooling. It is volume.
The arithmetic is identical for any rate. At 25 booked calls a week, a five-point move in show rate (60% to 65%) needs 1,467 calls per arm — 117 weeks, or two and a quarter years. At that volume a five-point move is not a result you can obtain; it is a result you can only assume.
If we can’t make you money, we don’t deserve yours.
Pay-Per-Result pricing — performance-based alignment.
The 75-reply rule: count replies, not sends
You cannot read a 50% improvement until the control arm has produced about 75 replies. That is the whole rule, and the useful thing about it is that the number barely moves with your reply rate — it is 77 replies at a 1% base rate, 73 at 5%, and 68 at 10%. Sends vary enormously; the replies you need do not.
| Lift you want to detect | Replies needed in the control arm | Total sends at a 5% reply rate |
|---|---|---|
| Doubling (5% → 10%) | ~22 | 864 |
| +50% (5% → 7.5%) | ~73 | 2,934 |
| +30% (5% → 6.5%) | ~189 | 7,546 |
| +20% (5% → 6%) | ~407 | 16,292 |
| +10% (5% → 5.5%) | ~1,560 | 62,392 |
Use it as a stopping rule you can apply from your inbox. Count replies in the control arm: under 22 you cannot read a doubling, under 75 you cannot read a 50% lift. Anyone declaring a winner before those counts is reading noise and rolling it into next quarter’s plan.
What should I test first when my volume is small?
Rank by effect size, because at low volume only large effects are legible. Four weeks at 500 sends a week gives 1,000 per arm, which can only detect a lift of about +62% or larger. Anything smaller is invisible however carefully you run it, so test only what could move the number by more than half.
| # | Lever | What it changes | Worth testing at 500 sends/week? |
|---|---|---|---|
| 1 | Who you send to | The audience itself — source, segment, recency, firmographic fit | Yes. Segment changes routinely clear +62% |
| 2 | The offer | The reason for contact and what the reply gets them | Yes |
| 3 | Channel mix | Email only vs email plus SMS or voice | Yes |
| 4 | Follow-up depth | Two touches vs six across a longer window | Marginal — 8+ weeks to read |
| 5 | Send timing | Day, hour, and response lag on inbound | No. Effect is real but too small to isolate here |
| 6 | Subject line or first sentence | The wording only | No. Needs 60+ weeks at this volume |
The ranking is our operating judgement from the campaigns we run, not a measured league table; the +62% threshold in the last column is arithmetic and you can check it. “We test our subject lines” is the most common answer to how a team tests, and at typical volumes it is the one test that cannot possibly conclude. Test the audience before you test the adjectives.
Readable wins compound, which is why this is worth the trouble. Four sequential links — contact, reply, booked, shown, the stages costed out one by one in our sales pipeline stages guide — each improved 15% gives 1.15⁴ = 1.75×: a 75% lift in bookings from four changes none of which looked impressive alone. Five links at 20% each gives 2.49×. That is the honest mechanism behind large multiples, set out in full in why small conversion gains compound.
Why checking the test every morning breaks it
Stopping a test the first morning it looks significant is not impatience, it is a method error that manufactures winners. Repeated testing of accumulating data inflates the false positive rate above the nominal level: Armitage, McPherson and Rowe put it plainly in 1969 — “if significance tests at a fixed level are repeated at stages during the accumulation of data the probability of obtaining a significant result when the null hypothesis is true rises above the nominal significance level” (Journal of the Royal Statistical Society Series A, 132(2), 235–244).
Practically: set the sample size before the first send, write it down, and do not look at the split until both arms reach it. Watch deliverability and opt-outs daily — those are safety checks, and a variant drawing complaints should be killed on sight. Do not watch the metric you are testing.
Our own published split test on whether a “reply STOP to opt out” line hurts booking rate is a worked example of what an honest read looks like: on roughly 40 sends per arm the opt-out gap was conclusive at p = 0.009, and the booking gap was not, at p = 0.085 — so we published it as a strong early signal rather than a finding. The full numbers are in our reply-STOP opt-out split test.
What does running this properly cost in hours, tooling and skill?
The method above is complete and you can run it yourself. Here is the honest bill, because the obstacle is not the statistics.
- Tooling: a sending platform that randomises at the contact level and keeps the assignment stable across the whole sequence. Most sequencers randomise per message, which contaminates both arms by touch three.
- Data: reply and booking events written back against the variant ID, deduplicated to unique contacts. This is the part that breaks, and it breaks silently.
- Hours: roughly a day to set up the first test properly, then an hour a week to keep the list and the exclusions clean. Not large — but it is a recurring hour, forever.
- Skill: one person who will hold the stopping rule when the number looks good on day four.
- Volume: the binding constraint. Nothing above substitutes for it.
That last line is the whole page. A team sending 250 a week and a team sending 5,000 a week can run identical method with identical discipline, and only one of them learns anything this quarter. So the first thing to fix is how much outreach there is — which is what an AI outbound sales system changes, and why AI appointment setting puts testing within reach of teams it was closed to. To prove the mechanism before scaling it, see how to run an AI outbound pilot that can fail.
Frequently asked questions
How much volume do I need to A/B test cold email?
At a 5% reply rate, 1,467 sends per variant to detect a 50% lift and 8,146 per variant to detect a 20% lift, at 95% confidence and 80% power. The formula is n = (Zα/2 + Zβ)² × [p₁(1−p₁) + p₂(1−p₂)] ÷ (p₁−p₂)², set out with the standard Z values of 1.96 and 0.84 in Hazra and Gogtay, “Biostatistics Series Module 5: Determining Sample Size”, Indian Journal of Dermatology (2016).
Can I split test if I only send 100 emails a week?
Not meaningfully. 100 a week takes 29 weeks to read a 50% lift and three years to read a 20% one. Run sequential changes instead: change one large thing, run it for a fixed block, and compare blocks knowing that the comparison is directional only. Say so out loud when you report it.
Is a 95% confidence level necessary for sales message testing?
No, and lowering it is a legitimate trade. Dropping to 90% confidence and 70% power cuts the required sample by about 40% (the leading term falls from 7.84 to 4.71), at the cost of more false winners. That is often the right call for a cheap, reversible change and the wrong call for one you will build a quarter around. State the threshold before the test, not after.
Should I test on reply rate or booked calls?
Test on replies, confirm on bookings. Booking rates run well below reply rates and lower rates need much larger samples — at a 1.5% booking rate, detecting a 50% lift needs 5,125 sends per arm against 1,467 on a 5% reply rate. Reply rate is the readable proxy; booked calls are the number that pays.
What if one variant wins on replies but loses on bookings?
Believe the bookings and re-run. It usually means the winning message widened the top of the funnel with people who were never going to buy — more replies, worse replies. Check opt-out rate at the same time, because a message can lift replies and lift exits together.
How many variants can I test at once?
Two, at the volumes most teams have. Each extra arm splits your volume again and adds a multiple-comparison problem: with four arms you need roughly double the total sends of a two-arm test and a correction to the threshold. Sequential two-arm tests beat one underpowered four-arm test.
Pay-Per-Result appointments
See if we’re a fit
We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.
- 50,769+ appointments booked without cold calling.
- Pay-Per-Result pricing — you pay for booked, qualified calls.
- Pick your own time on our live calendar, no phone tag.
