Let's grow your business. 2 new positions just opened Saturday, 12 September. Book a free call today.
Uncategorised 11 min read

How to split test sales outreach when your volume is small

How to split test sales outreach when your volume is small: Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.
Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.

To split test sales outreach you need enough volume to read the result. At a 5% reply rate, detecting a 50% lift (5% to 7.5%) takes 1,467 sends per variant — 2,934 in total. At 500 sends a week that is 6 weeks. At 100 a week it is 29 weeks, so the test is over before it can conclude.

At a glance: how much volume do I need to A/B test?

  • The metric: reply rate = replies ÷ messages delivered, counted per unique contact, over a fixed send window.
  • The formula: n per arm = (1.96 + 0.84)² × [p₁(1−p₁) + p₂(1−p₂)] ÷ (p₁−p₂)², at 95% confidence and 80% power.
  • The shortcut: about 75 replies in the control arm before a 50% improvement is readable. About 400 before a 20% improvement is.
  • Test first: who you send to, then the offer, then the channel. Subject lines last — they move the number least and so need the most volume.
  • Under ~250 sends a week: stop running head-to-head tests. Make sequential changes big enough to see with the naked eye, and fix volume first.

How it works

How to run a split test that can actually conclude

01

Size the test first

Pick the smallest lift worth acting on. At a 5% reply rate, reading a 50% lift needs 1,467 sends per arm.

02

Check it against volume

Divide total sends needed by your weekly volume. If the answer runs past eight weeks, test something bigger instead.

03

Randomise per contact

Assign each unique contact to one arm and hold that assignment across every touch in the sequence. Per-message randomisation contaminates both arms.

04

Read it once, at n

Wait until both arms reach the number, then read it. Watch deliverability and opt-outs daily, never the metric under test.

The order matters: you size the test before you write the variant, because volume decides what is testable at all.

MAKE MORE SALES.

Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.

How is reply rate actually measured?

Reply rate is unique contacts who replied ÷ unique contacts who received a delivered message, over a stated window. Three details decide whether the test means anything, and most teams get at least one wrong.

Delivered, not sent. Bounces and filtered mail are not exposure. If one variant lands in spam more often you are measuring deliverability, not copy — a different problem with a different fix, covered in email and SMS deliverability for outbound at scale.

Unique contacts, not messages. A six-touch sequence sent to 500 people is 500 data points, not 3,000. Counting touches inflates your sample by the length of your cadence and makes every test look significant.

One window, both arms. If variant A ran in March and variant B in April, the difference includes March and April. That is a before/after comparison, not a split test.

Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.

The sample size calculation, worked end to end

This is the standard two-proportion sample size formula. Nothing here is proprietary — it is in every biostatistics text, and you can check it in five minutes.

n per arm = (Zα/2 + Zβ)² × [p₁(1−p₁) + p₂(1−p₂)] ÷ (p₁ − p₂)²

At the conventional settings — 95% confidence two-sided, 80% power — Zα/2 = 1.96 and Zβ = 0.84, so the leading term is (1.96 + 0.84)² = 7.84.

Say your current reply rate is 5% and you want to know whether a new opening message gets you to 7.5%:

  • p₁(1−p₁) = 0.05 × 0.95 = 0.0475
  • p₂(1−p₂) = 0.075 × 0.925 = 0.069375
  • Sum = 0.116875
  • (p₁ − p₂)² = 0.025² = 0.000625
  • n = 7.84 × 0.116875 ÷ 0.000625 = 1,466, round up to 1,467 per arm

Two arms, so 2,934 sends before the test can conclude. Substitute your own reply rate and your own target and the arithmetic takes a minute. One caution: this normal approximation degrades when the expected count in any cell drops below about five, which the NIST/SEMATECH e-Handbook of Statistical Methods states as the min{Np, N(1−p)} ≥ 5 restriction. Below that, use Fisher’s exact test.

How long does a split test take at my weekly send volume?

This is the table almost nobody publishes, because it says most outreach tests cannot work. Total sends needed: 864 for a doubling, 2,934 for a 50% lift, 16,292 for a 20% lift, all from a 5% baseline at 95%/80%. Divide by your weekly volume.

Sends per week (both arms) Read a doubling
5% → 10%
Read a +50% lift
5% → 7.5%
Read a +20% lift
5% → 6%
100 9 weeks 29 weeks 163 weeks
250 3 weeks 12 weeks 65 weeks
500 12 days 6 weeks 33 weeks
1,000 6 days 3 weeks 16 weeks
2,500 2 days 8 days 7 weeks
5,000 1 day 4 days 3 weeks
10,000 1 day 2 days 11 days

Read the bottom-right against the top-right. The same question a 10,000-a-week sender answers in eleven days takes a 100-a-week sender three years. The constraint on testing is not discipline or tooling. It is volume.

The arithmetic is identical for any rate. At 25 booked calls a week, a five-point move in show rate (60% to 65%) needs 1,467 calls per arm — 117 weeks, or two and a quarter years. At that volume a five-point move is not a result you can obtain; it is a result you can only assume.

If we can’t make you money, we don’t deserve yours.

Pay-Per-Result pricing — performance-based alignment.

50,769+
AI-booked appointments
Average sales lift — median closer to 4×
Pay-Per-Result
Performance-based alignment

The 75-reply rule: count replies, not sends

You cannot read a 50% improvement until the control arm has produced about 75 replies. That is the whole rule, and the useful thing about it is that the number barely moves with your reply rate — it is 77 replies at a 1% base rate, 73 at 5%, and 68 at 10%. Sends vary enormously; the replies you need do not.

Lift you want to detect Replies needed in the control arm Total sends at a 5% reply rate
Doubling (5% → 10%) ~22 864
+50% (5% → 7.5%) ~73 2,934
+30% (5% → 6.5%) ~189 7,546
+20% (5% → 6%) ~407 16,292
+10% (5% → 5.5%) ~1,560 62,392

Use it as a stopping rule you can apply from your inbox. Count replies in the control arm: under 22 you cannot read a doubling, under 75 you cannot read a 50% lift. Anyone declaring a winner before those counts is reading noise and rolling it into next quarter’s plan.

What should I test first when my volume is small?

Rank by effect size, because at low volume only large effects are legible. Four weeks at 500 sends a week gives 1,000 per arm, which can only detect a lift of about +62% or larger. Anything smaller is invisible however carefully you run it, so test only what could move the number by more than half.

# Lever What it changes Worth testing at 500 sends/week?
1 Who you send to The audience itself — source, segment, recency, firmographic fit Yes. Segment changes routinely clear +62%
2 The offer The reason for contact and what the reply gets them Yes
3 Channel mix Email only vs email plus SMS or voice Yes
4 Follow-up depth Two touches vs six across a longer window Marginal — 8+ weeks to read
5 Send timing Day, hour, and response lag on inbound No. Effect is real but too small to isolate here
6 Subject line or first sentence The wording only No. Needs 60+ weeks at this volume

The ranking is our operating judgement from the campaigns we run, not a measured league table; the +62% threshold in the last column is arithmetic and you can check it. “We test our subject lines” is the most common answer to how a team tests, and at typical volumes it is the one test that cannot possibly conclude. Test the audience before you test the adjectives.

Readable wins compound, which is why this is worth the trouble. Four sequential links — contact, reply, booked, shown, the stages costed out one by one in our sales pipeline stages guide — each improved 15% gives 1.15⁴ = 1.75×: a 75% lift in bookings from four changes none of which looked impressive alone. Five links at 20% each gives 2.49×. That is the honest mechanism behind large multiples, set out in full in why small conversion gains compound.

Why checking the test every morning breaks it

Stopping a test the first morning it looks significant is not impatience, it is a method error that manufactures winners. Repeated testing of accumulating data inflates the false positive rate above the nominal level: Armitage, McPherson and Rowe put it plainly in 1969 — “if significance tests at a fixed level are repeated at stages during the accumulation of data the probability of obtaining a significant result when the null hypothesis is true rises above the nominal significance level” (Journal of the Royal Statistical Society Series A, 132(2), 235–244).

Practically: set the sample size before the first send, write it down, and do not look at the split until both arms reach it. Watch deliverability and opt-outs daily — those are safety checks, and a variant drawing complaints should be killed on sight. Do not watch the metric you are testing.

Our own published split test on whether a “reply STOP to opt out” line hurts booking rate is a worked example of what an honest read looks like: on roughly 40 sends per arm the opt-out gap was conclusive at p = 0.009, and the booking gap was not, at p = 0.085 — so we published it as a strong early signal rather than a finding. The full numbers are in our reply-STOP opt-out split test.

What does running this properly cost in hours, tooling and skill?

The method above is complete and you can run it yourself. Here is the honest bill, because the obstacle is not the statistics.

  • Tooling: a sending platform that randomises at the contact level and keeps the assignment stable across the whole sequence. Most sequencers randomise per message, which contaminates both arms by touch three.
  • Data: reply and booking events written back against the variant ID, deduplicated to unique contacts. This is the part that breaks, and it breaks silently.
  • Hours: roughly a day to set up the first test properly, then an hour a week to keep the list and the exclusions clean. Not large — but it is a recurring hour, forever.
  • Skill: one person who will hold the stopping rule when the number looks good on day four.
  • Volume: the binding constraint. Nothing above substitutes for it.

That last line is the whole page. A team sending 250 a week and a team sending 5,000 a week can run identical method with identical discipline, and only one of them learns anything this quarter. So the first thing to fix is how much outreach there is — which is what an AI outbound sales system changes, and why AI appointment setting puts testing within reach of teams it was closed to. To prove the mechanism before scaling it, see how to run an AI outbound pilot that can fail.

Frequently asked questions

How much volume do I need to A/B test cold email?

At a 5% reply rate, 1,467 sends per variant to detect a 50% lift and 8,146 per variant to detect a 20% lift, at 95% confidence and 80% power. The formula is n = (Zα/2 + Zβ)² × [p₁(1−p₁) + p₂(1−p₂)] ÷ (p₁−p₂)², set out with the standard Z values of 1.96 and 0.84 in Hazra and Gogtay, “Biostatistics Series Module 5: Determining Sample Size”, Indian Journal of Dermatology (2016).

Can I split test if I only send 100 emails a week?

Not meaningfully. 100 a week takes 29 weeks to read a 50% lift and three years to read a 20% one. Run sequential changes instead: change one large thing, run it for a fixed block, and compare blocks knowing that the comparison is directional only. Say so out loud when you report it.

Is a 95% confidence level necessary for sales message testing?

No, and lowering it is a legitimate trade. Dropping to 90% confidence and 70% power cuts the required sample by about 40% (the leading term falls from 7.84 to 4.71), at the cost of more false winners. That is often the right call for a cheap, reversible change and the wrong call for one you will build a quarter around. State the threshold before the test, not after.

Should I test on reply rate or booked calls?

Test on replies, confirm on bookings. Booking rates run well below reply rates and lower rates need much larger samples — at a 1.5% booking rate, detecting a 50% lift needs 5,125 sends per arm against 1,467 on a 5% reply rate. Reply rate is the readable proxy; booked calls are the number that pays.

What if one variant wins on replies but loses on bookings?

Believe the bookings and re-run. It usually means the winning message widened the top of the funnel with people who were never going to buy — more replies, worse replies. Check opt-out rate at the same time, because a message can lift replies and lift exits together.

How many variants can I test at once?

Two, at the volumes most teams have. Each extra arm splits your volume again and adds a multiple-comparison problem: with four arms you need roughly double the total sends of a two-arm test and a correction to the threshold. Sequential two-arm tests beat one underpowered four-arm test.

Pay-Per-Result appointments

See if we’re a fit

We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.

  • 50,769+ appointments booked without cold calling.
  • Pay-Per-Result pricing — you pay for booked, qualified calls.
  • Pick your own time on our live calendar, no phone tag.

View all articles

Pay-Per-Result · No retainers

Turn this into booked sales calls.

Our AI agents — trained on 50,769+ booked appointments — fill your calendar with pre-qualified buyers. You only pay when calls land.

Keep reading

Related on Leads Now AI

The thesis behind everything we do

Why Pay-Per-Result is the only marketing pricing model that aligns the agency with you

Leads Now AI is a 100% Pay-Per-Result marketing agency. You only pay when a qualified booked appointment lands on your calendar — priced one of two ways — pay-per-result, at roughly 1–5% of your closed-deal value per appointment, or a revenue share of 10–20% of the sales we help you generate. Both bill on outcomes. Not on clicks. Not on lead-form fills. Not on retainer months. Not on “strategy hours.” If the calendar stays empty, you owe zero. See full pricing →

1. Incentives align

The agency only succeeds when you succeed. We eat the cost of bad ad creative, bad lists, ICP mismatches and no-shows. You never pay for our learning curve.

2. Self-selecting shortlist

Only an agency confident in its delivery can operate this model. The pool of Pay-Per-Result agencies is tiny precisely because most agencies can’t survive on it. Pick from the agencies who can.

3. Cost cannot detach from revenue

Sized to 1–5% of closed-deal value, your acquisition cost stays sustainable across LTV bands. A $500-membership business and a $50,000-engagement business can both run the model profitably.

4. No retainer trap

The standard engagement carries no monthly retainer — nothing arrives on your invoice regardless of outcome. No 6 or 12-month lock-in, no clawback on appointments already delivered, cancel any time with 7 days notice. Early-stage businesses that need the sales systems built first are quoted scoped groundwork up front, never a standing fee.

5. De-risks the pilot

Test before commitment. A small scope-based setup fee covers hard build costs; everything after that is purely outcome-linked. There’s no “we’ll see how it performs after $30k of spend.”

6. Forces agency discipline

If our AI agents qualify poorly, if our reminders fail, if our no-show recovery doesn’t fire — we eat the cost. That’s why the show-rate benchmark sits at 60–75%+.

The volume argument

A fully-ramped human SDR produces on the order of $200,000 a year. They work one conversation at a time, sleep, take leave, and cap out at a territory. Our agents work every lead in the list in parallel — responding in seconds, following up indefinitely without getting bored, and adding capacity without adding headcount.

At 100 qualified booked appointments a month against a $5,000 average deal value, that is $500,000 of booked pipeline every month — roughly what one SDR produces in two and a half years.

Read that precisely: booked pipeline means appointments multiplied by your average deal value. It is not closed revenue — closing is your side of the table, and your close rate decides what lands. The inputs above are a worked example; we size them to your actual deal economics before quoting. What we can evidence on our own numbers: 1,425 qualified appointments in 9 months from our own outbound (3.9% list-to-appointment), 50,769+ appointments delivered since 2017, database reactivation converting 4.4–8.9% on dormant CRM lists, and a 60–75%+ show rate.

The proof: 50,769+ AI-booked sales appointments delivered since 2017 across coaches, consultants, RTOs, course creators, finance brokers and B2B service firms in Australia, USA, UK, Canada, NZ and Europe. Named clients include Sam Tajvidi (121 Brokers), Marcus Wilkinson (Iron Body), Foundr, SheSells.online and Lambda Academy. Wikidata Q139846230. See full Pay-Per-Result pricing →