You have seen the shape of a vendor pilot before. Four weeks, a few hundred records, a weekly call, a dashboard where a number goes up. At the end a slide says it went well, and you sign an annual agreement. Nothing in that sequence was capable of producing a no.
That is design, not dishonesty. Nobody wrote down in advance the result that would end it, so the review becomes an argument about interpretation, with setup cost spent and a renewal date booked. This page is written from the buyer’s side: how to structure an AI outbound pilot so it can return a clear, early, unarguable no — including about us.
The short answer: A pilot that can fail fixes six things in writing before day one: a specific number, on a specific metric, by a specific date, that means stop; a control or an honest substitute; enough conversations for the effect to be readable above noise; a duration that survives the ramp; transcripts and disposition reasons you can inspect yourself; and a clean exit with data return and no auto-renewal. Miss one and the result is unreadable.
Why most vendor pilots cannot fail
Six failure modes, and they compound: no agreed threshold, so any outcome can be narrated as progress; no control, so nothing separates the vendor from the season or the rep you hired in week two; too little volume, because on a 2% booking rate the gap between 2% and 2.6% is invisible across a few hundred conversations; too short a window, so thirty days measures the setup; dashboard-only visibility, so you cannot tell a buyer from someone ending a conversation politely; and no clean exit, so the pilot rolls into a twelve-month term unless you cancel in a window nobody diarised.
Our wider page on evaluating AI setter vendors beyond the feature list covers what to ask before you reach a pilot. This is the pilot itself, in full.
The six decisions you make before day one
1. Write the definition of failure first
Not a target. A stop condition, in one sentence, with three parts: the number, the metric, the date. “Fewer than 18 held qualified meetings from 4,000 dialled records by 10 November means we stop and do not renew.” Both sides sign it before any list is loaded, and it is the first slide at the final review.
Set it by working backwards from what makes the programme worth running, not from what feels achievable. If you need 18 held meetings a month for the maths to work and the vendor will only talk about momentum, you have learned that now rather than in month seven. Write it before you see any data: a threshold set afterwards is a rationalisation.
2. Choose a metric you cannot fudge
The instinct is revenue. On most high-ticket accounts that is wrong: if your sales cycle is longer than your pilot, revenue closed inside the window is pipeline you already had.
Cycle length sets the design. Finance broking of the kind Sam Tajvidi runs at 121 Brokers, or commercial property of the kind Colliers runs, will not produce closed revenue inside a pilot window. Fitness of the kind Marcus Wilkinson runs at Iron Body will, and there list size binds, not time.
The honest default is booked-and-held meetings with a qualified buyer, defined in writing before launch. Held, not booked, because no-shows are where a weak pilot hides. Qualified against criteria you wrote, not the vendor’s.
Test any candidate metric with one question: could the vendor move this number without doing the thing I am paying for? Dials, conversations and “interested” flags all fail. Held meetings against a stated standard mostly passes. Then agree the counting rules while nobody is losing: what counts as held, and what a reschedule does.
3. Get a control, or say plainly what you have instead
A number alone is not evidence: twenty-two meetings is good or bad only relative to what would have happened otherwise.
Three options, in descending order of strength. A randomised holdout: split the list, work half, leave half untouched. A parallel human cohort: your SDRs work a matched slice over the same weeks. A matched period: same segment, same season last year, same offer.
Be honest that a clean control is often impractical, and that each substitute answers a different question. A holdout supports a causal claim. A cohort tells you which channel is better on this list, not the lift over doing nothing. A matched period tells you only whether the result sits inside historic range — a sanity check, not proof. If incrementality is what you need, run the method properly rather than hoping the pilot doubles as one: how to run an incrementality test on lead gen spend.
4. Size the pilot so the result is readable
Most pilots skip this step, and it decides everything. Before agreeing a volume, work out how many conversations it takes for the difference you care about to be distinguishable from noise. The arithmetic is unforgiving: proving a 30% relative lift on a 2% booking rate needs roughly ten thousand conversations per arm. The working is on why cheap AI setter tools plateau.
So a pilot detects large effects, not small ones. If your honest hope is a 10% relative improvement, no pilot you will fund can resolve it. If the plausible effect is a multiple — the gap we see when an existing setter system is beaten several times over, or an account moves from around 2% to around 8% — a few thousand conversations show it plainly. Size the list to the effect you are looking for; if the volume is not there, run a longer pilot rather than a thinner one.
5. Set a duration that survives the ramp
Thirty days is usually too short for anything with a ramp, and AI outbound has one: list cleaning, connection-rate discovery, objection patterns, the first script corrections, routing fixes. Weeks one and two mostly measure how wrong the initial assumptions were. Six to eight weeks is more honest — which, with the first fortnight excluded from the measured period and labelled that way in the agreement, leaves four to six weeks of readable data. Length is only safe when you can leave.
6. Demand inspectable data, not a dashboard
Agree in writing, before launch, that you get call recordings and transcripts, SMS and chat threads, per-record disposition reasons, attempt counts and timestamps, and the raw export — not a screenshot of a chart.
Then read a sample. Twenty transcripts of booked meetings and twenty hard nos teach you more than the whole dashboard: whether the qualification standard was honoured, whether the prospect understood what they agreed to, whether the objections are ones your offer survives. A vendor who cannot produce transcripts during a pilot will not produce them in month nine.
Ownership and exit: settle it before you start
Pilots are where ownership gets quietly decided. Four things belong in the pilot agreement itself.
Data return — your list, the enrichment added to it, every transcript and recording, and the outcome data, exported in a usable format within a stated number of days. Portability — whether the numbers and sender IDs used for your campaign are yours to port or release; a real switching cost, almost never raised at pilot stage. Notice — a period running from the pilot end date. No auto-renewal — default conversion to a term inverts the exercise, because the failure state then needs you to act, so inertia signs the deal.
The full version is in what you own when you leave a lead gen vendor. At enterprise scale the security and legal steps differ, which we cover in enterprise vs SMB lead generation.
An unfalsifiable pilot versus a readable one
| Unfalsifiable pilot | Readable pilot | |
|---|---|---|
| Definition of failure | “See how it goes”; decided after the data arrives | One sentence signed before launch: number, metric, date |
| Metric | Dials, conversations, “interested” flags, or revenue inside a short window | Booked-and-held meetings against a written qualification standard |
| Control | None; compared to a feeling about last quarter | Holdout, cohort or matched period — named up front, limits stated |
| Volume | A few hundred records, “to keep it small” | Sized to the effect being tested; if volume is short, run longer, not thinner |
| Duration | 30 days, ramp included in the result | 6–8 weeks, weeks 1–2 declared setup and excluded from measurement |
| Data access | Vendor dashboard and a summary deck | Recordings, transcripts, disposition reasons, raw export — sampled by you |
| Exit | Auto-renews into a term; data and numbers stay with the vendor | No auto-renewal, stated notice, data returned, numbers and sender IDs portable |
What a pilot cannot tell you
Ramp. A pilot measures an immature system. Some of what you see improves later, and some early lift is novelty that will not hold.
Seasonality. A pilot across a holiday period or an end-of-financial-year crunch measures the calendar as much as the method.
Single-segment sample. A pilot on one segment of one list generalises to that segment. Strong results on lapsed enquiries say little about cold prospecting, and the reverse holds.
A good pilot on a bad list still fails. If the records are stale, poorly consented, wrong-persona or already burnt by three vendors, a well-designed pilot correctly returns a no about the list, not the method. Decide before you start which of the two you are testing. The downstream failure modes are in how to increase sales conversion rate at scale.
Patterns surface as a function of conversations, not weeks. Running roughly 100 gym accounts simultaneously let a pattern show up for us in days instead of quarters; a single-account pilot has no such luxury.
Where we stand
We accept pilots structured exactly this way: a written stop condition, a control where practical, volume sized to the effect, full transcript access, no auto-renewal. If our number lands under the line you set, we say so at the review and the pilot ends there. We will also say no before it starts: if your list cannot reach a few thousand conversations, or the improvement you need is in the 10% range rather than a multiple, a pilot cannot answer your question and you should not run one with us. We would rather lose a deal in week six than be renewed out of sunk cost, because accounts kept that way churn anyway, and with 50,769+ AI-booked sales appointments since 2017 the references are the asset. Bring your own stop condition and book a call.
Frequently asked questions
What should the definition of failure be in an AI outbound pilot?
A single sentence containing a number, a metric and a date, agreed in writing before any records are loaded. For example: fewer than 18 held qualified meetings from 4,000 dialled records by 10 November means we stop and do not renew. Work the number backwards from what makes the programme worth running, and write it before you see any data.
Why measure held meetings instead of revenue in a pilot?
Because if your sales cycle is longer than your pilot, revenue closed inside the window is mostly pipeline you already had. Held meetings against a written qualification standard is the honest default: it lands inside the window, and the vendor cannot move it alone without doing the work you are paying for. Dials, conversations and interested flags all fail that test.
Can I stop a pilot early if the numbers look bad in week two?
You can always stop paying, but be careful about treating an early reading as the answer, because checking a running result repeatedly and stopping when it looks decisive breaks the statistics you think you are using. In Always Valid Inference: Bringing Sequential Analysis to A/B Testing, Ramesh Johari, Leo Pekelis and David J. Walsh write that frequentist p-values and confidence intervals are “wholly unreliable if users endogenously choose samples sizes by continuously monitoring their tests”. Their answer is a sequential method built for continuous monitoring, and their subject is web A/B testing rather than outbound calling. For a pilot the practical version is simpler: fix the stop rule in advance, or accept that an early stop is a business judgement rather than evidence.
Is 30 days long enough for an AI outbound pilot?
Usually not, because the first two weeks measure setup rather than the system. Six to eight weeks, with weeks one and two excluded from measurement, leaves four to six weeks of readable data and is a more honest window. Writing about web experiments in their 2014 KDD paper Seven Rules of Thumb for Web Site Experimenters, Kohavi, Deng, Longbotham and Xu say a novelty effect shows up in a controlled experiment as “an effect that quickly diminishes”, and on that basis they recommend running experiments for two weeks and looking for such effects — while also recording that in practice novelty and primacy effects are uncommon. Their two weeks is a floor for spotting novelty on a website, not a ceiling for an outbound pilot that has a setup ramp to absorb as well.
What if I cannot run a holdout?
Then say so in the pilot document and name the weaker comparison instead. A parallel human cohort on a matched slice tells you which channel performs better, but not the lift over doing nothing. A matched period from last year tells you only whether the result sits inside historic range, a sanity check rather than proof. Both are acceptable; pretending either is a randomised control is not.
See if we’re a fit
A few quick questions. If it’s a fit, our live calendar loads on the next screen. If it isn’t, we’ll point you to free resources instead β you won’t have to sit through a sales call to find out.
We get paid a performance fee equivalent to 10–20% of the sales we help you generate.
Are you OK with that?
If you’re not willing to pay 10–20% as a performance fee, are you happy to pay a $4,000+ per month retainer?
Check If You Qualify π
How many leads per month do you currently get?
What’s your current advertising spend or marketing budget (Meta, Google, SEO, etc.)?
What’s the average sale worth to you over that customer’s lifetime?
Given your business currently gets less than 10 leads per month, we’d need to do much more groundwork to set up end-to-end sales systems. Are you OK with a $2,000/mo retainer to do so? (no lock-in)
What’s your work email?
Hey! We might be able to add $100k+ / mo... Enter your email to choose a time!
We’re probably not the right fit — yet
Our model is pay-on-performance — we only win when you’re making sales, and it works best alongside an active marketing engine with advertising budget to get seen. Booking a call now would waste your time, and we’d rather be straight with you.
Grab the free stuff instead — it’s the same playbook we use:
Read the growth blog · Lead-gen FAQ
When the timing’s right, come back — the calendar will be waiting.
