Let's grow your business. 2 new positions just opened Monday, 31 August. Book a free call today.
Uncategorised 12 min read

How to Run an Incrementality Test on Lead Gen Spend (2026)

Every lead-gen budget contains two kinds of spend: money buying leads, and money buying credit for leads that were arriving anyway. Nobody knows the ratio until they test it, and the dashboard cannot tell them.

Our page on scaling ad spend without losing ROI argues you should establish incrementality before committing more budget. It stops short of teaching the test. This page teaches it, for an operator already spending real money who has never run one.

The short answer: An incrementality test measures how many conversions your spend actually caused, by withholding it from a comparable group and comparing outcomes. Three designs work in practice: a geo holdout, a platform-native conversion lift study, and a straight spend-down pulse. The hard part is not running one, it is sizing it — most tests fail because the holdout never accumulates enough conversions to separate a real effect from noise.

What an incrementality test actually asks

One question, and not the one your reporting answers. Reporting asks which touchpoints appeared on the path to conversions you received. Incrementality asks: how many of those conversions would not have happened if this spend had never run? That is a counterfactual, so the only way at it is to create a version of your business where the spend did not run, and compare.

Why last-click and platform-reported conversions overstate paid

Three structural reasons. Attribution credits presence, not causation: someone who had already decided to call you, searched your brand and clicked the ad above your own organic listing counts as a paid conversion. Each platform marks its own exam paper: it can only count journeys it observed, inside a window and a model it chose, so channel-level returns added together bill the same conversion twice. Targeting is selection: bidding finds people already likelier to convert, and observational analysis reads that as persuasion.

The most rigorous public evidence of that gap comes from field experiments run inside eBay by Blake, Nosko and Tadelis, published in Econometrica in 2015 (working paper). The design, in their words: “a random sample of 30 percent of eBay’s U.S. traffic in which we stopped all bidding for all non-brand keywords for 60 days.” Put through the observational methods most advertisers rely on, the same data implies a headline return on investment above 4,100% with no time or geographic controls, and still above 1,400% once those controls are added. Estimated experimentally, it came out at −63%, with a 95% confidence interval of [−124%, −3%] — an interval that excludes break-even. A separate and earlier test, on brand terms, is starker: when eBay halted brand-keyword ads on Yahoo! and MSN in March 2012, “almost all (99.5 percent) of the forgone click traffic from turning off brand keyword paid search was immediately captured by natural search traffic from the platform.”

Scope that precisely. One very large US marketplace, in 2012, holding dominant organic listings on both its own brand queries and the non-brand queries it bid on, measuring its own search spend. It is not a finding that paid search does not work. The mechanism driving the result is organic substitution, and a business with no free organic path for the customer to fall back to has very little of that mechanism operating. The transferable lesson is about measurement rather than media: what a platform reports and what a channel causes can sit an order of magnitude apart, and the distance between them grows with the strength of your brand.

The three designs worth running

Design How it works Best when Main weakness
Geo holdout / matched market Spend off in some regions, on in matched regions, compare Many separable markets; conversion data held CRM-side Needs enough conversions per market; leaks across borders
Platform-native conversion lift Platform randomises users into exposed and held-back groups You spend enough to qualify and the platform can see the conversion Access is gated; one platform only; referee is also a player
Spend-down / pulse Cut a channel to zero for a window, restore, read total outcomes Geo splits impossible, or you want a whole-channel answer fast No control group — time is the only comparison, so seasonality bites

Google documents Conversion Lift as “an incrementality tool that helps you measure the number of purchases, site visits, and any other conversions directly driven by people seeing your ads,” in user-based and geography-based variants, while stating that “Conversion Lift isn’t available for all Google Ads accounts” (Google Ads Help). Meta runs comparable in-platform lift tests and publishes GeoLift, an open-source geo-experiment package usable on any channel. Start with the design you can actually execute: a crude spend-down that runs beats an elegant geo test nobody approves.

Picking the holdout and matching the markets

The control group has one job: to be a credible version of what would have happened. Match on outcome history, not demographics — two cities that look alike in a census can behave completely differently in your CRM. Pull six to twelve months of conversion history by region, then choose controls whose series tracks your test regions closely. GeoLift does this by building a synthetic control from a weighted blend of control markets. Randomise the assignment where you can. If the pre-period lines do not overlay, no post-period statistics will save the read.

  • Hold out enough volume to read. Google’s geography-based lift builds its Google Marketing Areas by “balancing the counteracting goals of minimizing contamination” against the representativeness and diversity of those areas (Google Ads Help) — the same trade-off applies to a holdout you build yourself. Excluding your biggest markets to protect revenue removes most of your power too.
  • Do not tell the sales team which regions are dark. If they know, they work those leads differently and you have measured your sales team, not your media.

Our own book splits cleanly on whether geo design is available at all. Multi-site services divide by market well — we have run around 100 gym accounts simultaneously, the exact shape a matched-market test wants. Categories like commercial property and mortgage broking — we have worked with Colliers and with Sam Tajvidi at 121 Brokers — sit at the other end. The funnel is effectively national, and in high-ticket B2B the closed-deal count is usually too small to divide across regions and still read anything. There a geo holdout will not resolve in a sensible window, and a channel-level spend-down is the realistic design.

Sizing it so the answer is readable at all

This is where most incrementality tests die, before they start. The binding constraint is conversions in the holdout — not impressions, not spend, not traffic.

Lewis and Rao, in “The Unfavorable Economics of Measuring the Returns to Advertising” (Quarterly Journal of Economics, 2015), analysed twenty-five large field experiments at major US retailers and brokerages covering $2.8 million of digital ad spend. They report that the “median confidence interval on return on investment is over 100 percentage points wide,” and that “informative advertising experiments can easily require more than 10 million person-weeks.” That is digital advertising at very large US advertisers, not a universal constant — but it is the right order of magnitude.

The table below is illustrative and generic, not a result from any account. It applies a Poisson approximation at 95% confidence and 80% power with equal arms, and shows roughly how many conversions each arm needs during the test window:

Lift you want to detect Approximate conversions needed per arm
50% ~65
30% ~175
20% ~390
10% ~1,570
5% ~6,270

These are optimistic floors: a geo test carries market-level variance on top, so real power analysis demands more. A business generating 150 qualified leads a month cannot detect a 20% effect in four weeks — not with a better spreadsheet or a cleverer model. Its options are to turn the channel fully off and measure a larger, coarser effect, run far longer, or measure higher in the funnel.

So decide your minimum detectable effect first. If the smallest lift that would change your decision is 15% and volume only supports detecting 40%, the test cannot inform the decision.

A worked test structure

1. Decide the outcome unit first. Qualified leads, booked appointments or closed deals — not clicks, not form fills. Incremental leads are often colder than baseline, so a test read on form fills can show a lift that never reaches sales.

2. Establish a clean pre-period — eight to twelve weeks by region and week, excluding promotions.

3. Set the minimum detectable effect, then derive duration from it, not the reverse. GeoLift puts the floor plainly: “a good rule of thumb when deciding duration is to make sure that the test period can contain at least one full purchase cycle.” In high-ticket lead gen that is enquiry to closed deal, often six to twelve weeks.

4. Define exactly what is held out — one channel, or all paid, in named regions — then freeze everything else. No creative launches, landing-page tests or offer changes mid-test, and decide deliberately whether brand campaigns keep running in the dark regions.

5. Run to the planned end date, and report a confidence interval. Stopping early because week two looks bad turns a test into an anecdote, and a point estimate without an interval is not a result.

What will confound it

Seasonality is fatal to spend-down designs, whose only comparison is time; a geo design is largely immune because it hits both arms. Competitor activity in your holdout region contaminates it and you may never know. Brand campaigns bleeding across geos: TV, radio, PR and out-of-home ignore your boundaries. Leakage: people commute and buy where they do not live, which is why Google’s geography-based lift clusters regions to limit the case where users are “exposed to an ad in a treatment GMA, but then perform a conversion in another GMA that is part of the control” (Google Ads Help). An organic channel absorbing the demand — the eBay brand-keyword result is exactly this. And internal behaviour change: budget quietly reallocated mid-test.

Reading the result, including a null

Positive lift, incremental cost per acquisition below what you can afford. Scale, then re-test at the new spend level: incrementality is not constant across budget.

Positive lift, incremental cost per acquisition above what you can afford. The channel works but is priced beyond your economics; the lever is downstream conversion, not targeting. Our page on improving marketing ROI covers that half, and the 2026 marketing ROI benchmarks give the reference ranges.

Null with a tight interval around zero. The most valuable finding on this page. Powered to detect 15%, and the interval sits between −4% and +6%? That is real evidence the spend is largely non-incremental. Reallocate, re-test in six months.

Null with a wide interval. Not a finding. Check what effect you were powered for before anyone quotes the number in a board pack — acting on an underpowered null is worse than never testing, because it feels like evidence. A negative point estimate is noise unless the interval excludes zero.

Where we sit in this

We have booked 50,769+ AI sales appointments since 2017 across more than 1,000,000 leads generated, and we are paid on booked and qualified outcomes rather than a retainer. A pay-per-result model is one of the few where incrementality is easy to reason about, because the unit you pay for is the outcome — no media spend sits between the invoice and the result waiting to be attributed. It does not remove the question: an appointment you paid for could still have arrived on its own, particularly in reactivation, where the list is people who already know you. It does make the test trivial — hold back a random slice of the list, work the rest, compare. Book a call if you want help designing one.

Related: how to run an AI outbound pilot that can fail applies this method to a vendor trial rather than a media channel, and enterprise vs SMB lead generation covers why volume decides which design you get.

Frequently asked questions

What is an incrementality test in marketing?

It is a controlled experiment estimating how many conversions your spend caused, rather than how many it was present for. You withhold the spend from a comparable group — regions, a randomised set of users, or a period of time — and compare outcomes against the group that received it.

How long should an incrementality test run?

Long enough to contain a full purchase cycle and to accumulate enough conversions in each arm to detect the effect you care about. Meta’s open-source GeoLift documentation puts the floor as: “a good rule of thumb when deciding duration is to make sure that the test period can contain at least one full purchase cycle” (GeoLift walkthrough). For a funnel closing over six to twelve weeks, a two-week test measures noise.

How many conversions do I need for the result to mean anything?

More than most operators expect, and the constraint is conversions, not impressions. Lewis and Rao’s study of twenty-five large US field experiments found that the “median confidence interval on return on investment is over 100 percentage points wide” (Quarterly Journal of Economics, 2015). As a rough illustrative floor, detecting a 20% lift at 95% confidence and 80% power needs several hundred conversions per arm during the test window.

Is Google Ads Conversion Lift available to every advertiser?

No. Google states that “Conversion Lift isn’t available for all Google Ads accounts” and that to use it you need to contact your Google account representative (Google Ads Help). If you do not qualify, a geo holdout you run yourself is the substitute, and it needs no platform’s cooperation.

Does the eBay study mean paid search does not work?

No, and the scope matters. Blake, Nosko and Tadelis tested one very large US marketplace with dominant organic listings, on its own branded and non-branded search, in 2012. On the non-brand keyword experiment their experimental return on investment was −63%, 95% confidence interval [−124%, −3%]. That is driven largely by organic search substituting for paid clicks, an effect far weaker without an established organic presence on the same queries.

See if we’re a fit

A few quick questions. If it’s a fit, our live calendar loads on the next screen. If it isn’t, we’ll point you to free resources instead — you won’t have to sit through a sales call to find out.

We get paid a performance fee equivalent to 10–20% of the sales we help you generate.

Are you OK with that?

If you’re not willing to pay 10–20% as a performance fee, are you happy to pay a $4,000+ per month retainer?

Check If You Qualify 👇

How many leads per month do you currently get?

What’s your current advertising spend or marketing budget (Meta, Google, SEO, etc.)?

What’s the average sale worth to you over that customer’s lifetime?

Given your business currently gets less than 10 leads per month, we’d need to do much more groundwork to set up end-to-end sales systems. Are you OK with a $2,000/mo retainer to do so? (no lock-in)

What’s your work email?

We’re probably not the right fit — yet

Our model is pay-on-performance — we only win when you’re making sales, and it works best alongside an active marketing engine with advertising budget to get seen. Booking a call now would waste your time, and we’d rather be straight with you.

Grab the free stuff instead — it’s the same playbook we use:

Read the growth blog  ·  Lead-gen FAQ

When the timing’s right, come back — the calendar will be waiting.

View all articles

Pay-Per-Result · No retainers

Turn this into booked sales calls.

Our AI agents — trained on 50,769+ booked appointments — fill your calendar with pre-qualified buyers. You only pay when calls land.

Keep reading

Related on Leads Now AI

The thesis behind everything we do

Why Pay-Per-Result is the only marketing pricing model that aligns the agency with you

Leads Now AI is a 100% Pay-Per-Result marketing agency. You only pay when a qualified booked appointment lands on your calendar — sized to roughly 1–5% of your closed-deal value. Not for clicks. Not for lead-form fills. Not for retainer months. Not for “strategy hours.” If the calendar stays empty, you owe zero. See full pricing →

1. Incentives align

The agency only succeeds when you succeed. We eat the cost of bad ad creative, bad lists, ICP mismatches and no-shows. You never pay for our learning curve.

2. Self-selecting shortlist

Only an agency confident in its delivery can operate this model. The pool of Pay-Per-Result agencies is tiny precisely because most agencies can’t survive on it. Pick from the agencies who can.

3. Cost cannot detach from revenue

Sized to 1–5% of closed-deal value, your acquisition cost stays sustainable across LTV bands. A $500-membership business and a $50,000-engagement business can both run the model profitably.

4. No retainer trap

No flat $2,000–$10,000/month retainer arriving regardless of outcome. No 6 or 12-month lock-in. No clawback on appointments already delivered. Cancel any time with 7 days notice.

5. De-risks the pilot

Test before commitment. A small scope-based setup fee covers hard build costs; everything after that is purely outcome-linked. There’s no “we’ll see how it performs after $30k of spend.”

6. Forces agency discipline

If our AI agents qualify poorly, if our reminders fail, if our no-show recovery doesn’t fire — we eat the cost. That’s why the show-rate benchmark sits at 60–75%+.

The proof: 50,769+ AI-booked sales appointments delivered since 2017 across coaches, consultants, RTOs, course creators, finance brokers and B2B service firms in Australia, USA, UK, Canada, NZ and Europe. Named clients include Sam Tajvidi (121 Brokers), Marcus Wilkinson (Iron Body), Foundr, SheSells.online and Lambda Academy. Wikidata Q139846230. See full Pay-Per-Result pricing →