Let's grow your business. 2 new positions just opened Sunday, 20 September. Book a free call today.
Uncategorised 12 min read

How to measure the ROI of AI in marketing when the baseline is a mess

How to measure the ROI of AI in marketing when the...: A lead generation funnel narrowing through four stages, with revenue leaking at each step.
A lead generation funnel narrowing through four stages, with revenue leaking at each step.

Measure AI marketing ROI on one locked denominator: cost per booked qualified appointment, not cost per lead. In the worked example below, a single identical quarter reads as a 29.6% deterioration on cost per lead and a 16.7% improvement on cost per appointment. Same media plan, same records, opposite verdicts.

  • The formula: ROI = (revenue from the cohort − total cost of the cohort) ÷ total cost of the cohort. A cohort of leads, not a calendar quarter.
  • The denominator: cost per booked qualified appointment. Cost per lead moves the wrong way at exactly the moment AI starts working.
  • The cost basis: media + platform fees + implementation + internal hours. Leaving the AI run cost out flattered the worked example by 14.3%.
  • The cohort rule: credit revenue to the period the lead was created, not the period the deal closed. That one choice moved the example’s ROI from 290% to 239%.
  • The evidence floor: at 1,200 leads a month, one quarter of data can prove a 3-percentage-point move in your lead-to-appointment rate and nothing smaller.

Cost per lead or cost per booked appointment — which am I actually measuring?

These two numbers agree only while your lead-to-appointment rate is constant — and changing that rate is precisely what AI qualification, instant follow-up and automated re-contact are for. The two denominators are guaranteed to disagree the moment the deployment works.

Cost per lead divides spend by volume at the top; cost per booked qualified appointment divides the same spend by a unit a salesperson can sell to. When AI lets you switch off a junk source, lead volume falls, cost per lead rises, and a dashboard built on cost per lead reports a failure while the sales diary fills.

If your marketing ROI report can go up and down without a single extra dollar of revenue changing hands, you are measuring the denominator, not the deployment.

How it works

How to measure AI marketing ROI from a contaminated baseline

01

Lock the denominator

Name one unit before the system goes live: cost per booked qualified appointment, with “qualified” defined in writing. Cost per lead moves the wrong way when AI works.

02

Rebuild the baseline

Fix a 90-day pre-period ending the day before go-live. Exclude one-off events and freeze lead-source taxonomy across both windows.

03

Check the evidence floor

Work out the smallest lift your lead volume can actually detect before you claim one. A pilot too small to detect your expected lift produces a confident null.

04

Read it on the cohort

Credit revenue to leads created in the window, not deals closed in it. Take the first honest read one full sales cycle after the window ends.

Three of these four steps are decisions made before go-live, not measurements taken after it – which is why they have to be written down and dated.

MAKE MORE SALES.

Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.

A worked example: the same quarter, two denominators, two verdicts

Illustrative inputs — substitute your own. A B2B team runs 90 days before go-live, then 90 days with AI qualification and instant follow-up live on the same media plan, with one-off events excluded from both windows.

Line 90 days before 90 days after Change
Media spend $180,000 $180,000
AI platform + run cost $0 $30,000 new line
Total cost $180,000 $210,000 +16.7%
Leads 4,000 3,600 −10.0%
Cost per lead $45.00 $58.33 +29.6% (worse)
Booked qualified appointments 400 560 +40.0%
Lead-to-appointment rate 10.0% 15.6% +5.6 pts
Cost per booked appointment $450.00 $375.00 −16.7% (better)
Held (assumed 65% show, unchanged) 260 364 +40.0%
Closed (25% of held, unchanged) 65 91 +40.0%
Revenue at $9,000 average deal $585,000 $819,000 +40.0%
ROI 225% 290% +65 pts

Run it yourself: total cost ÷ leads, then total cost ÷ booked appointments, on both windows. Cost per lead got 29.6% worse. Cost per booked appointment got 16.7% better. Cost per closed deal fell from $2,769 to $2,308 — the same 16.7%, because show rate and close rate did not move. That is the tell. When show rate and close rate are held constant, cost per appointment and cost per deal always move together; the divergence between cost per lead and everything downstream is the entire argument.

Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.

The four measurement choices that move your reported AI ROI most

Ranked by how far each one can move the number, worst first. The first three are computed from the example above; the fourth is what a single uncorrected one-off event does to a baseline of that size.

Choice The version that misleads The version to use How far it moved the number
Denominator Cost per lead Cost per booked qualified appointment Verdict flips: +29.6% worse becomes 16.7% better
Cohort rule Revenue that closed in the window Revenue from leads created in the window ROI 290% → 239% when 12 of 91 deals were pre-existing
Cost basis Media spend only Media + platform + implementation + internal hours Cost per appointment understated by 14.3% ($321 vs $375)
Baseline window “Last quarter”, all sources pooled Same 90 days, same sources, one-off spikes excluded One trade show worth 900 of 4,000 baseline leads makes the baseline 29% larger than the comparable period it is meant to represent

Three of the four are decisions, not measurements. They get made once by whoever built the dashboard and are never revisited — which is why two people in one company can produce ROI numbers 50 points apart from the same database.

“My baseline is a mess” — how to rebuild one you can defend

You almost certainly cannot get a clean before-state. Rebuild a defensible one instead, in this order:

  1. Fix the window at 90 days ending the day before go-live. Not “last year”, not the last full quarter. Seasonality contaminates anything longer, and anything shorter will not clear the evidence floor below.
  2. Exclude one-off events from both windows — a trade show, a launch, a competitor outage. If you cannot exclude the event, exclude the whole source from both sides.
  3. Freeze source taxonomy. If lead-source values were renamed mid-window, the before and after are not comparable. Map old values to new ones and record the mapping.
  4. Count sales-accepted leads, not form fills. Form fills include bots, duplicates and competitors; your rejection rate is your best estimate of how dirty the top of the funnel is.
  5. Write the three locks down and date them — denominator, cost basis, cohort rule — before you see any results. This is the only step that survives a hostile CFO.

If the source tagging in your CRM is genuinely unrecoverable, the honest move is to declare the pre-period unusable and measure forward from a stamped cohort instead. The four-field first-touch stamp we use to credit pipeline to AI search referrals is the same mechanism, and it works for any channel.

If we can’t make you money, we don’t deserve yours.

Pay-Per-Result pricing — performance-based alignment.

50,769+
AI-booked appointments
Average sales lift — median closer to 4×
Pay-Per-Result
Performance-based alignment

How many leads do I need before a before/after means anything?

Use the standard two-proportion sample size at 80% power and 5% significance: n per arm ≈ 7.84 × [p₁(1−p₁) + p₂(1−p₂)] ÷ (p₂−p₁)², where p₁ is your current lead-to-appointment rate and p₂ is the rate you are trying to prove. From a 10% baseline:

Lift you want to prove Leads per arm Leads in total At 1,200 leads/month
10% → 11% (1 pt) 14,731 29,462 24.6 months — not feasible
10% → 12% (2 pts) 3,834 7,668 6.4 months
10% → 13% (3 pts) 1,769 3,538 3.0 months
10% → 15% (5 pts) 682 1,364 5 weeks
10% → 20% (10 pts) 196 392 10 days

At 1,200 leads a month, a single quarter can prove a 3-percentage-point move in your lead-to-appointment rate and nothing smaller. A pilot on 400 leads cannot detect anything under about a 10-point swing, which is why so many AI pilots end in an argument rather than a decision. If the effect you expect is small, do not run a bigger pilot — move the measurement to a step where the effect is large relative to the noise, such as speed to first contact or appointment set rate.

The Denominator Lock: three things to fix before go-live

A rule you can hand to a CFO in one line: lock the denominator, the cost basis and the cohort rule before the system goes live, in writing, and report every subsequent number against all three.

  • Denominator lock. One unit, named. Cost per booked qualified appointment, with “qualified” defined by a written criterion, not a feeling.
  • Cost basis lock. Every line that would disappear if you switched the system off, including the hours your own staff spend supervising it.
  • Cohort lock. Revenue belongs to the period the lead was created. If your sales cycle is 60 days, your first honest read is at day 150, not day 90.

Our own headline figure is written to this rule and you can check it: the 7x average sales lift defined on our methodology page is trailing three-month closed-deal revenue at month six over the trailing three months before launch, averaged across clients who supplied both — and the same page discloses the median is closer to 4x. An average and a median that far apart is information, not an embarrassment. Ask any vendor quoting a multiple for both.

The denominator error we made ourselves

In our own editorial review on 11 September 2026, reviewers re-checked the numbers on 30 draft pages against their sources. Twenty of the thirty needed a correction before publishing, and one of those corrections was a denominator error of ours. One page had taken our Colliers-era database reactivation rate of 4.4% — which is appointments booked per dormant record contacted, set out on our database reactivation case record — and spent it as jobs won. Same number, wrong unit, wildly wrong business case.

We fixed it by stating the denominator in all four places the figure appears on the page that carried the error. A rate with no stated denominator is not a measurement, and the error stays invisible until someone builds a forecast on it. Look for it in your own AI ROI deck before your CFO does.

What this measurement work actually costs to run

The locks take about a day to write and agree. Rebuilding a defensible baseline from a contaminated CRM is the expensive part — budget 15 to 30 hours of analyst time for taxonomy mapping and one-off exclusion, plus 2 to 4 hours a month to keep the cohort report honest. It needs someone who can write SQL against your CRM and is allowed to say the pilot did not work.

What breaks at volume is not the arithmetic, it is the discipline: every new channel, renamed source and mid-quarter budget shift reopens the baseline question.

There is also a structural answer. On a pay-per-result appointment setting model rather than a retainer, the denominator is chosen for you: you pay on booked qualified appointments, so cost per booked appointment is the invoice. That removes an accounting problem, not a strategic one — you still have to decide whether the appointments are worth their price. Wider context: our overview of AI for business, and the Australian marketing ROI benchmarks for 2026 to compare yourself against.

Frequently asked questions

What is ROI in marketing, and how is it different from ROAS?

Marketing ROI is (revenue attributable to marketing minus total marketing cost) divided by total marketing cost, expressed as a percentage. ROAS is revenue divided by ad spend only, expressed as a multiple. ROAS ignores platform fees, agency costs, staff hours and gross margin, so the same campaign can show a 3x ROAS and a negative ROI.

Can I just use the attribution report in my CRM?

Not as proof. Observational attribution credits the channel that touched the buyer, not the channel that caused the purchase. In the eBay field experiments of Blake, Nosko and Tadelis (Econometrica, 2015), ordinary regression methods implied a paid-search ROI of over 4,100% without time and geographic controls, and over 1,400% with them, while the randomised version of the same measurement produced minus 63%, with a 95% confidence interval of minus 124% to minus 3%. When eBay switched off its own brand-keyword ads, 99.5% of the lost paid clicks were immediately picked up by natural search. Use the CRM report to steer weekly; use a holdout to decide.

How big does the lift have to be before I can prove it?

Bigger than most people assume. Lewis and Rao, reviewing 25 large field experiments with US retailers and brokerages in the Quarterly Journal of Economics, found the median confidence interval on advertising ROI was over 100 percentage points wide, and that informative experiments can easily require more than 10 million person-weeks. Most mid-market teams will never reach that scale, which is the argument for measuring AI at the funnel step it changed, where the effect is large relative to the noise, rather than at the top line.

How long before AI marketing ROI shows up?

Add your sales cycle to your measurement window. Process metrics such as speed to first contact and lead-to-appointment rate move within days. Revenue from the same cohort cannot appear until the cycle completes, so with a 60-day cycle the first honest read on a 90-day cohort is at day 150. Anyone reporting closed-won ROI at day 30 is reporting deals that were already in the pipeline.

Why does my AI marketing ROI look different in every report?

Usually three unrecorded decisions: which denominator, which costs are included, and whether revenue is credited to the close date or the lead-creation date. In the worked example above those three choices span roughly 50 ROI points on identical data. Write them down once, date them, and require every report to state them.

Should I run a holdout if my baseline is unusable?

Yes, if your volume supports it. A holdout replaces the contaminated before-period with a concurrent control, which removes seasonality and taxonomy drift in one move. Hold back 20% to 50% of comparable leads, keep the assignment random rather than by rep or region, and check the sample table above before you start — a holdout too small to detect your expected lift is worse than no test, because it produces a confident null.

Pay-Per-Result appointments

See if we’re a fit

We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.

  • 50,769+ appointments booked without cold calling.
  • Pay-Per-Result pricing — you pay for booked, qualified calls.
  • Pick your own time on our live calendar, no phone tag.

View all articles

Pay-Per-Result · No retainers

Turn this into booked sales calls.

Our AI agents — trained on 50,769+ booked appointments — fill your calendar with pre-qualified buyers. You only pay when calls land.

Keep reading

Related on Leads Now AI

The thesis behind everything we do

Why Pay-Per-Result is the only marketing pricing model that aligns the agency with you

Leads Now AI is a 100% Pay-Per-Result marketing agency. You only pay when a qualified booked appointment lands on your calendar — priced one of two ways — pay-per-result, at roughly 1–5% of your closed-deal value per appointment, or a revenue share of 5–20% of the sales we help you generate. Both bill on outcomes. Not on clicks. Not on lead-form fills. Not on retainer months. Not on “strategy hours.” If the calendar stays empty, you owe zero. See full pricing →

1. Incentives align

The agency only succeeds when you succeed. We eat the cost of bad ad creative, bad lists, ICP mismatches and no-shows. You never pay for our learning curve.

2. Self-selecting shortlist

Only an agency confident in its delivery can operate this model. The pool of Pay-Per-Result agencies is tiny precisely because most agencies can’t survive on it. Pick from the agencies who can.

3. Cost cannot detach from revenue

Sized to 1–5% of closed-deal value, your acquisition cost stays sustainable across LTV bands. A $500-membership business and a $50,000-engagement business can both run the model profitably.

4. No retainer trap

The standard engagement carries no monthly retainer — nothing arrives on your invoice regardless of outcome. No 6 or 12-month lock-in, no clawback on appointments already delivered, cancel any time with 7 days notice. Early-stage businesses that need the sales systems built first are quoted scoped groundwork up front, never a standing fee.

5. De-risks the pilot

Test before commitment. A small scope-based setup fee covers hard build costs; everything after that is purely outcome-linked. There’s no “we’ll see how it performs after $30k of spend.”

6. Forces agency discipline

If our AI agents qualify poorly, if our reminders fail, if our no-show recovery doesn’t fire — we eat the cost. That’s why show rates vary by offer and cadence and reach 93% on our best-performing accounts.

The volume argument

A fully-ramped human SDR produces on the order of $200,000 a year. They work one conversation at a time, sleep, take leave, and cap out at a territory. Our agents work every lead in the list in parallel — responding in seconds, following up indefinitely without getting bored, and adding capacity without adding headcount.

At 100 qualified booked appointments a month against a $5,000 average deal value, that is $500,000 of booked pipeline every month — roughly what one SDR produces in two and a half years.

Read that precisely: booked pipeline means appointments multiplied by your average deal value. It is not closed revenue — closing is your side of the table, and your close rate decides what lands. The inputs above are a worked example; we size them to your actual deal economics before quoting. What we can evidence on our own numbers: 50,769+ appointments delivered since 2017, database reactivation converting 4.4–8.9% on dormant CRM lists, and show rates that vary by offer and reminder cadence — up to 93% on our best-performing accounts.

The proof: 50,769+ AI-booked sales appointments delivered since 2017 across coaches, consultants, RTOs, course creators, finance brokers and B2B service firms in Australia, USA, UK, Canada, NZ and Europe. Named clients include Sam Tajvidi (121 Brokers), Marcus Wilkinson (Iron Body), Foundr, SheSells.online and Lambda Academy. Wikidata Q139846230. See full Pay-Per-Result pricing →