Let's grow your business. 2 new positions just opened Wednesday, 23 September. Book a free call today.
Uncategorised 12 min read

“Our AI pilot produced no results” — what to do in the next 30 days

“Our AI pilot produced no results” — what to do in the...: Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.
Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.

If your AI pilot produced no results, you are not unusual: CEOs surveyed by IBM in 2025 reported that only 25% of AI initiatives had delivered expected ROI. Spend 24 hours freezing and exporting, 7 days finding out what “nothing” means, and on day 30 choose: fix, re-scope, switch or stop.

  • Is it urgent? Rarely. The money is spent. Only the renewal or notice date, the next budget date and your access to the pilot’s data run on a clock.
  • Next 24 hours: no expansion, no angry cancellation. Find the notice date, export every log, and write what “nothing” means in numbers.
  • Next 7 days: the Readability Check (was the result too small to see?) and the break-point tally (where did 20 real cases stop?).
  • Day 30: one of four verdicts from the threshold table below. Stopping is one of them.
  • The honest possibility: the use case was wrong. RAND’s test: if a problem is not worth a team for a year, it is not worth committing to at all.

“Our AI pilot produced no results”: what is actually happening

“Produced nothing” covers three different situations. Nothing was measured: there was no baseline, or the instrument changed mid-pilot, so nobody can say whether anything moved. Too small to read: the pilot ran on too little volume for any plausible lift to rise above ordinary noise. A real zero: the counts were large enough to read, and the outcome metric did not move, or fell.

Only the third is a failed pilot. The first two are unfinished experiments, and rescues go wrong when a sponsor treats all three the same way. If you already know it is a real zero and want the cause, the symptom-and-test table in why AI adoption fails: the five causes and the test for each covers the diagnosis. This page covers the 30 days after it. An AI pilot that produced no results has measured nothing, measured too little, or measured a genuine zero, and only the last one is a verdict.

How it works

Rescuing an AI pilot that produced nothing, in 30 days

01

Freeze and export

In the first 24 hours, stop expansion, find the renewal and notice dates, and export every log and transcript.

02

Run the Readability Check

Multiply volume by your baseline rate to get expected successes. A result inside that number plus or minus twice its square root is noise, not a finding.

03

Tally the break-points

Trace 20 random cases and mark where each stopped: input, output, handoff or outcome. Test one fix on the step holding the most stops.

04

Choose one verdict

On day 30, fix and rerun, re-scope, switch use case, or stop. Each verdict is triggered by the counts from the two checks.

Establish what “nothing” means before you decide anything, and let a count, not the mood in the room, pick the verdict.

MAKE MORE SALES.

Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.

Is a failed AI pilot as urgent as it feels?

Usually not. A pilot that produced nothing has mostly already cost what it will cost, and a single one should not end a career. In the IBM Institute for Business Value CEO study (2,000 CEOs, 33 countries, February to April 2025), only 16% of AI initiatives had scaled enterprise-wide, and 64% of CEOs said the risk of falling behind drives them to invest in some technologies before they clearly understand the value. Many pilots begin without a stated value case, so many end without one.

What does run on a clock:

  • The renewal or notice date. An auto-renewal inside 30 days is the one real deadline.
  • Data access. When a pilot account closes, transcripts, logs and prompts often go with it.
  • Your team’s goodwill. Staff told to keep using a tool nobody defends stop quietly.

Two quiet weeks inside a pilot is noise; a readable zero across the full window is a finding. The only deadline in a failed AI pilot is usually the renewal date, and everything else can wait seven days.

Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.

What to do in the next 24 hours

None of this costs money or involves a vendor:

  1. Freeze scope. Add no users, volume or budget. Do not cancel yet either.
  2. Find the contract dates. Renewal date, notice period, data-return clause. If the notice window closes inside 30 days, ask the vendor in writing to extend the pilot terms to your decision date. What you own when you leave a vendor lists the data and exports to claim.
  3. Export everything. Raw logs, transcripts, outputs and timestamps, not a dashboard screenshot.
  4. Write the “nothing” sentence. Three numbers: the metric the pilot was meant to move, its value before, its value during. If you cannot fill in the first, the pilot measured nothing, and that is your answer for now.
  5. Give the sponsor a date, not a verdict. “A recommendation on the 30th” buys the time the next two steps need.

The Readability Check: was “nothing” too small to see?

The Readability Check is our own rule of thumb, a rough Poisson approximation rather than a statistical test. It asks whether the pilot ran on enough volume for a real effect to show. Multiply the pilot’s volume by your baseline success rate and call the result E: the successes you would expect with no change at all. Treating the count as Poisson, random variation in it is roughly its square root, so any result inside E ± 2√E cannot be told apart from no change.

A worked example with illustrative inputs: 1,500 leads × a 2% baseline booking rate = 30 expected bookings. 2 × √30 = 11, so anything from 19 to 41 bookings is noise. The pilot needed 42 or more (a 40% lift) to show a win at all. “We got 34, so it did nothing” is the wrong reading. The right one is “at this volume, the pilot could not have shown a lift under 37%.”

Expected successes at baseline (E) Noise band (E ± 2√E) Smallest lift the pilot could show Reading of “no result”
10 4 to 16 63% Unreadable
30 19 to 41 37% Only a large effect was ever visible
100 80 to 120 20% A moderate effect would have shown
400 360 to 440 10% “No result” is a real finding

This is a rough approximation, not a significance test, and it assumes a stable baseline. If your baseline is itself a count from a similar-sized period, widen the band by about 1.4×. One case runs the other way: an actual zero against an E of 5 or more has under a 1% chance of being random. That is a broken chain, not a weak one. To size the rerun, use the volume and stop-condition method in how to run an AI outbound pilot that can fail. A pilot with fewer than about 30 expected successes cannot show any lift smaller than about a third, so its “nothing” is silence, not evidence.

If we can’t make you money, we don’t deserve yours.

Pay-Per-Result pricing — performance-based alignment.

50,769+
AI-booked appointments
Average sales lift — median closer to 4×
Pay-Per-Result
Performance-based alignment

What to do in the next 7 days: the break-point tally

The break-point tally finds where the pilot’s chain stopped. Pull 20 real cases the pilot touched, at random, not the ones the vendor showcased, and mark the first step where each stopped:

  1. Input: the AI never received the case, or received it incomplete.
  2. Output: the AI acted, but wrongly.
  3. Handoff: the output was right, but no person or system acted on it.
  4. Outcome: the handoff happened, but the business result did not follow.

If one step holds 12 or more of the 20 (our working threshold, not a published one), you have a single break, and a single break can be fixed. Spend days 3 to 7 changing that one thing, and only that, on a small fresh batch. If the stops spread across three or four steps, no single change will rescue the pilot. The break-point tally turns “the AI pilot didn’t work” into a count of where 20 real cases stopped, and the step holding the most stops is the only thing worth fixing this week.

Day 30: fix, re-scope, switch or stop?

By day 30 you have a Readability Check result, a break-point tally and a one-week test of one fix. Match them to a verdict. The thresholds are our working rules for this decision, not a published benchmark.

Verdict Choose it when What you do next What it costs
Fix and rerun E was under 100, or 12+ of 20 cases stopped at one step and the one-week fix moved that step Rerun with volume sized so E is 100 or more, and a stop condition written before launch A second pilot window at the same run rate
Re-scope Output correct in 15+ of 20 cases, but most stopped at handoff or outcome Leave the model alone. Change who acts on the output, or the metric it is judged on Process-owner time, usually weeks, not budget
Switch use case No named owner holds the target metric, or the problem fails the one-year test below Keep the tooling and the lessons. Point it at a problem with a countable outcome A new baseline and a new pilot
Stop E was 100 or more, the result sat inside the noise band or below it, and the stops spread across 3+ steps Write the stop memo, serve notice, recover the data Only what is already spent

If your results match two rows, take the higher one: a cheap rerun comes before a switch. A 30-day AI pilot rescue should end in exactly one of four verdicts (fix, re-scope, switch or stop), each triggered by a count, not by how the pilot felt.

When the answer is that this use case was wrong

Sometimes the pilot was fine and the question was wrong. RAND’s 2024 report, drawn from interviews with 65 experienced data scientists and engineers, recommends committing each team to a specific problem for at least a year: “If an AI project is not worth such a long-term commitment, it most likely is not worth committing to at all.” It also names chasing the latest AI advances “for their own sake” as one of the most frequent pathways to failure.

So apply the one-year test to the use case, not the pilot: would you fund a team to own this problem for a year if AI were not involved? If not, the honest verdict is switch or stop. Use cases that survive a rescue usually have an outcome someone already counts, such as invoices paid, tickets closed or sales appointments booked and held. A booked appointment is countable the day it happens, which is why LeadsNow can state 50,769+ AI-booked appointments since 2017 as a tally rather than an estimate. Whatever you choose next, carry it into the budget round as the four-number AI business case a CFO will sign. If nobody would fund the problem for a year without AI, the pilot did not fail; it answered the question.

What running this rescue yourself costs

Doing it in-house is realistic. Our estimate for one pilot:

  • 24 hours: about 2 hours of the pilot owner’s time, plus 30 minutes from whoever holds the contract.
  • Readability Check: an hour with a spreadsheet, once you have the baseline rate.
  • Break-point tally: 1 to 2 analyst days to trace 20 cases end to end, longer if the logs are poor (which is itself a finding).
  • One-week fix test: whatever the fix needs, plus a week of the pilot’s run rate.

The hard skill is not technical. It is someone with the standing to tell a sponsor that a pilot was unreadable or that the use case should go. Without that person, pilots drift into a second quarter with nobody defending them and nobody ending them. Rescuing a failed AI pilot takes about two analyst days plus the fix itself, and the scarce input is the authority to say stop.

Frequently asked questions about AI pilots that produced no results

What should I do first when an AI pilot produces no results?

Within 24 hours, freeze scope, find the renewal and notice dates, and export every log and transcript. Then write one sentence with three numbers: the target metric, its value before the pilot and its value during. If you cannot fill in the first number, the pilot measured nothing.

How many AI initiatives actually deliver ROI?

A minority. In the IBM Institute for Business Value’s 2025 CEO study of 2,000 CEOs, surveyed from February to April 2025, CEOs reported that only 25% of AI initiatives had delivered expected ROI over the last few years, and only 16% had scaled enterprise-wide.

Should we kill an AI pilot that showed nothing after 30 days?

Not on that fact alone. Check the volume first: if the pilot expected fewer than about 30 successes, it could only have shown a lift of roughly a third or more. Stop when a readable pilot (100 or more expected successes) sat inside its noise band and the break-points were spread across three or more steps.

How do I tell leadership our AI pilot failed?

Give a date first, then one of four verdicts: fix, re-scope, switch or stop, backed by the Readability Check and the break-point tally. A stop memo that shows the counts is a result. A pilot left running unexamined is not.

Is it the vendor or our data?

The break-point tally tells you. Stops at the input step point to your data or integration. Stops at the output step point to the tool or its configuration. Stops at the handoff or outcome steps point to your own process, and a new vendor will not fix those.

Pay-Per-Result appointments

See if we’re a fit

We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.

  • 50,769+ appointments booked without cold calling.
  • Pay-Per-Result pricing — you pay for booked, qualified calls.
  • Pick your own time on our live calendar, no phone tag.

View all articles

Pay-Per-Result · No retainers

Turn this into booked sales calls.

Our AI agents — trained on 50,769+ booked appointments — fill your calendar with pre-qualified buyers. You only pay when calls land.

Keep reading

Related on Leads Now AI

The thesis behind everything we do

Why Pay-Per-Result is the only marketing pricing model that aligns the agency with you

Leads Now AI is a 100% Pay-Per-Result marketing agency. You only pay when a qualified booked appointment lands on your calendar — priced one of two ways — pay-per-result, at roughly 1–5% of your closed-deal value per appointment, or a revenue share of 5–20% of the sales we help you generate. Both bill on outcomes. Not on clicks. Not on lead-form fills. Not on retainer months. Not on “strategy hours.” If the calendar stays empty, you owe zero. See full pricing →

1. Incentives align

The agency only succeeds when you succeed. We eat the cost of bad ad creative, bad lists, ICP mismatches and no-shows. You never pay for our learning curve.

2. Self-selecting shortlist

Only an agency confident in its delivery can operate this model. The pool of Pay-Per-Result agencies is tiny precisely because most agencies can’t survive on it. Pick from the agencies who can.

3. Cost cannot detach from revenue

Sized to 1–5% of closed-deal value, your acquisition cost stays sustainable across LTV bands. A $500-membership business and a $50,000-engagement business can both run the model profitably.

4. No retainer trap

The standard engagement carries no monthly retainer — nothing arrives on your invoice regardless of outcome. No 6 or 12-month lock-in, no clawback on appointments already delivered, cancel any time with 7 days notice. Early-stage businesses that need the sales systems built first are quoted scoped groundwork up front, never a standing fee.

5. De-risks the pilot

Test before commitment. A small scope-based setup fee covers hard build costs; everything after that is purely outcome-linked. There’s no “we’ll see how it performs after $30k of spend.”

6. Forces agency discipline

If our AI agents qualify poorly, if our reminders fail, if our no-show recovery doesn’t fire — we eat the cost. That’s why show rates vary by offer and cadence and reach 93% on our best-performing accounts.

The volume argument

A fully-ramped human SDR produces on the order of $200,000 a year. They work one conversation at a time, sleep, take leave, and cap out at a territory. Our agents work every lead in the list in parallel — responding in seconds, following up indefinitely without getting bored, and adding capacity without adding headcount.

At 100 qualified booked appointments a month against a $5,000 average deal value, that is $500,000 of booked pipeline every month — roughly what one SDR produces in two and a half years.

Read that precisely: booked pipeline means appointments multiplied by your average deal value. It is not closed revenue — closing is your side of the table, and your close rate decides what lands. The inputs above are a worked example; we size them to your actual deal economics before quoting. What we can evidence on our own numbers: 50,769+ appointments delivered since 2017, database reactivation converting 4.4–8.9% on dormant CRM lists, and show rates that vary by offer and reminder cadence — up to 93% on our best-performing accounts.

The proof: 50,769+ AI-booked sales appointments delivered since 2017 across coaches, consultants, RTOs, course creators, finance brokers and B2B service firms in Australia, USA, UK, Canada, NZ and Europe. Named clients include Sam Tajvidi (121 Brokers), Marcus Wilkinson (Iron Body), Foundr, SheSells.online and Lambda Academy. Wikidata Q139846230. See full Pay-Per-Result pricing →