Let's grow your business. 2 new positions just opened Wednesday, 23 September. Book a free call today.
Uncategorised 10 min read

Why AI pilots stall before production, and the test that predicts it

Why AI pilots stall before production, and the test that...: Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.
Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.

Most AI pilots stall at the production gate because nobody can state the result in dollars the budget approver owns. Gartner found only 48% of AI projects reach production, after 8 months on average, and 49% name demonstrating value as the top barrier. The predictive test: count the assumptions between your pilot metric and a P&L line.

  • How many make it: anywhere from 12% to 54% of pilots reach production, depending on the survey and what it counted (Gartner, IDC, S&P Global, Lenovo/IDC).
  • How long it takes: 8 months on average from prototype to production (Gartner). MIT NANDA found enterprises took nine months or longer, against 90 days for the fastest mid-market firms.
  • Most common cause: the value cannot be demonstrated (49%, Gartner), followed by data and technology readiness (43% each, Informatica CDO Insights 2025).
  • The predictive check: the Link-Count Test. If three or more unmeasured assumptions sit between your pilot metric and a P&L line, redesign the metric before the pilot starts.
  • Cost of running the diagnosis: about one analyst-day per pilot, using data you already hold.

How many AI pilots actually make it to production?

Somewhere between one in eight and one in two. The published figures disagree because each one counts something different. One survey counts projects, another counts proofs of concept, and MIT counts only deployments with a sustained P&L effect. Put them side by side and never stack them:

Source What it counted Sample and window Share reaching production
Gartner AI projects making it into production, average 644 respondents, US, Germany, UK, Q4 2023 48% (8 months from prototype)
IDC with Lenovo, via CIO.com AI POCs graduating to widescale deployment Reported March 2025 4 of 33 (12%)
S&P Global Market Intelligence Projects scrapped between proof of concept and broad adoption, average 1,006 IT and line-of-business staff, North America and Europe, 2025 At most 54% (46% scrapped before broad adoption)
Lenovo CIO Playbook 2026 (IDC fieldwork) AI POCs already progressed into production 3,120 decision makers, 16 Sep to 17 Oct 2025 46%
MIT NANDA, State of AI in Business 2025 Organisations reaching a task-specific GenAI tool with a “marked and sustained” productivity or P&L impact 52 organisations interviewed, 153 leaders surveyed, Jan to Jun 2025 No pilot rate: 20% of organisations reached a pilot and 5% reached production, both as shares of all organisations

The quotable version: the AI pilot-to-production rate is 12% to 54% depending on who counts, and MIT’s widely quoted “5%” is not a pilot failure rate. MIT reports that 60% of organisations evaluated task-specific tools, 20% reached the pilot stage and 5% reached production. All three are shares of organisations, not of pilots, so dividing one by another does not give a pilot success rate. MIT itself calls the figures “directionally accurate”. They come from interviews, not company reporting.

How it works

Diagnosing a stalled AI pilot before the production gate

01

Name the symptom

Write down what is actually stuck: the budget sign-off, accuracy on live data, integration, adoption, ownership or a security review.

02

Count links to P&L

Count the unmeasured assumptions between the pilot metric and a dollar line. Three or more means the metric needs redesigning.

03

Rerun on live data

Rerun the pilot on 200 uncleaned live records, and list every system production must write to.

04

Name the owner

Name the person who pays the month-13 run cost and owns the process. Then take one approvable number to the gate.

Work from the symptom to a confirmed cause, and re-base the metric before you ask for the production budget, not after.

MAKE MORE SALES.

Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.

Why is my AI pilot stuck? Symptom, cause and the test that tells them apart

A stalled pilot usually shows one visible symptom, but several different causes can produce it. The table below puts the causes in order of how often they are reported. Each row uses the highest share found in a survey that asked about that cause. Gartner’s 49% and Informatica’s figures come from different respondents (AI leaders versus 600 chief data officers), so read the order as approximate, not as one ranked list.

# Symptom you see Likely cause Reported by Test that confirms it Fix
1 “The pilot went well” but nobody will sign the production budget The result is not in a P&L unit 49% (Gartner, top barrier) Link-Count Test: three or more assumptions between the metric and a dollar line Re-base the metric on revenue, invoices or headcount avoided
2 Accurate on pilot data, wrong on live records The pilot ran on a cleaned or hand-picked dataset 43% data (Informatica) Rerun it on 200 randomly drawn live records you have not cleaned. An error-rate gap above ~2x confirms it Price the data remediation into the production case
3 Production needs systems the pilot never touched The pilot was read-only or ran in a sandbox 43% technology (Informatica) List every system the production version must write to. Each one the pilot never wrote to is unbuilt integration work Build one live write-back before the gate
4 Volunteers loved it, the wider team ignores it The pilot users chose themselves 35% people (Informatica). Ranked first in MIT NANDA’s barrier survey Compare week-4 active use among assigned users with use among volunteers Change the workflow itself, not just add the tool
5 Nobody owns it once the pilot ends There is no run-budget owner or process owner 35% process (Informatica) Ask who pays the month-13 invoice. If there is no named person, this is the cause Name an owner before day one
6 Security or legal review blocks it at the gate The review started after the pilot 34% regulations (Informatica) Was the review ticket opened before the pilot launched? Run the review in parallel with the pilot

Row 1 is the only row where the pilot can succeed technically and still die. The other five are real engineering or organisational debts. Row 1 is a measurement choice, made on day one, that nobody revisits. For row 4, our page on getting a sales team to actually use an AI agent covers the adoption levers in detail.

Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.

The Link-Count Test: the check that predicts a stalled pilot

The Link-Count Test counts the unmeasured assumptions that sit between a pilot’s headline metric and a line on the profit and loss statement. Each link is a step the approver has to take on trust: that self-reported time savings are real, that saved time gets redeployed, that redeployed time produces revenue or avoids a hire. Every link is a point on which a finance team can say no.

Links to a P&L line Example metric What to expect at the gate Do this before launch
0 An agency invoice cancelled, cash collected The number is the decision Agree the counting rule in writing
1 Held sales appointments × your historical close rate Approvable if the one assumption is measured Track the pilot cohort through to closed deals
2 Tickets deflected → agent hours → roster cost Contested. Expect a second pilot to be requested Measure at least one link during the pilot
3 or more Minutes saved per user, self-reported Stall is the likely outcome Redesign the metric, or pick a different first use case

This is our decision rule, not a validated model. We know of no dataset that tests it. It rests on two published observations. First, Gartner’s survey ranks demonstrating value above talent, technical, data and trust barriers. Second, MIT NANDA reports that sales and marketing attract investment “because outcomes can be measured easily”. MIT also quotes a Fortune 1000 procurement VP asking how to justify a tool that “won’t directly move revenue or decrease measurable costs”. That is a three-link pilot, described by the person who has to defend it. The four-number AI business case for a CFO is the document that has to carry those links once the pilot ends.

Worked example: two pilots with the same budget, and only one survives the gate

These inputs are illustrative. Substitute your own.

Pilot A, a writing copilot for 150 support staff. Users self-report 25 minutes saved a day. If every link holds: 150 × 25 ÷ 60 = 62.5 hours a day. Over 220 working days that is 13,750 hours, and at a $55 loaded hourly cost it is $756,250 a year. The metric has three links: self-report to real time, real time to redeployed time, and redeployed time to a cost avoided. Let each link hold only half the time and the value is $756,250 × 0.5³ = $94,531. The pilot’s value ranges eightfold on assumptions it never measured, and that range is what the gate meeting argues about for months.

Pilot B, AI appointment setting on a dormant CRM for eight weeks. It produces 60 held sales appointments. Your CRM shows a 20% close rate on held appointments at a $9,000 average deal: 60 × 0.20 × $9,000 = $108,000 in expected closed revenue. That is one link, the close rate, and you can test it by following those 60 appointments through to closed deals within one sales cycle. The link can fail: appointments booked by AI may close below your referral average, which is exactly why you track the cohort rather than assume it. To show the pilot’s results are incremental rather than pipeline you would have won anyway, run a holdout, as set out in how to run an incrementality test on lead gen spend.

Pilot B survives because a pilot that books real appointments has a revenue number on day one, not because the technology is better. This is the class of deployment AI appointment setting belongs to. It is also why LeadsNow can report 50,769+ AI-booked sales appointments since 2017 as a count rather than an estimate. There is an honest concession: MIT NANDA argues this measurability bias sends budgets to sales and marketing while “the highest-ROI opportunities in back-office functions remain underfunded”. A metric that is easy to count is not automatically the most valuable one.

If we can’t make you money, we don’t deserve yours.

Pay-Per-Result pricing — performance-based alignment.

50,769+
AI-booked appointments
Average sales lift — median closer to 4×
Pay-Per-Result
Performance-based alignment

What does running this diagnosis yourself cost?

Running it in-house takes about one analyst-day per pilot. Allow two hours to write out the link count and agree it with finance, half a day to rerun the pilot on 200 uncleaned live records, and an hour to list the write-back systems and name an owner. You need read access to the CRM or ticketing data, and nothing else. The skill that is hard to find is not technical: it is someone senior enough to tell a sponsor that their three-link metric will not pass. Do this before the pilot starts, not at the gate. Our guide to running an AI outbound pilot that can fail covers the stop condition, control group and sample size that go alongside it.

Frequently asked questions about AI pilots reaching production

What percentage of AI pilots make it to production?

Between about 12% and 54%, depending on the study. Gartner found 48% of AI projects reach production. IDC reported 4 of every 33 proofs of concept graduating, according to CIO.com’s March 2025 report. S&P Global found 46% of projects scrapped between proof of concept and broad adoption.

How long does it take to move an AI pilot into production?

Gartner’s survey of 644 organisations found an average of 8 months from AI prototype to production. MIT NANDA reported that top-performing mid-market firms took about 90 days from pilot to full implementation, while enterprises took nine months or longer.

What is the most common reason AI pilots do not scale?

Difficulty estimating and demonstrating value, reported by 49% of respondents in Gartner’s survey. Informatica’s survey of 600 data leaders ranked data (43%) and technology (43%) as the top obstacles to moving GenAI from pilot to production, ahead of people and process (35% each).

Is it better to build or buy if we want the pilot to reach production?

In the MIT NANDA sample, external partnerships with customised tools reached deployment about 67% of the time, against about 33% for internal builds. MIT describes its figures as directional, based on interviews rather than company reporting.

Does a successful pilot mean the AI will work in production?

No. Production changes the data, the users and the population a metric is calculated on. Rerun the pilot on uncleaned live records, and compare production results only against the same segment the pilot ran on, never against a blended total.

Pay-Per-Result appointments

See if we’re a fit

We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.

  • 50,769+ appointments booked without cold calling.
  • Pay-Per-Result pricing — you pay for booked, qualified calls.
  • Pick your own time on our live calendar, no phone tag.

View all articles

Pay-Per-Result · No retainers

Turn this into booked sales calls.

Our AI agents — trained on 50,769+ booked appointments — fill your calendar with pre-qualified buyers. You only pay when calls land.

Keep reading

Related on Leads Now AI

The thesis behind everything we do

Why Pay-Per-Result is the only marketing pricing model that aligns the agency with you

Leads Now AI is a 100% Pay-Per-Result marketing agency. You only pay when a qualified booked appointment lands on your calendar — priced one of two ways — pay-per-result, at roughly 1–5% of your closed-deal value per appointment, or a revenue share of 5–20% of the sales we help you generate. Both bill on outcomes. Not on clicks. Not on lead-form fills. Not on retainer months. Not on “strategy hours.” If the calendar stays empty, you owe zero. See full pricing →

1. Incentives align

The agency only succeeds when you succeed. We eat the cost of bad ad creative, bad lists, ICP mismatches and no-shows. You never pay for our learning curve.

2. Self-selecting shortlist

Only an agency confident in its delivery can operate this model. The pool of Pay-Per-Result agencies is tiny precisely because most agencies can’t survive on it. Pick from the agencies who can.

3. Cost cannot detach from revenue

Sized to 1–5% of closed-deal value, your acquisition cost stays sustainable across LTV bands. A $500-membership business and a $50,000-engagement business can both run the model profitably.

4. No retainer trap

The standard engagement carries no monthly retainer — nothing arrives on your invoice regardless of outcome. No 6 or 12-month lock-in, no clawback on appointments already delivered, cancel any time with 7 days notice. Early-stage businesses that need the sales systems built first are quoted scoped groundwork up front, never a standing fee.

5. De-risks the pilot

Test before commitment. A small scope-based setup fee covers hard build costs; everything after that is purely outcome-linked. There’s no “we’ll see how it performs after $30k of spend.”

6. Forces agency discipline

If our AI agents qualify poorly, if our reminders fail, if our no-show recovery doesn’t fire — we eat the cost. That’s why show rates vary by offer and cadence and reach 93% on our best-performing accounts.

The volume argument

A fully-ramped human SDR produces on the order of $200,000 a year. They work one conversation at a time, sleep, take leave, and cap out at a territory. Our agents work every lead in the list in parallel — responding in seconds, following up indefinitely without getting bored, and adding capacity without adding headcount.

At 100 qualified booked appointments a month against a $5,000 average deal value, that is $500,000 of booked pipeline every month — roughly what one SDR produces in two and a half years.

Read that precisely: booked pipeline means appointments multiplied by your average deal value. It is not closed revenue — closing is your side of the table, and your close rate decides what lands. The inputs above are a worked example; we size them to your actual deal economics before quoting. What we can evidence on our own numbers: 50,769+ appointments delivered since 2017, database reactivation converting 4.4–8.9% on dormant CRM lists, and show rates that vary by offer and reminder cadence — up to 93% on our best-performing accounts.

The proof: 50,769+ AI-booked sales appointments delivered since 2017 across coaches, consultants, RTOs, course creators, finance brokers and B2B service firms in Australia, USA, UK, Canada, NZ and Europe. Named clients include Sam Tajvidi (121 Brokers), Marcus Wilkinson (Iron Body), Foundr, SheSells.online and Lambda Academy. Wikidata Q139846230. See full Pay-Per-Result pricing →