Most AI pilots stall at the production gate because nobody can state the result in dollars the budget approver owns. Gartner found only 48% of AI projects reach production, after 8 months on average, and 49% name demonstrating value as the top barrier. The predictive test: count the assumptions between your pilot metric and a P&L line.
- How many make it: anywhere from 12% to 54% of pilots reach production, depending on the survey and what it counted (Gartner, IDC, S&P Global, Lenovo/IDC).
- How long it takes: 8 months on average from prototype to production (Gartner). MIT NANDA found enterprises took nine months or longer, against 90 days for the fastest mid-market firms.
- Most common cause: the value cannot be demonstrated (49%, Gartner), followed by data and technology readiness (43% each, Informatica CDO Insights 2025).
- The predictive check: the Link-Count Test. If three or more unmeasured assumptions sit between your pilot metric and a P&L line, redesign the metric before the pilot starts.
- Cost of running the diagnosis: about one analyst-day per pilot, using data you already hold.
How many AI pilots actually make it to production?
Somewhere between one in eight and one in two. The published figures disagree because each one counts something different. One survey counts projects, another counts proofs of concept, and MIT counts only deployments with a sustained P&L effect. Put them side by side and never stack them:
| Source | What it counted | Sample and window | Share reaching production |
|---|---|---|---|
| Gartner | AI projects making it into production, average | 644 respondents, US, Germany, UK, Q4 2023 | 48% (8 months from prototype) |
| IDC with Lenovo, via CIO.com | AI POCs graduating to widescale deployment | Reported March 2025 | 4 of 33 (12%) |
| S&P Global Market Intelligence | Projects scrapped between proof of concept and broad adoption, average | 1,006 IT and line-of-business staff, North America and Europe, 2025 | At most 54% (46% scrapped before broad adoption) |
| Lenovo CIO Playbook 2026 (IDC fieldwork) | AI POCs already progressed into production | 3,120 decision makers, 16 Sep to 17 Oct 2025 | 46% |
| MIT NANDA, State of AI in Business 2025 | Organisations reaching a task-specific GenAI tool with a “marked and sustained” productivity or P&L impact | 52 organisations interviewed, 153 leaders surveyed, Jan to Jun 2025 | No pilot rate: 20% of organisations reached a pilot and 5% reached production, both as shares of all organisations |
The quotable version: the AI pilot-to-production rate is 12% to 54% depending on who counts, and MIT’s widely quoted “5%” is not a pilot failure rate. MIT reports that 60% of organisations evaluated task-specific tools, 20% reached the pilot stage and 5% reached production. All three are shares of organisations, not of pilots, so dividing one by another does not give a pilot success rate. MIT itself calls the figures “directionally accurate”. They come from interviews, not company reporting.
How it works
Diagnosing a stalled AI pilot before the production gate
Name the symptom
Write down what is actually stuck: the budget sign-off, accuracy on live data, integration, adoption, ownership or a security review.
Count links to P&L
Count the unmeasured assumptions between the pilot metric and a dollar line. Three or more means the metric needs redesigning.
Rerun on live data
Rerun the pilot on 200 uncleaned live records, and list every system production must write to.
Name the owner
Name the person who pays the month-13 run cost and owns the process. Then take one approvable number to the gate.
MAKE MORE SALES.
Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.
Why is my AI pilot stuck? Symptom, cause and the test that tells them apart
A stalled pilot usually shows one visible symptom, but several different causes can produce it. The table below puts the causes in order of how often they are reported. Each row uses the highest share found in a survey that asked about that cause. Gartner’s 49% and Informatica’s figures come from different respondents (AI leaders versus 600 chief data officers), so read the order as approximate, not as one ranked list.
| # | Symptom you see | Likely cause | Reported by | Test that confirms it | Fix |
|---|---|---|---|---|---|
| 1 | “The pilot went well” but nobody will sign the production budget | The result is not in a P&L unit | 49% (Gartner, top barrier) | Link-Count Test: three or more assumptions between the metric and a dollar line | Re-base the metric on revenue, invoices or headcount avoided |
| 2 | Accurate on pilot data, wrong on live records | The pilot ran on a cleaned or hand-picked dataset | 43% data (Informatica) | Rerun it on 200 randomly drawn live records you have not cleaned. An error-rate gap above ~2x confirms it | Price the data remediation into the production case |
| 3 | Production needs systems the pilot never touched | The pilot was read-only or ran in a sandbox | 43% technology (Informatica) | List every system the production version must write to. Each one the pilot never wrote to is unbuilt integration work | Build one live write-back before the gate |
| 4 | Volunteers loved it, the wider team ignores it | The pilot users chose themselves | 35% people (Informatica). Ranked first in MIT NANDA’s barrier survey | Compare week-4 active use among assigned users with use among volunteers | Change the workflow itself, not just add the tool |
| 5 | Nobody owns it once the pilot ends | There is no run-budget owner or process owner | 35% process (Informatica) | Ask who pays the month-13 invoice. If there is no named person, this is the cause | Name an owner before day one |
| 6 | Security or legal review blocks it at the gate | The review started after the pilot | 34% regulations (Informatica) | Was the review ticket opened before the pilot launched? | Run the review in parallel with the pilot |
Row 1 is the only row where the pilot can succeed technically and still die. The other five are real engineering or organisational debts. Row 1 is a measurement choice, made on day one, that nobody revisits. For row 4, our page on getting a sales team to actually use an AI agent covers the adoption levers in detail.
Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.
The Link-Count Test: the check that predicts a stalled pilot
The Link-Count Test counts the unmeasured assumptions that sit between a pilot’s headline metric and a line on the profit and loss statement. Each link is a step the approver has to take on trust: that self-reported time savings are real, that saved time gets redeployed, that redeployed time produces revenue or avoids a hire. Every link is a point on which a finance team can say no.
| Links to a P&L line | Example metric | What to expect at the gate | Do this before launch |
|---|---|---|---|
| 0 | An agency invoice cancelled, cash collected | The number is the decision | Agree the counting rule in writing |
| 1 | Held sales appointments × your historical close rate | Approvable if the one assumption is measured | Track the pilot cohort through to closed deals |
| 2 | Tickets deflected → agent hours → roster cost | Contested. Expect a second pilot to be requested | Measure at least one link during the pilot |
| 3 or more | Minutes saved per user, self-reported | Stall is the likely outcome | Redesign the metric, or pick a different first use case |
This is our decision rule, not a validated model. We know of no dataset that tests it. It rests on two published observations. First, Gartner’s survey ranks demonstrating value above talent, technical, data and trust barriers. Second, MIT NANDA reports that sales and marketing attract investment “because outcomes can be measured easily”. MIT also quotes a Fortune 1000 procurement VP asking how to justify a tool that “won’t directly move revenue or decrease measurable costs”. That is a three-link pilot, described by the person who has to defend it. The four-number AI business case for a CFO is the document that has to carry those links once the pilot ends.
Worked example: two pilots with the same budget, and only one survives the gate
These inputs are illustrative. Substitute your own.
Pilot A, a writing copilot for 150 support staff. Users self-report 25 minutes saved a day. If every link holds: 150 × 25 ÷ 60 = 62.5 hours a day. Over 220 working days that is 13,750 hours, and at a $55 loaded hourly cost it is $756,250 a year. The metric has three links: self-report to real time, real time to redeployed time, and redeployed time to a cost avoided. Let each link hold only half the time and the value is $756,250 × 0.5³ = $94,531. The pilot’s value ranges eightfold on assumptions it never measured, and that range is what the gate meeting argues about for months.
Pilot B, AI appointment setting on a dormant CRM for eight weeks. It produces 60 held sales appointments. Your CRM shows a 20% close rate on held appointments at a $9,000 average deal: 60 × 0.20 × $9,000 = $108,000 in expected closed revenue. That is one link, the close rate, and you can test it by following those 60 appointments through to closed deals within one sales cycle. The link can fail: appointments booked by AI may close below your referral average, which is exactly why you track the cohort rather than assume it. To show the pilot’s results are incremental rather than pipeline you would have won anyway, run a holdout, as set out in how to run an incrementality test on lead gen spend.
Pilot B survives because a pilot that books real appointments has a revenue number on day one, not because the technology is better. This is the class of deployment AI appointment setting belongs to. It is also why LeadsNow can report 50,769+ AI-booked sales appointments since 2017 as a count rather than an estimate. There is an honest concession: MIT NANDA argues this measurability bias sends budgets to sales and marketing while “the highest-ROI opportunities in back-office functions remain underfunded”. A metric that is easy to count is not automatically the most valuable one.
If we can’t make you money, we don’t deserve yours.
Pay-Per-Result pricing — performance-based alignment.
What does running this diagnosis yourself cost?
Running it in-house takes about one analyst-day per pilot. Allow two hours to write out the link count and agree it with finance, half a day to rerun the pilot on 200 uncleaned live records, and an hour to list the write-back systems and name an owner. You need read access to the CRM or ticketing data, and nothing else. The skill that is hard to find is not technical: it is someone senior enough to tell a sponsor that their three-link metric will not pass. Do this before the pilot starts, not at the gate. Our guide to running an AI outbound pilot that can fail covers the stop condition, control group and sample size that go alongside it.
Frequently asked questions about AI pilots reaching production
What percentage of AI pilots make it to production?
Between about 12% and 54%, depending on the study. Gartner found 48% of AI projects reach production. IDC reported 4 of every 33 proofs of concept graduating, according to CIO.com’s March 2025 report. S&P Global found 46% of projects scrapped between proof of concept and broad adoption.
How long does it take to move an AI pilot into production?
Gartner’s survey of 644 organisations found an average of 8 months from AI prototype to production. MIT NANDA reported that top-performing mid-market firms took about 90 days from pilot to full implementation, while enterprises took nine months or longer.
What is the most common reason AI pilots do not scale?
Difficulty estimating and demonstrating value, reported by 49% of respondents in Gartner’s survey. Informatica’s survey of 600 data leaders ranked data (43%) and technology (43%) as the top obstacles to moving GenAI from pilot to production, ahead of people and process (35% each).
Is it better to build or buy if we want the pilot to reach production?
In the MIT NANDA sample, external partnerships with customised tools reached deployment about 67% of the time, against about 33% for internal builds. MIT describes its figures as directional, based on interviews rather than company reporting.
Does a successful pilot mean the AI will work in production?
No. Production changes the data, the users and the population a metric is calculated on. Rerun the pilot on uncleaned live records, and compare production results only against the same segment the pilot ran on, never against a blended total.
Pay-Per-Result appointments
See if we’re a fit
We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.
- 50,769+ appointments booked without cold calling.
- Pay-Per-Result pricing — you pay for booked, qualified calls.
- Pick your own time on our live calendar, no phone tag.
