AI Transformation|Ravi Kiran, VP of Engineering|October 9, 2026|9 min read

Why AI Pilots Stall Before Production

MIT, IDC, S&P Global, Gartner and McKinsey all measured the same gap between AI pilots and production. What the numbers say, and the five places pilots die.

Why do AI pilots stall before production

We’ve spent the last few years being called in after a pilot stopped moving. The pattern is consistent enough that we can usually guess where it stopped before we see the code. The industry research says the same thing with bigger samples, so here is what it actually measures, where the numbers are weaker than the headlines suggest, and what we do about it.

The numbers, side by side

SourceSampleFinding
McKinsey, State of AI 2026 (Aug 2026)1,719 respondents, 97 countries37% report any EBIT impact from AI, flat on 2025. About 6% are “high performers” (5%+ of EBIT from AI), also flat
McKinsey, same surveySame40% of $1B+ companies are scaling AI agents, up from 27%. Smaller organisations: flat at 22%
MIT NANDA, The GenAI Divide (Jul 2025)150 interviews, 350 employees surveyed, 300 public deployments95% of organisations saw no measurable P&L return from GenAI. Partner-built tools reached deployment about 67% of the time, internal builds about a third as often
S&P Global Market Intelligence (2025)1,000+ respondents, North America and Europe42% of companies abandoned most of their AI initiatives, up from 17% in 2024. The average company scrapped 46% of proofs of concept
IDC with Lenovo (2025)~3,000 IT and business leadersFor every 33 AI proofs of concept, 4 reached production
Gartner (Jun 2025)ForecastMore than 40% of agentic AI projects will be cancelled by the end of 2027 over cost, unclear value or weak risk controls

Before you quote any of these

The 95% figure from MIT gets repeated as “95% of AI projects fail.” That’s not what it measured. It measured whether organisations could point to profit-and-loss impact from generative AI over a short window in early 2025, and the partner-versus-internal comparison is a correlation in self-reported data, which the authors said themselves. S&P’s 42% counts companies that abandoned most initiatives, which includes some that were sensibly killing bad ideas. Gartner’s number is a forecast.

So don’t build a board deck on any single figure. Look at what they agree on. Five different methods, five different samples, and every one of them lands in the same place: a large majority of AI work never reaches the point where it changes how the business runs. McKinsey’s 2026 data adds the uncomfortable part. Adoption keeps rising (44% of organisations now say AI is scaling across the enterprise, up from 38%), but the share seeing EBIT impact hasn’t moved. More AI is not turning into more results.

Five places pilots die

This is the field view, from the engagements we’ve been brought into. None of these is about the model.

1. The pilot ran on a clean extract

Someone exported six months of data into a tidy file, the model did well on it, and everyone celebrated. Production data lives in four systems, arrives late and has a field called “notes” that holds half the information. The fix is unglamorous: build the pilot against live connections from week one, even if that makes the demo slower.

2. The workflow stayed the same

Teams bolt AI onto the existing process and wonder why nothing changes. McKinsey found that nearly three-quarters of high performers fundamentally redesigned workflows around AI, compared with about a quarter of everyone else. That’s the single biggest gap in the survey. If the AI drafts a document that still goes through the same five approvals, you’ve saved the drafting time and nothing else.

3. Nobody owns it after the demo

The pilot had a sponsor, usually from IT or innovation. Production needs an owner in the business, someone whose numbers improve if it works and who will chase adoption when the novelty wears off. If you can’t name that person before the pilot starts, don’t start the pilot.

4. Security and permissions show up in month four

The pilot ran with broad access because it was a pilot. Then the security review arrives, asks who can see what, and the architecture has to be rebuilt around role-based access. S&P’s respondents named data privacy and security among their top obstacles. Bring the security team in during week one. They are much friendlier about a design than about a finished system.

5. Success was never written down in money

“Improve productivity” can’t be defended in a budget review. Eighty percent of McKinsey’s respondents say AI improved their own productivity, yet only 37% see it in EBIT. Somewhere between the person and the P&L, the value leaks. Pick one number the business already tracks, such as hours per tender, cost per resolved ticket or days to close, and measure the baseline before you build anything.

Why mid-sized companies have it harder

The McKinsey split is worth sitting with. Large companies went from 27% to 40% scaling AI agents in a year. Smaller organisations didn’t move at all. Big companies have platform teams, data engineers and someone whose full-time job is AI governance. A $200 million manufacturer usually has a capable IT team that is already fully booked.

That’s also why MIT’s partner finding matters for the mid-market, even with its caveats. If you don’t have a bench of engineers who have taken AI to production before, borrowing one is often faster than building one. The point isn’t outsourcing for its own sake. It’s that someone in the room should have seen these five failure points before.

A production-readiness test for any AI pilot

Run this before you approve a pilot, and again before you approve production. If you can’t answer yes to at least eight, you’re funding a demo.

#QuestionWhy it matters
1Is there a named business owner who isn’t in IT?Adoption needs someone whose numbers depend on it
2Have we written down one metric and its current baseline?You can’t show impact against a number you never measured
3Does the pilot use live data connections, not an extract?Clean samples hide most production problems
4Has security reviewed the access model?Late reviews force rebuilds
5Have we mapped the workflow as it will run with AI in it?Bolted-on AI saves minutes, not money
6Do we know what happens when the AI is wrong?Every system needs a human fallback and an escalation path
7Is there a monitoring plan for accuracy and cost after launch?Token and running costs now constrain 1 in 5 companies
8Do the people who’ll use it know it’s coming, and have they tried it?Surprised users become resistant users
9Is there a date when we decide to scale or stop?Pilots without a deadline drift for quarters
10Does someone on the build team have prior production deployments?Experience is the cheapest risk control there is

The cost point in question 7 comes from the same McKinsey survey: about 20% of respondents said AI operating costs, including tokens, had limited their AI use. That number will rise as more systems move into production.

What to do in the next 30 days

If you have a pilot that’s been “nearly ready” for a while, don’t start another one. Take the stalled pilot, run it through the ten questions above, and find which of the five failure points it hit. In our experience it’s usually two or three of them at once, and usually the same two: no business owner and no baseline. Both can be fixed in a fortnight without touching the code.

Frequently asked questions

What percentage of AI projects fail?

It depends what you count as failure. IDC found 4 of every 33 AI proofs of concept reached production. S&P Global found 42% of companies abandoned most AI initiatives in 2025. McKinsey’s 2026 survey found 37% of organisations report any EBIT impact from AI.

Is the MIT 95% AI failure statistic accurate?

It’s accurate for what it measured: organisations reporting no measurable P&L impact from generative AI in early 2025, in a sample of interviews, surveys and public deployments. It doesn’t mean 95% of AI projects technically failed.

Why do AI pilots fail to scale?

Mostly because of integration with real data and systems, workflows that were never redesigned, missing business ownership, late security reviews and success that was never defined in financial terms. Model quality is rarely the main cause.

How long should an AI pilot take to reach production?

Set a decision date before the pilot starts. In our engagements, going from the first workshop to production typically takes around six weeks.

Should we build AI in-house or use a partner?

Build in-house if you have engineers who have shipped AI to production before and the capacity to own it long-term. If not, MIT’s research found partner-built tools reached deployment about twice as often, though that comparison is correlational.

Ravi Kiran
About the author
Ravi Kiran
VP of Engineering, KnackLabs

Ravi Kiran is the VP of Engineering at KnackLabs and a LangChain Ambassador. He writes The Last Mile, a series on what it takes to get AI systems from pilot to production.

Have a pilot that’s stuck?

We’ll run the 10-question test with you in 30 minutes.

Book a 30-min call →

Sources

  1. McKinsey, The state of AI in 2026: On the road to ROI (25 Aug 2026)
  2. Fortune on MIT NANDA, The GenAI Divide (18 Aug 2025)
  3. CIO Dive on S&P Global Market Intelligence (14 Mar 2025)
  4. Gartner press release, agentic AI cancellations (25 Jun 2025)
  5. IDC with Lenovo, CIO Playbook 2025 (PDF)
  6. CIO.com on the IDC/Lenovo study (25 Mar 2025)
  • AI pilots
  • Enterprise AI
  • AI ROI
  • Production AI
←All field notes
ShareLinkedInX
Stay in the loop

New field notes, when we ship something worth writing about.

No cadence, no filler. Just the engineering and case studies as they go live.