We’ve spent the last few years being called in after a pilot stopped moving. The pattern is consistent enough that we can usually guess where it stopped before we see the code. The industry research says the same thing with bigger samples, so here is what it actually measures, where the numbers are weaker than the headlines suggest, and what we do about it.
The numbers, side by side
| Source | Sample | Finding |
|---|---|---|
| McKinsey, State of AI 2026 (Aug 2026) | 1,719 respondents, 97 countries | 37% report any EBIT impact from AI, flat on 2025. About 6% are “high performers” (5%+ of EBIT from AI), also flat |
| McKinsey, same survey | Same | 40% of $1B+ companies are scaling AI agents, up from 27%. Smaller organisations: flat at 22% |
| MIT NANDA, The GenAI Divide (Jul 2025) | 150 interviews, 350 employees surveyed, 300 public deployments | 95% of organisations saw no measurable P&L return from GenAI. Partner-built tools reached deployment about 67% of the time, internal builds about a third as often |
| S&P Global Market Intelligence (2025) | 1,000+ respondents, North America and Europe | 42% of companies abandoned most of their AI initiatives, up from 17% in 2024. The average company scrapped 46% of proofs of concept |
| IDC with Lenovo (2025) | ~3,000 IT and business leaders | For every 33 AI proofs of concept, 4 reached production |
| Gartner (Jun 2025) | Forecast | More than 40% of agentic AI projects will be cancelled by the end of 2027 over cost, unclear value or weak risk controls |
Before you quote any of these
The 95% figure from MIT gets repeated as “95% of AI projects fail.” That’s not what it measured. It measured whether organisations could point to profit-and-loss impact from generative AI over a short window in early 2025, and the partner-versus-internal comparison is a correlation in self-reported data, which the authors said themselves. S&P’s 42% counts companies that abandoned most initiatives, which includes some that were sensibly killing bad ideas. Gartner’s number is a forecast.
So don’t build a board deck on any single figure. Look at what they agree on. Five different methods, five different samples, and every one of them lands in the same place: a large majority of AI work never reaches the point where it changes how the business runs. McKinsey’s 2026 data adds the uncomfortable part. Adoption keeps rising (44% of organisations now say AI is scaling across the enterprise, up from 38%), but the share seeing EBIT impact hasn’t moved. More AI is not turning into more results.
Five places pilots die
This is the field view, from the engagements we’ve been brought into. None of these is about the model.
1. The pilot ran on a clean extract
Someone exported six months of data into a tidy file, the model did well on it, and everyone celebrated. Production data lives in four systems, arrives late and has a field called “notes” that holds half the information. The fix is unglamorous: build the pilot against live connections from week one, even if that makes the demo slower.
2. The workflow stayed the same
Teams bolt AI onto the existing process and wonder why nothing changes. McKinsey found that nearly three-quarters of high performers fundamentally redesigned workflows around AI, compared with about a quarter of everyone else. That’s the single biggest gap in the survey. If the AI drafts a document that still goes through the same five approvals, you’ve saved the drafting time and nothing else.
3. Nobody owns it after the demo
The pilot had a sponsor, usually from IT or innovation. Production needs an owner in the business, someone whose numbers improve if it works and who will chase adoption when the novelty wears off. If you can’t name that person before the pilot starts, don’t start the pilot.
4. Security and permissions show up in month four
The pilot ran with broad access because it was a pilot. Then the security review arrives, asks who can see what, and the architecture has to be rebuilt around role-based access. S&P’s respondents named data privacy and security among their top obstacles. Bring the security team in during week one. They are much friendlier about a design than about a finished system.
5. Success was never written down in money
“Improve productivity” can’t be defended in a budget review. Eighty percent of McKinsey’s respondents say AI improved their own productivity, yet only 37% see it in EBIT. Somewhere between the person and the P&L, the value leaks. Pick one number the business already tracks, such as hours per tender, cost per resolved ticket or days to close, and measure the baseline before you build anything.
Why mid-sized companies have it harder
The McKinsey split is worth sitting with. Large companies went from 27% to 40% scaling AI agents in a year. Smaller organisations didn’t move at all. Big companies have platform teams, data engineers and someone whose full-time job is AI governance. A $200 million manufacturer usually has a capable IT team that is already fully booked.
That’s also why MIT’s partner finding matters for the mid-market, even with its caveats. If you don’t have a bench of engineers who have taken AI to production before, borrowing one is often faster than building one. The point isn’t outsourcing for its own sake. It’s that someone in the room should have seen these five failure points before.
A production-readiness test for any AI pilot
Run this before you approve a pilot, and again before you approve production. If you can’t answer yes to at least eight, you’re funding a demo.
| # | Question | Why it matters |
|---|---|---|
| 1 | Is there a named business owner who isn’t in IT? | Adoption needs someone whose numbers depend on it |
| 2 | Have we written down one metric and its current baseline? | You can’t show impact against a number you never measured |
| 3 | Does the pilot use live data connections, not an extract? | Clean samples hide most production problems |
| 4 | Has security reviewed the access model? | Late reviews force rebuilds |
| 5 | Have we mapped the workflow as it will run with AI in it? | Bolted-on AI saves minutes, not money |
| 6 | Do we know what happens when the AI is wrong? | Every system needs a human fallback and an escalation path |
| 7 | Is there a monitoring plan for accuracy and cost after launch? | Token and running costs now constrain 1 in 5 companies |
| 8 | Do the people who’ll use it know it’s coming, and have they tried it? | Surprised users become resistant users |
| 9 | Is there a date when we decide to scale or stop? | Pilots without a deadline drift for quarters |
| 10 | Does someone on the build team have prior production deployments? | Experience is the cheapest risk control there is |
The cost point in question 7 comes from the same McKinsey survey: about 20% of respondents said AI operating costs, including tokens, had limited their AI use. That number will rise as more systems move into production.
What to do in the next 30 days
If you have a pilot that’s been “nearly ready” for a while, don’t start another one. Take the stalled pilot, run it through the ten questions above, and find which of the five failure points it hit. In our experience it’s usually two or three of them at once, and usually the same two: no business owner and no baseline. Both can be fixed in a fortnight without touching the code.
Frequently asked questions
What percentage of AI projects fail?
It depends what you count as failure. IDC found 4 of every 33 AI proofs of concept reached production. S&P Global found 42% of companies abandoned most AI initiatives in 2025. McKinsey’s 2026 survey found 37% of organisations report any EBIT impact from AI.
Is the MIT 95% AI failure statistic accurate?
It’s accurate for what it measured: organisations reporting no measurable P&L impact from generative AI in early 2025, in a sample of interviews, surveys and public deployments. It doesn’t mean 95% of AI projects technically failed.
Why do AI pilots fail to scale?
Mostly because of integration with real data and systems, workflows that were never redesigned, missing business ownership, late security reviews and success that was never defined in financial terms. Model quality is rarely the main cause.
How long should an AI pilot take to reach production?
Set a decision date before the pilot starts. In our engagements, going from the first workshop to production typically takes around six weeks.
Should we build AI in-house or use a partner?
Build in-house if you have engineers who have shipped AI to production before and the capacity to own it long-term. If not, MIT’s research found partner-built tools reached deployment about twice as often, though that comparison is correlational.

We’ll run the 10-question test with you in 30 minutes.
Sources
- McKinsey, The state of AI in 2026: On the road to ROI (25 Aug 2026)
- Fortune on MIT NANDA, The GenAI Divide (18 Aug 2025)
- CIO Dive on S&P Global Market Intelligence (14 Mar 2025)
- Gartner press release, agentic AI cancellations (25 Jun 2025)
- IDC with Lenovo, CIO Playbook 2025 (PDF)
- CIO.com on the IDC/Lenovo study (25 Mar 2025)
- AI pilots
- Enterprise AI
- AI ROI
- Production AI



