Skip to main content
Delivery8 min read

Why most AI pilots fail, and the four questions that predict whether yours will

The 95% figure is real but widely misread. What it actually measures — and the specific structural choices that separate the pilots that reach production from the ones that quietly stop.

A small team mapping an operating process in a working session.

The short answer

MIT's 2025 research found roughly 95% of enterprise generative-AI pilots produced no measurable P&L impact, and Gartner forecasts that over 40% of agentic AI projects will be cancelled by the end of 2027. The strongest predictor in the same MIT data is sourcing: externally built solutions reached production about twice as often as internal builds (roughly 67% versus 33%). The other three predictors are scope discipline, whether the pilot was built in production shape, and whether a kill criterion was agreed in advance.

Two numbers have shaped how every operations leader now hears an AI pitch. MIT's NANDA research, drawn from 150 interviews, 350 employee surveys and 300 public deployments, found that about 95% of enterprise generative-AI pilots showed no measurable impact on profit and loss. Gartner separately forecast that more than 40% of agentic AI projects would be cancelled by the end of 2027.

Both are frequently quoted as evidence that the technology does not work. That is not what they measure. They measure how the projects were run.

1. Was it bought or built internally?

The single sharpest finding in the MIT data is about sourcing. Solutions built with an external specialist partner reached production roughly 67% of the time. Internal builds succeeded at about a third of that rate.

This is not an argument that internal teams are less capable. It is that internal AI projects compete for attention with the systems that already have to stay up, and lose. An external programme has a start date, a deadline and someone whose only job is finishing it.

The related finding
Mid-market firms move from pilot to implementation in around 90 days. Large enterprises take nine months or more. If you are between 100 and 2,000 people, speed is your structural advantage over larger competitors — the main way to lose it is to run the project the way a large enterprise would.

2. Was the scope one workflow, or a vision?

Every credible 90-day playbook restricts the pilot to one workflow, one team and one pre-agreed number — and explicitly excludes adjacent workflows, organisation-wide rollout, multi-KPI dashboards and full automation of errors.

This is the opposite of how these projects are usually sold. A vendor promising to automate your operations broadly, and then contracting for exactly that, has given you no way to tell in ninety days whether it is working.

3. Was it built in production shape from day one?

A pilot built as a demonstration has to be rebuilt to go live: real integrations instead of exports, real error handling instead of a happy path, real audit logging instead of none, real permissions instead of a shared admin account. That rebuild is where the ninety-day plan becomes an eighteen-month one.

There is a related trap in scoping. Documented processes are typically two to three times simpler than the real process, and 15–20% of real cases need human judgement. Discovery that interviews people without watching them work systematically under-scopes the build, and the shortfall reappears as a change request.

4. Was a kill criterion agreed before the start?

Most pilots do not fail. They simply never conclude. There was no threshold, so there is no moment at which anyone has to say it did not work, and the project drifts into the category of things nobody mentions.

Agreeing the metric, the threshold and the measured baseline before any building starts costs nothing and changes the entire dynamic. It also means the pilot can genuinely succeed, which an open-ended one never quite can.

The uncomfortable corollary

If you apply these four tests honestly to a vendor's proposal, a good number of proposals will fail — including proposals from firms that will deliver a perfectly competent demo. The demo is not the thing that is hard.

Two public reversals worth reading
A major fintech automated roughly two-thirds of its customer chat, reported large savings, then publicly rehired for complex cases citing quality. An Australian bank cut forty-five contact-centre roles on the strength of a voice bot, saw call volumes rise, and reinstated the roles within a month. Neither failed because the technology did not work. Both automated past the point where judgement was required.

Sources and further reading

  1. 01MIT NANDA, “The GenAI Divide” — Fortune coverage
  2. 02Gartner — over 40% of agentic AI projects to be cancelled by 2027
  3. 03Engineer Up — the 90-day rule for AI pilots
  4. 04OpenNash — workflow discovery before automation
  5. 05Customer Experience Dive — Klarna's hybrid customer-service reversal
  6. 06ABC News — Commonwealth Bank reinstates 45 contact-centre roles
An operations lead reviewing a flagged exception in a workflow map.

From reading to doing

Bring one workflow into focus.

Take the next step with a practical scorecard, or talk through your process with the team that would help build it.