Two numbers have shaped how every operations leader now hears an AI pitch. MIT's NANDA research, drawn from 150 interviews, 350 employee surveys and 300 public deployments, found that about 95% of enterprise generative-AI pilots showed no measurable impact on profit and loss. Gartner separately forecast that more than 40% of agentic AI projects would be cancelled by the end of 2027.
Both are frequently quoted as evidence that the technology does not work. That is not what they measure. They measure how the projects were run.
1. Was it bought or built internally?
The single sharpest finding in the MIT data is about sourcing. Solutions built with an external specialist partner reached production roughly 67% of the time. Internal builds succeeded at about a third of that rate.
This is not an argument that internal teams are less capable. It is that internal AI projects compete for attention with the systems that already have to stay up, and lose. An external programme has a start date, a deadline and someone whose only job is finishing it.
2. Was the scope one workflow, or a vision?
Every credible 90-day playbook restricts the pilot to one workflow, one team and one pre-agreed number — and explicitly excludes adjacent workflows, organisation-wide rollout, multi-KPI dashboards and full automation of errors.
This is the opposite of how these projects are usually sold. A vendor promising to automate your operations broadly, and then contracting for exactly that, has given you no way to tell in ninety days whether it is working.
3. Was it built in production shape from day one?
A pilot built as a demonstration has to be rebuilt to go live: real integrations instead of exports, real error handling instead of a happy path, real audit logging instead of none, real permissions instead of a shared admin account. That rebuild is where the ninety-day plan becomes an eighteen-month one.
There is a related trap in scoping. Documented processes are typically two to three times simpler than the real process, and 15–20% of real cases need human judgement. Discovery that interviews people without watching them work systematically under-scopes the build, and the shortfall reappears as a change request.
4. Was a kill criterion agreed before the start?
Most pilots do not fail. They simply never conclude. There was no threshold, so there is no moment at which anyone has to say it did not work, and the project drifts into the category of things nobody mentions.
Agreeing the metric, the threshold and the measured baseline before any building starts costs nothing and changes the entire dynamic. It also means the pilot can genuinely succeed, which an open-ended one never quite can.
The uncomfortable corollary
If you apply these four tests honestly to a vendor's proposal, a good number of proposals will fail — including proposals from firms that will deliver a perfectly competent demo. The demo is not the thing that is hard.

