The 95% Was Never a Measurement
The most-quoted number in enterprise AI comes from a single yes-or-no question. Almost nobody who quotes it has read the appendix.
Ninety-five percent of AI pilots fail. You have seen it on a slide this month. It carries MIT’s name, and that is usually enough.
Here is where it comes from.
One question. The MIT NANDA report prints its interview script in the appendix. Question 12: “Have you observed measurable ROI from any GenAI deployment?” One yes-or-no question, put to one executive, about their whole company, answered from memory. The 95% is the share who said no.
Nobody examined a pilot. Nobody counted one.
Everyone wrote their own exam. Question 13 asks which metric was used — and lets each executive pick their own. Cost savings, productivity, retention, whatever came to mind. The report’s own limitations concede that success metrics “vary significantly across organizations and industries, limiting direct comparisons.” The most-quoted number in enterprise AI compares things its authors say cannot be compared.
The authors flagged it themselves. One limitation reads that the six-month window “may be insufficient,” and is “potentially understating success rates.” The researchers said their own failure rate was probably too high. Almost nobody has quoted that sentence.
Now the reckoning.
MIT gets a pass. They labeled it Preliminary Findings, version 0.1. They published the sample limits. They printed the questionnaire. Fifty-two interviews is a good way to generate hypotheses, and they never claimed it was anything else. You cannot fault researchers for being transparent about work other people chose to misuse.
Fortune knew better. On 18 August 2025 the headline read: “MIT report: 95% of generative AI pilots at companies are failing.” The report said 95% of organizations were getting zero return. Organizations are not pilots. One word, swapped for traffic, and a caveated finding became a fact.
Consultants and vendors are the real problem. They took the headline, never opened the PDF, and built decks on it. It is a perfect sales number: big, alarming, attributed to MIT, and pointing straight at whatever they happened to be selling. None of them needed four pages of appendix to see it would not hold. They needed it to hold.
The companion statistic is worse. Eighty-eight percent of agent pilots never reach production, we are told. It also circulates as 78, 80 and 89 percent, credited to four different sources. That is not a finding. It is a rumor with a decimal point.
None of this means enterprise AI is going well. Much of it is not. But the useful question is the one nobody asked: what were these pilots trying to achieve, and did they achieve it?
If you have put the 95% on a slide in the last year, you owe it ten minutes. The appendix is four pages.
Don’t take my word for it
Skip my analysis. Paste this into whichever AI you already use, and read what comes back:
Find the MIT Project NANDA report "The GenAI Divide: State of AI in Business 2025" (version 0.1). Read Appendix sections 8.2 and 8.3.
Answer using verbatim quotes from the report, not paraphrase:
1. What is the exact wording of the interview question that produces the 95% figure?
2. Does the 95% refer to pilots, or to organizations?
3. How were "success" and "ROI" defined and measured, and over what period?
4. Did each respondent choose their own success metric, or was one imposed?
5. Quote every stated limitation in full.
6. What was the total sample, and is any sampling frame or response rate given?
Then tell me what this study can and cannot support.
If it comes back with anything materially different from what you have just read, I want to know.
You cannot sell rigor with a number you never checked.