one year on
MIT report finds 95% of enterprise generative AI pilots deliver no measurable P&L return
A data-backed study of 300+ deployments finds the failure is organizational, not technical — generic tools like ChatGPT stall because they cannot learn or integrate into workflows.
A new report from MIT’s NANDA initiative, published today, pours cold water on the narrative that enterprise generative AI is driving immediate returns. The study — an analysis of 300 public AI deployments, based on 150 interviews with leaders and a survey of 350 employees — finds that 95% of enterprise AI pilots deliver zero measurable impact on the P&L.
The failure, according to lead author Aditya Challapally, is not a model-quality problem but an organization one. Generic tools like ChatGPT excel for individuals because of their flexibility, but they stall in enterprise use since they don’t learn from or adapt to workflows — a phenomenon the report calls a “learning gap.” Only about 5% of organizations are getting a pilot across to real, measurable operational or financial impact, and the ones that do tend to buy from specialized vendors rather than build in-house. “Almost everywhere we went, enterprises were trying to build their own tool,” Challapally said, but the data showed building succeeds only one-third as often as buying.
The report also reveals a misalignment in resource allocation: more than half of generative AI budgets go to sales and marketing tools, while the biggest ROI comes from back-office automation, such as eliminating business process outsourcing and cutting external agency costs. Workforce disruption is already underway, mostly through attrition rather than layoffs, concentrated in customer support and administrative roles.
The report also highlights the widespread use of ‘shadow AI’ — unsanctioned tools like ChatGPT — and the ongoing challenge of measuring AI’s impact on productivity and profit.
The record
Lead author of the report, Challapally said successful startups 'have seen revenues jump from zero to $20 million in a year' by focusing on one pain point, executing well, and partnering smartly, while most enterprises fail due to a learning gap.
One year later — open only if you can handle spoilers
The "95%" became one of the most-cited — and most-contested — AI statistics of the era. Within days it was being blamed for helping feed a late-August 2025 wobble in AI-linked stocks and a wider "is this a bubble?" argument, even as critics (Futuriom among them) attacked the report's methodology, its shifting sample numbers, and its loose definition of "return." A year on the headline figure is still quoted in nearly every enterprise-AI ROI debate and still disputed; NANDA's deeper claim — that the bottleneck is organizational learning, not model quality — has aged better than the number that made it famous.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy