The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

10JUL2025replayed
one year on
researchMETR · Cursor · Anthropic · OpenAI

METR randomized trial finds experienced developers 19% slower when using AI coding tools

A rigorous real-world study from the non-profit METR finds that experienced open-source developers took 19% longer on real tasks when using early-2025 AI tools, contradicting their own beliefs and expert forecasts of large productivity gains.

SAN FRANCISCO — A randomized controlled trial published today by the non-profit research group METR finds that experienced open-source developers who used early-2025 AI coding tools took 19% longer to complete real tasks than when working without AI, even though the same developers believed AI made them roughly 20% faster.

The study recruited 16 developers from large open-source repositories averaging 22,000+ stars and 1 million+ lines of code, and assigned them 246 real issues — bug fixes, features, and refactors — randomly split between AI-allowed and AI-forbidden conditions. When AI was allowed, developers primarily used Cursor Pro with Anthropic’s Claude 3.5 and 3.7 Sonnet, frontier models at the time. Developers self-reported completion times while being paid $150 per hour.

METR’s analysis rules out many experimental artifacts: developers used frontier models, complied with treatment assignment, submitted similar-quality pull requests in both conditions, and did not differentially drop hard issues. The researchers identify five likely contributing factors, including time spent prompting AI, waiting for responses, and AI struggles with large, complex codebases.

The result clashes sharply with developer self-assessments: participants forecasted a 24% speedup before the study, and even after experiencing the slowdown, still believed AI had sped them up by 20%. It also runs counter to benchmark scores on tests like SWE-Bench Verified and widespread anecdotal claims of AI productivity gains. METR explicitly says the study does not show AI never helps — only that in this particular setting with experienced developers on familiar codebases, the measured effect is negative.

The research group, which studies whether AI systems might threaten catastrophic harm, plans to repeat the methodology to track changes over time as AI tools evolve. The paper notes that learning effects from using tools like Cursor may only appear after hundreds of hours, and that AI capabilities may be lower in settings with very high quality standards. In the developer community, the finding is already fueling debate about whether benchmarks and anecdotal reports systematically overestimate real-world AI usefulness.

M
METR researchers (Joel Becker, Nate Rush, Beth Barnes, David Rein)

The researchers state they were broadly expecting to see positive speedup and that scientific integrity compelled them to share results regardless of outcome. They caution against overgeneralizing: 'We do not provide evidence that AI systems do not currently speed up many or most software developers.'

view the original post →
One year later — open only if you can handle spoilers

The 19%-slower headline made this one of the most-cited empirical checks on AI-coding hype through 2025 and into 2026, repeatedly invoked against benchmark-driven productivity claims. When METR revisited the cohort in early 2026, it found developers had grown so reliant on AI that a clean no-AI control group was no longer feasible; its revised estimate pointed to a modest speedup of roughly 18% but with a very wide confidence interval (about -38% to +9%), and in February 2026 the group announced it was redesigning the experiment. A separate METR survey in May 2026 found technical workers self-reporting far larger gains than measured studies could confirm — leaving the perception-versus-measurement gap this study exposed unresolved.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy