Hacker News new | ask | show | jobs
by ianbutler 30 days ago
2025 is such old news that this just isn't relevant.

METR already redid the study at a later date and now finds a likely 18% speedup

"For the subset of the original developers who participated in the later study, we now estimate a speedup of -18% with a confidence interval between -38% and +9%" (note their use of - and + here could be slightly confusing but they do mean 18% faster per the post)

https://metr.org/blog/2026-02-24-uplift-update/

7 comments

Their followup study essentially says the followup study itself is possibly broken because developers will now not participate in some of the non-AI tasks and because the study pays less.

I would not, at all, suggest that this second study corrects or debunks the first.

Instead what it shows (if anything, i.e. if you can even put aside the regrettable choice to change the payment level, which affects applicant recruitment) is that the mindset shift has already happened: developers now don’t want to attempt some tasks without AI.

What that tells you is not (with any confidence at least) that they are faster, but perhaps that we are beyond the point that this can be meaningfully measured. AI could still be making developers slower, but developers aren’t going to be willing or perhaps able to help you find out.

Basically the job is different now.

What this does for me, perhaps, is vindicate my feelings. I can do agentic coding; I have learned the principles and some tools and I could learn more. But if this study is really reflective of how other developers feel now, I am done.

The original study itself had at least one developer who later revealed that he had filtered out tasks he prefered not to do without AI: https://xcancel.com/ruben_bloom/status/1943536052037390531 -- given the N was 16, and he seems to have been one of the more AI-experienced devs, and we don't know if the other devs did this, the results of the first study itself could be questioned.
Most useful comment in the thread — participant-level selection is exactly what METR's update flags as the reason their new data is weak. Curious which direction the filtering ran: "AI won't help here" and "I don't want to do this one manually" corrupt the estimate in opposite directions.
I am not at all suggesting the first study is good, or that I believe its conclusions.

(Or that the failure of the second study validates the conclusions of the first.)

I am just saying that people here who think the second study overturned, debunked or corrected the findings of the first are explicitly wrong, because even its authors admit it is a broken study.

It would take a non-broken study to do that, and it may not actually be possible anymore, which is perhaps the most useful finding of the second study.

Either way, it's not a dramatic improvement. Thankfully I work in an environment where with little bureaucracy so my time is actually spent doing technical work.

I do think AI has been a huge boon to productivity in many ways, but looking at feature timelines, I think it's pretty clear the 'critical shortest path' of key features hasn't been sped up by that much.

They also say "Wider adoption of AI has made it more difficult to measure task-level productivity"

I think there is a simple reason for that. If you automate something, you make the measureable/predictable thing faster. So the hard to measure/predict part of the job will take more share of the time, and overall difficulty to measure/predict goes up.

I think this is what happened with Agile Scrum - as developers became more productive (for unrelated reasons, two main sources of SW developer productivity before AI were compilers and open source), the bureacracy (amount of meetings) increased, because the ratio of hard to measure vs easy to measure went up. Bureacracy is hard to measure, so it went up (as a share of work). I expect this only getting worse with more automation, such as AI. So I predict an increase in share of bureacracy compared to pre-AI world.

Either way, IMHO main point is automation has the opposite effect on human job predictability, it lowers it. Tasks we can easily automate are those that are easy to predict.

I've held this stance on agile for a long time - it coincided with mainstream adoption of ssds, windows with memory protection and google search - all of which sped delivery despite agile, not because of agile.
Even not touching the laughable sample size for both studies - almost halved sample size between 2025 and 2026? Sounds like a massive selection bias, and not in the way they're implying.
That post literally says the results are unreliable...
...and in particular it says that one of the reasons is that developers are refusing to participate in the non-AI branch, and when they do, changing what tasks they select to those where AI would be less useful.

Overall this suggests to them that the current speedup is likely greater than what the study could measure.

It might suggest that but they can’t back it up because the study is broken and perhaps forever unrepeatable.

Like, what people are saying is, “That old study was wrong! They did a new broken study that overturned it!”

Models might be better today, but the takeaway that persists is the delusion around things being magically faster.
And to be specific, the METR study was using the Cursor harness with Claude Sonnet 3.5/3.7, along with other models of that era of the participant’s choosing.

Which is ancient at this point, and half a year older than the November 2025 inflection point when agentic coding got really good.

The original article is from August 2025, and the overall message to not trust ‘how it feels’ and rather measure outcomes seems right to me despite the outdated figures. On my team at least, we are seeing a noticeable inflection in work shipped with AI according to Weave.