Hacker News new | ask | show | jobs
by dofm 30 days ago
Their followup study essentially says the followup study itself is possibly broken because developers will now not participate in some of the non-AI tasks and because the study pays less.

I would not, at all, suggest that this second study corrects or debunks the first.

Instead what it shows (if anything, i.e. if you can even put aside the regrettable choice to change the payment level, which affects applicant recruitment) is that the mindset shift has already happened: developers now don’t want to attempt some tasks without AI.

What that tells you is not (with any confidence at least) that they are faster, but perhaps that we are beyond the point that this can be meaningfully measured. AI could still be making developers slower, but developers aren’t going to be willing or perhaps able to help you find out.

Basically the job is different now.

What this does for me, perhaps, is vindicate my feelings. I can do agentic coding; I have learned the principles and some tools and I could learn more. But if this study is really reflective of how other developers feel now, I am done.

1 comments

The original study itself had at least one developer who later revealed that he had filtered out tasks he prefered not to do without AI: https://xcancel.com/ruben_bloom/status/1943536052037390531 -- given the N was 16, and he seems to have been one of the more AI-experienced devs, and we don't know if the other devs did this, the results of the first study itself could be questioned.
Most useful comment in the thread — participant-level selection is exactly what METR's update flags as the reason their new data is weak. Curious which direction the filtering ran: "AI won't help here" and "I don't want to do this one manually" corrupt the estimate in opposite directions.
I am not at all suggesting the first study is good, or that I believe its conclusions.

(Or that the failure of the second study validates the conclusions of the first.)

I am just saying that people here who think the second study overturned, debunked or corrected the findings of the first are explicitly wrong, because even its authors admit it is a broken study.

It would take a non-broken study to do that, and it may not actually be possible anymore, which is perhaps the most useful finding of the second study.