Hacker News new | ask | show | jobs
by intrepidkarthi 30 days ago
Most useful comment in the thread — participant-level selection is exactly what METR's update flags as the reason their new data is weak. Curious which direction the filtering ran: "AI won't help here" and "I don't want to do this one manually" corrupt the estimate in opposite directions.