|
|
|
|
|
by intrepidkarthi
30 days ago
|
|
Most useful comment in the thread — participant-level selection is exactly what METR's update flags as the reason their new data is weak. Curious which direction the filtering ran: "AI won't help here" and "I don't want to do this one manually" corrupt the estimate in opposite directions. |
|