Hacker News new | ask | show | jobs
by rbuccigrossi 9 days ago
The others are an IQ test (TrackingAI), a test of Ph.D. questions across multiple domains (Humanity's Last Exam), and graphical pattern matching (ARC-AGI-2).

What's interesting is that while they are rather different in nature (yes it is odd that METR measures clock time as opposed to iterations etc.) but the behavior of the resulting improvement curves are extremely close.

That's the punchline: 4 independent measures point to the same conclusion. That increases the chance that the conclusion is correct.