Hacker News new | ask | show | jobs
by Maro 149 days ago
> This repo contains a version of Anthropic's original performance take-home, before Claude Opus 4.5 started doing better than humans given only 2 hours.

Was the screening format here that this problem was sent out, and candidates had to reply with a solution within 2 hours?

Or, are they just saying that the latest frontier coding models do better in 2 hours than human candidates have done in the past in multiple days?

2 comments

4 hours
Oh, I thought candidates got 2 hours but now I am confused too