|
|
|
|
|
by Maro
149 days ago
|
|
> This repo contains a version of Anthropic's original performance take-home, before Claude Opus 4.5 started doing better than humans given only 2 hours. Was the screening format here that this problem was sent out, and candidates had to reply with a solution within 2 hours? Or, are they just saying that the latest frontier coding models do better in 2 hours than human candidates have done in the past in multiple days? |
|