Y
Hacker News
new
|
ask
|
show
|
jobs
by
spongebobstoes
5 days ago
this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt
test Codex, not Sol. test Claude code, not Opus
1 comments
ChrisLTD
5 days ago
There are other benchmarks for that
link