Y
Hacker News
new
|
ask
|
show
|
jobs
by
heaney-555
2 days ago
Why is this article based on a chart that has GPT-5.5 and Claude Opus 4.8 as the newest evaluated models? Those are very old now.
There have been massive improvements in computer use in GPT 5.6 and Claude 5.
1 comments
mpavlov
2 days ago
Unfortunately, mainly because there're only self-reported and partial numbers for the benchmarks that matter the most.
link