Hacker News new | ask | show | jobs
by heaney-555 2 days ago
Why is this article based on a chart that has GPT-5.5 and Claude Opus 4.8 as the newest evaluated models? Those are very old now.

There have been massive improvements in computer use in GPT 5.6 and Claude 5.

1 comments

Unfortunately, mainly because there're only self-reported and partial numbers for the benchmarks that matter the most.