Hacker News new | ask | show | jobs
by gertlabs 45 days ago
GLM 5.2 is the first model we've tested that is unambiguously on par with, or better than Opus 4.6 (although as usual, we have GLM 5.2 and most other Chinese models a bit below most other benchmarks with more vulnerable test methodologies).

Data at https://gertlabs.com/rankings

1 comments

I really have to take your score with a grain of salt because Opus 4.5 does better than Opus 4.6
They're within confidence intervals of each other, but remember how much discussion there was that Opus 4.6 had been nerfed in March. We averaged samples over the entire lifetime of Opus 4.6, which likely served many different underlying checkpoints. Even the best version of Opus 4.6 was hardly an upgrade.

We find a lot of interesting anomalies with our benchmark that hold up under large sample sizes.