Hacker News new | ask | show | jobs
by blovescoffee 13 days ago
"Current GPT/Claude/Gemini" is not a meaningful statement about perf. There's many different models from each of those providers and there's a considerable gap between the best of anthropic and open ai compared to gemini.

Benchmarks have GLM 5.2 somewhere underneath Sol and Fable and closer to now last-gen openai and anthropic models.

2 comments

One error: GLM 5.2 beats the best public Gemini model, 3.5 pro.

There's 2 caveats with the rest. First, GLM 5.2 matches those models in "xhigh" effort modes, which has a very low quota on the subscriptions, especially for Claude.

Second, last-gen GPT/Claude means what they release in April/May of 2026. Or to be even more complete/fair:

GLM 5.2 beats what OpenAI released in March 2026 (GPT 5.5 xxhigh), and what Anthropic released in April 2026 (Opus 4.7 xhigh). It is beaten by what OpenAI released in April of 2026 (GPT 5.6 Sol xxhigh) and Anthropic released in May 2026 (Opus 4.8 (the same as "Fable" ?), xhigh effort)

GLM 5.2 was released on Jun 16 and if OpenAI and Anthropic hadn't done those quick releases they would have been beaten on their best available models ...

So great news! Open source now has SOTA performance 3 months after OpenAI/Anthropic/Google. Wow.

The gap has been steadily closing over time.

Opus 4.8 (May) to Kimi K3 (July) has apparently just dropped it to two months.

China also does efficiency improvements. Qwen 3.6 27B is better than Sonnet 4.5 and you can run it on a couple of gaming video cards. That's incredible. I can do real actual work with this!

As Google said in 2023, none of them have a moat, open weight models will win.

> As Google said in 2023, none of them have a moat, open weight models will win.

Google, who probably canceled the release of Gemini 3.5 pro to avoid having their best new model perform worse than BOTH GLM 5.2 AND Kimi K3?

I mean how is this anything but an incredible defeat of Google's supposed inventiveness?

My previous comment from 24 hours ago is now irrelevant.

Newly released Kimi K3 is benching better than Claude Opus 4.8. The only better models are Claude Fable and GPT 5.6 Sol Max Effort.