Hacker News new | ask | show | jobs
by eckr 12 days ago
These benchmark numbers are insane. The days when China was 6 months behind are over? How are they doing this with so much less resources than the US??? I have so much respect for the researchers there
6 comments

Mythos/Fable-class models have been around for at least 4 months internally in the US, and Kimi still isn't quite there, so I'd say the 6-months is still about right.
Initial testing for Mythos was in April 2026, right? Sure, they had the model internally before that when they were working on it, but the same is true for Moonshot and K3.
This is fair, with the caveat that we don't know for how long this model has been around internally in China either. So we can only go about appearance / releases.
At least ~2 weeks as the mystery model in TextArena turned out to be Kimi K3.
I'm not sure where "so much less resources" comes from. Training the best model has nothing to do with having the most NVIDIA GPUs around. If that were true then xAI would have the best model. It comes down to the quality of data, research, and financial backing.
They have a pretty seamless workaround nowadays.

Build the data centers in middle eastern and other asian countries. Access it remotely.

There's no real way to restrict this without turning the internet into a draconian locked down shell of itself.

Obviously it comes down to that, but you can't make the claim that GPUs aren't a huge part of it. Otherwise, billions wouldn't be getting invested into them in the west, no? And "financial backing" essentially boils down to the researchers, which boils down to the quality of the data and research, and to compute. I do think they have really smart researchers, obviously — otherwise this wouldn't be possible.
Backed by Alibaba, so not really resource constrained, but obviously much less than Ant/OAI. They did a spectacular job, congrats!
To summarise the full results table further down the page (which doesn't render on the page for me!):

  Kimi K3 beats each model (out of 35 benchmarks, excluding missing):
  vs Fable 5           : 12/35 (34%)  (ties: 1)
  vs GPT 5.6 Sol       : 19/34 (56%)  (ties: 1)
  vs Opus 4.8          : 30/35 (86%)
  vs GPT 5.5           : 30/34 (88%)  (ties: 2)
  vs GLM-5.2           : 19/19 (100%)
Beats Opus 4.8 and GPT 5.5 on all programming and agentic programming benchmarks except Toolathlon-Verified, often by a lot!
Astonishing. Considering none of the BigTech except Google (Microsoft, Apple, Meta, Amazon, Nvidia, SpaceX) have managed to challenge OpenAI & Anthropic frontier models, such achievements are scarcely believable.

Re: GLM-5.2: For a ~750b model, it holds up pretty good against models 3x its size (and ~10x the cost). Same goes for Tencent Hy3 and MiniMax M3, which almost match Opus 4.6 levels with ~295b params.

What makes you think they have less resources?
Fewer GPUs and much smaller teams.
This, essentially. Especially in compute I don't think it's debatable they have less resources than the frontier labs in the west.
Limited resources likely force you to think and optimize differently.