These benchmark numbers are insane. The days when China was 6 months behind are over? How are they doing this with so much less resources than the US??? I have so much respect for the researchers there
Mythos/Fable-class models have been around for at least 4 months internally in the US, and Kimi still isn't quite there, so I'd say the 6-months is still about right.
Initial testing for Mythos was in April 2026, right? Sure, they had the model internally before that when they were working on it, but the same is true for Moonshot and K3.
This is fair, with the caveat that we don't know for how long this model has been around internally in China either. So we can only go about appearance / releases.
I'm not sure where "so much less resources" comes from. Training the best model has nothing to do with having the most NVIDIA GPUs around. If that were true then xAI would have the best model. It comes down to the quality of data, research, and financial backing.
Obviously it comes down to that, but you can't make the claim that GPUs aren't a huge part of it. Otherwise, billions wouldn't be getting invested into them in the west, no? And "financial backing" essentially boils down to the researchers, which boils down to the quality of the data and research, and to compute. I do think they have really smart researchers, obviously — otherwise this wouldn't be possible.
To summarise the full results table further down the page (which doesn't render on the page for me!):
Kimi K3 beats each model (out of 35 benchmarks, excluding missing):
vs Fable 5 : 12/35 (34%) (ties: 1)
vs GPT 5.6 Sol : 19/34 (56%) (ties: 1)
vs Opus 4.8 : 30/35 (86%)
vs GPT 5.5 : 30/34 (88%) (ties: 2)
vs GLM-5.2 : 19/19 (100%)
Beats Opus 4.8 and GPT 5.5 on all programming and agentic programming benchmarks except Toolathlon-Verified, often by a lot!
Astonishing. Considering none of the BigTech except Google (Microsoft, Apple, Meta, Amazon, Nvidia, SpaceX) have managed to challenge OpenAI & Anthropic frontier models, such achievements are scarcely believable.
Re: GLM-5.2: For a ~750b model, it holds up pretty good against models 3x its size (and ~10x the cost). Same goes for Tencent Hy3 and MiniMax M3, which almost match Opus 4.6 levels with ~295b params.