Hacker News new | ask | show | jobs
by scrollop 12 days ago
Nah:

https://www.youtube.com/watch?v=LSlV206xPqM

These real world examples show it's one tier away.

2 comments

If Chinese AI companies can train a model that's slightly worse than the frontier, then there's no reason why they can't train a model that is slightly better than the frontier.

Everybody can agree that K3 doesn't clearly surpass Fable. However, inevitably there will be a time in the future when a Chinese AI company releases a model that's better than any US model.

K3 isn't the knockout blow but it's the 2nd knockdown that makes everyone in the arena realize that the fighter is not winning the fight.

Anthropic is arguably still better in tooling and integrating model and tooling. Good habit beats raw intelligence.

For code editing Cursor editor tooling is even better.

These "real world" examples are nothing like the way I use LLMs from within a harness. GPT 5.6 Sol and Fable are clearly more impressive, but how does this translate to interactive agent use, or use under an agent orchestration framework?
This is a question I am going to get an answer tomorrow with evals. Extremely interesting...