Hacker News new | ask | show | jobs
by frabcus 12 days ago
Umm, Fable only really came out 2 weeks ago, and GPT-5.6 Sol only 1 week ago.

Yes, Kimi K3 appears a touch below them both, but above all other models. So I'd say a few weeks behind, not months now...

6 comments

Granted I've only used it for a few hours, but to me K3 still appears below Sonnet, even below 4.5.

I have their highest subscription, so it's not that I can't find uses for it and it has sides to it I really like, and it performs really well in some situations, but 2.7 also gets totally lost on tasks Sonnet and Opus has no problems with, and it looks like that is still the case with K3.

That said, I'm doing things with these models that are a lot more complex than the average app people will throw these models at, so I'm sure there are lots of use cases where it will perform better than what I'm seeing.

I don't know about Kimi, but in my experience other providers like z.ai and minimax don't serve the same models via their subscription plans as they do with api pricing.

They're clearly quantised - spelling mistakes, wonky thinking section dividers, missing whitespace are the immediately obvious tells but along with that comes degraded quality, and it seems to vary based on time of day.

I wouldn't be surprised if Kimi did similar things with their subscription plans - try using it through openrouter and see if you notice different behaviour.

Frontier labs release frontier models to the public only if there is market pressure to do so. Anthropic is not even hiding that they have been using Mythos internally for months now.

I wouldn’t be surprised if OpenAI (so much for “open”) is using GPT-6 internally already.

It appears that peasants like us are not going to get access to frontier AI anymore at any price.

> Anthropic is not even hiding that they have been using Mythos internally for months now.

This would be more impressive if their software and delivery quality was higher.

Anthropic had Mythos-Preview for many months internally, but from available sources was an active work in progress, and it seems they started releasing it via Project Glasswing to partners before the final checkpoint was available.
Maybe or maybe not. Anthropic made it a marketting thing.
very unlikely that they are holding models back. models are very quickly depreciating in value and have internal cutoff dates. that would really not make sense from a business perspective. and doing an entire training run without commercializing it also seems like a huge waste.
That assumes progress is linear, which it certainly is not. We’re also assuming Chinese companies are releasing their best stuff when our own government is dictating what American companies can and cannot release.

I don’t know for sure that they are weeks or months behind. I doubt anyone outside of 3 letter agencies knows that. The pace of AI is crazy fast and China is notoriously secretive. We could be comfortably ahead or China could have the top model by the end of the year.

I don’t think any of us have enough information to know what’s really going on, but I suspect it’s a very tight race.

Mythos/Fable was trained 6 months ago.
Should we compare release date to release date? Or training date to training date?
Normally I would say release date to release date, but Mythos/Fable release was significantly delayed because of security concerns.
Mythos/Fable release was significantly delayed because Dario is pushing for regulatory capture. It also plays into the hype because there was no one else releasing a model better than Opus. The same thing was with the scarcity crap. They said Fable will not run on subscription, then they will run just for a little while and then forever... Once you get the game they are playing it starts to become quite laughable. If you check his history he was doing the safety/terminator stuff since the days at OpenAI when the models couldnt even calculate 1+1 reliably
Is there any actual proof of this? Everything suggests it's a sincerely held belief and there's been no whistleblowers claiming otherwise. "Since before they could add 1+1" works against your argument because the models had no commercial value, no guarantee they would ever have commercial value, and plotting regulatory capture was about as serious an idea as regulatory capture of antimatter engines.
I'm not defending Dario. That's not the point I'm making.
All three models were quite likely undergoing training at the same time, with Ant a couple of months ahead. Glasswing was announced 3 months ago, Sol and Kimi were already being trained at that point. They are taking snapshots and doing experiments the whole way through anyway and all models will continue to be trained after release. Kimi 3.1 could be better than Fable 5.1
Yeah, I think the answer is somewhere in between. Anthropic has been a few months ahead of everybody in terms of internal capabilities.

So it's not 6 months but it's also not a few weeks.

If they were a few weeks behind, Moonshot would release a Fable level model in a couple of weeks.
I used fable more than a month ago, what are you talking about?
Both mythos and fable were available in a preview version first.

The preview models were not as good as the final release. Thus training must have continued after the initial announcement

Nah that’s not how that works. There are literally an endless stream of examples of the model providers tweaking backend stuff and improving prompt responses without using new or retrained data. And even if they added new training data to an existing model, it isn’t true to say that model ‘just came out’. But even then, when providers release better models based on new training data they release it as a new dot version, which boosts their sales from the announcement.

Fable has been out for more than a month - I didn’t have the preview version and I was using it around June 10th, when my Claude subscription expired. Saying “Fable only really came out 2 weeks ago” is just factually incorrect all around.