Hacker News new | ask | show | jobs
by credit_guy 28 days ago
> A given model has a shelf-life, which these days is measured in months, not years.

Not all new models are trained from scratch. ChatGPT 5.3 to 5.4 (and likely 5.5) was basically the same model, but probably trained a bit more, not a new model from scratch.

> The "someday" when frontier model providers can enjoy their current high inference margins without the burden of significant training costs is never going to arrive.

That is debatable. I believe the moat for the frontier model providers is the compute. At the level of 10 trillion parameters (that Fable/Mythos are rumored to have), you need serious compute to serve inference, and you also need serious compute to train. Will DeepSeek, Qwen, Kimi, GLM come up with a 10T new model anytime soon? I doubt that. People keep saying that the Chinese labs are catching up to the US big 3, and measured in months the gap is now only 4-6 months. I doubt a Chinese version of Fable/Mythos will be released in the next 12 months.

1 comments

>ChatGPT 5.3 to 5.4 (and likely 5.5) was basically the same model, but probably trained a bit more, not a new model from scratch.

Then those models have an eroding moat and will be quickly driven down to commodity pricing. The only thing propping up inference margins are the cap-ex costs of training. That's the moat. That's why there's no way to win this game. You cannot have low training/infrastructure cost and high margins (such as would justify today's valuations).

>I doubt a Chinese version of Fable/Mythos will be released in the next 12 months.

I would take that bet.