Hacker News new | ask | show | jobs
by blintz 11 days ago
Are frontier models actually astronomically expensive to train? GLM 5.2 was trained on ~30T tokens, so ~10^25 FLOPs. Say you get B300's for $5/h (pretty high), and you get 50% MFU; that's ~$15M. Of course there's also a bunch of risk that the training itself goes badly, post-training, etc. But still, compared to the inference spend after, it's not that crazy.
1 comments

Nobody spends 15M on inference only to check at the end that their money was wasted.
Are there resources I could read to learn more about how the labs work?

I am curious about what changed since GPT 4.5 and other unsuccessful attempts to train large models, compared to now.