Hacker News new | ask | show | jobs
by torginus 18 days ago
Yes. Anyone who doesn't acknowledge the efficiency difference between pretraining vs RL and assume that since we've run out of data for the former, we have to do the latter, is not making a serious attempt at modelling the future:

https://www.tobyord.com/writing/inefficiency-of-reinforcemen...

This is similar to that other exponential, which happened with CPUs - we ran out of true geometric scaling in the mid 2000s, and everything else supporting Moore's Law has been cleverness that arrived in the nick of time, supported by a bit of marketing, and very optimizable benchmarks, far from guaranteed gains coming from making a single physical metric better.