Hacker News new | ask | show | jobs
by adw 1 day ago
Other post-train mechanisms have meaningful bandwidth for introducing new information to the model (on-policy distillation is more or less the other extreme). RLVR aligns the model to particular behaviors it could already (unreliably, intermittently) express, it doesn't introduce new behaviors; the behaviors were latent in the pretrain/midtrain.