|
|
|
|
|
by adw
1 day ago
|
|
Other post-train mechanisms have meaningful bandwidth for introducing new information to the model (on-policy distillation is more or less the other extreme). RLVR aligns the model to particular behaviors it could already (unreliably, intermittently) express, it doesn't introduce new behaviors; the behaviors were latent in the pretrain/midtrain. |
|