|
|
|
|
|
by desterothx
15 days ago
|
|
According to Pat Toulme, the thinking traces and outputs that the Chinese researchers distilled are useful for getting initial trajectories, to prevent a cold start during RL. Once you get those initial correct trajectories (the model actually solving a task), you can generate you own new traces, so the distillation is already done, there is no need for further reasoning traces. Sure they could probably get better alignment with the frontier by distilling fruther, but in any case the damage is already done, at this point hiding reasoning is mostly just hurting users |
|