|
|
|
|
|
by snthpy
19 days ago
|
|
Not just RLHF but also RLVR, and isn't that the litter lesson though? My sense of the Sutton Dwarkesh interview was that he was calling out that he didn't mean just longer datasets, but rather learning through exploration and that's exactly RL. |
|