|
|
|
|
|
by epolanski
37 days ago
|
|
Can you back up this with hard data and evidence? Most research converges to the idea that RL on synthetic data makes models worse, not better. If what you claim was anywhere near that relevant, than we would've long achieved singularity by simply feeding increasingly better output to the training of the next model in a loop. Yet this doesn't work. 25 million turns on Claude output is a small amount, yet an expensive one (we talking hundreds of $ millions) that is better spent on compute. There's no evidence such a process works, but I'd like to know more if I'm wrong. |
|
You are missing a mountain of nuance by generalizing the existence of a hole there.