|
|
|
|
|
by Danau5tin
14 days ago
|
|
I am also curious! The inner-RL-loop models are only trained once, then disgarded. But the outer-RL-loop model is trained on the same tasks over and over again. I imagine it would overfit after many more steps, but perhaps with a larger set of diverse tasks, the model would simply improve. |
|