|
|
|
|
|
by jimsimmons
1196 days ago
|
|
In RL case it’s rewarded because the supervision signal is generated post hoc. You can do the same with pure supervised learning and no RL. HF is the key, not RL. Yoav misses the nuance John had. RL is not bringing something fundamental to the table. It is just a better way to do things at the moment |
|