Hacker News new | ask | show | jobs
by janalsncm 13 days ago
It is a huge improvement to PPO because you don’t need a separate critic model which cuts memory costs in half and stabilizes training.
1 comments

Yes, but monte carlo estimating the critic model is not new.