Hacker News new | ask | show | jobs
by m_ke 2 days ago
On Policy Self Distillation and Active Learning. Anything that increases sample efficiency by providing a richer more dense feedback signal and is more efficient at exploration / sampling.