Hacker News new | ask | show | jobs
by red75prime 4 hours ago
Yeah, I should have said RL, not RLVR. The point is RL interacts with external world, while autoregressive pretraining is limited to the passive ingestion, and RLHF relies on a model of human preferences that has no access to truth sources besides the limited training data it was built upon.