Hacker News new | ask | show | jobs
by visarga 18 days ago
I think it's more RLVR (reinforcement learning from verified rewards). The RLHF is just to align models to human preferences, meaning to behave nice.
2 comments

More accurate to say RLHF aligns models to human preferences, most significantly to be helpful.
What makes you say that