Hacker News new | ask | show | jobs
by froh 19 days ago
rlhf = reinforcement learning from human feedback

(had to look it up)

1 comments

I think it's more RLVR (reinforcement learning from verified rewards). The RLHF is just to align models to human preferences, meaning to behave nice.
More accurate to say RLHF aligns models to human preferences, most significantly to be helpful.
What makes you say that