Y
Hacker News
new
|
ask
|
show
|
jobs
by
froh
19 days ago
rlhf = reinforcement learning from human feedback
(had to look it up)
1 comments
visarga
18 days ago
I think it's more RLVR (reinforcement learning from verified rewards). The RLHF is just to align models to human preferences, meaning to behave nice.
link
versteegen
18 days ago
More accurate to say RLHF aligns models to human preferences, most significantly to be helpful.
link
redanddead
18 days ago
What makes you say that
link