Y
Hacker News
new
|
ask
|
show
|
jobs
by
wwind123
30 days ago
This kind of approach would generally still need human guidance, otherwise these models might get stuck in weird niche corners of the problem space that would not be relevant to any real world project.
1 comments
ben_w
30 days ago
We could call this "reinforcement learning from human feedback" (RLHF) :)
https://en.wikipedia.org/wiki/Reinforcement_learning_from_hu...
link
https://en.wikipedia.org/wiki/Reinforcement_learning_from_hu...