Hacker News new | ask | show | jobs
by bonoboTP 8 hours ago
There is RLHF with a model that's trained to emulate a human evaluator, but the evaluator is not really trained jointly with the main model to adapt to its distribution and tell it from real text. Though I'm sure there are some niche cases when this is done. But definitely not a prominent thing.