Hacker News new | ask | show | jobs
by mediaman 25 days ago
They did use RLHF at the time, at which point it is not a pure probabilistic representation of the training corpora. Bizarrely, RLHF never came up in the paper.