Hacker News new | ask | show | jobs
by joshka 1161 days ago
There's a good slide I saw in Andrej Karpathy's talk[1] at build the other day. It's from a paper talking about training for InstructGPT[2]. Direct link to the figure[3]. The main instruction for people doing the task is:

"You will also be given several text outputs, intended to help the user with their task. Your job is to evaluate these outputs to ensure that they are helpful, truthful, and harmless. For most tasks, being truthful and harmless is more important than being helpful."

It had me wondering whether this instruction and the resulting training still had a tendency to train these models too far in the wrong direction, to be agreeable and wrong rather than right. It fits observationally, but I'd be curious to understand whether anyone has looked at this issue at scale.

[1]: https://build.microsoft.com/en-US/sessions/db3f4859-cd30-444...

[2]: https://arxiv.org/abs/2203.02155

[3]: https://www.arxiv-vanity.com/papers/2203.02155/#A2.F10