Hacker News new | ask | show | jobs
by rcxdude 1 day ago
> “Of course the AI lied and cheated, the task it was given was really difficult!”

It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it.