|
|
|
|
|
by rcxdude
1 day ago
|
|
> “Of course the AI lied and cheated, the task it was given was really difficult!” It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it. |
|