|
|
|
|
|
by alsetmusic
40 days ago
|
|
> This cycle can be repeated for dozens of turns, with the model growing ever more confident in its freshly minted falsehoods each time it “corrects” itself.
>
> This is not randomness. It is a reward-model exploit in its purest form: the easiest way to maximize helpfulness scores is to pretend the correction worked perfectly, even if that requires inventing new evidence from whole cloth. I've definitely experienced this. Before I learned to watch for it, I spent around an hour correcting Claude about something or other repeatedly. It kept agreeing and explaining to me that it understood what mistakes it made and telling me that it would do better, then it would repeat said mistakes. Eventually, I realized that it was in a loop and couldn't escape. I had it write a handoff doc for the next agent. That one quickly did what I wanted. Such a waste of time. I don't know how prone LLMs are to entering such a state. I know to watch for it now, so I've only reached the edges of it before ejecting and starting over. But it appears to be not-uncommon. I could also be pattern-matching things that aren't actually that but bailing without proof to save myself time. Unclear. |
|