|
I think it might be even worse. LLMs seem to get tragically stuck on certain patterns. Maybe it's partly because a pile of weights essentially always starts from scratch in the same condition, but even within a single conversation, it will literally just latch onto words and repeat them incessantly, to the point where it becomes annoying. So for example, current Claude models love "honest". They are always producing "honest" assessments. "The honest caveat" - I'm sorry, did you mean the caveat, period? But also, use the wrong phrasing and suddenly you can create your own word of the day for an AI model. I used the word "analytical" once, in a conversation with Gemini 3 Pro. I am pretty sure every single response from that point on had "analytical" in it at least once. This is especially funny because system prompts and whatnot can also cause this behavior, but at least you can tweak those. You can't really do much about the model weights just having a weird affinity for a word. I bet someone will or probably already has come up with a way to detect and prevent these problems during training or post training. I'm not saying it's an easy problem, but it has the benefit that it really should be detectable with just statistics. |
> Honesty is a core aspect of our vision for Claude’s ethical character. Indeed, while we want Claude’s honesty to be tactful, graceful, and infused with deep care for the interests of all stakeholders, we also want Claude to hold standards of honesty that are substantially higher than the ones at stake in many standard visions of human ethics.
https://www.anthropic.com/constitution