|
|
|
|
|
by jsw97
20 days ago
|
|
I am wondering about the author's allegation that there is a user filter, not just a prompt filter. Of course it could also be the case that it is just a prompt filter, but Fable sees memories from the authors' prior sessions that cause a rejection. I wonder if the author could control for this is in some way, if Claude lets you run isolated session without memory access. |
|
The obvious failure mode is that trying to fix an innocent prompt to pass an over-sensitive classifier looks like a bad actor trying to jailbreak the model. I don't really see how Anthropic can fix this. Jailbreaking is a fundamental weakness endemic to LLMs, so 'smarter' models aren't the answer.
I suspect they're being so stringent because, at least some at Anthropic, genuinely believe LLMs are already an existential risk to humanity. However, it's clear other frontier competitors rank that risk lower and are taking a more nuanced, pragmatic position on safety. To the extent Anthropic's fears continue to make them less useful to customers, competitors are going to bypass them.