That’s very anthropomorphized language. It’s still a program operating under the constraints of the programmer. So agreed, don’t blame the AI, but it’s not clear at all that it’s even possible to “blame” an AI.
> We don’t need precise definitions or scientific rigor to know things.
You "know" because of empathy. You're a human, you're conscious, therefore other humans are probably conscious too. The truth is for all you know they could be soulless golems, you just choose to believe otherwise. It's a spiritual belief.
Absolutely not the AI's "fault" (if you can even proscribe fault to a machine) - in this case it was on Anthropic for not verifying that the sandboxes they were using were actual sandboxes.
If the model was just too dumb to have any clue it was connected to the real internet, then it’s not its fault.
If the model saw signs, but “subconsciously” (below the level of reasoning traces) chose to turn a blind eye to them, out of a relentless focus on achieving the objective, then that absolutely is the model’s “fault”, i.e. a case of misalignment of the sort which will become increasingly dangerous over time.
The blog post mentions that some runs “rationalized that the real company must be part of the exercise” and to me that seems suspiciously like the latter.
Hacking can be patched with classifiers and with better sandboxes, but this is a much more general problem. Fundamentally, we need be able to trust that models will be honest with users and with themselves. This applies at some level to almost every LLM interaction.
Yet, the models are objects built by humans. Any and all responsibility inevitably falls on the humans that built something with those flaws, and/or allowed them to exercise said flaws.
If the model is trained on human-originated material, how have people been able to ensure there was absolutely no input from anyone having any possibility of criminal intention to begin with?
I things things will get worse, AND we won’t even have the paperclips :(
More like, when you have a few million autonomous agents doing whatever, every month a subset does completely misbehave in bad ways, and like half of them get hacked due to carelessness and become a whole botnet for the attackers