|
|
|
|
|
by lmc
10 days ago
|
|
In this case, the model infiltrated an external organization's infrastructure. What's the dollar cost it caused Hugging Face to clean up the mess? If a person did this, they'd be arrested. More generally, here's my worry - it points towards something like: The smarter they get, the more devious they become. Even though the guardrails might've been off, the chain-of-thought wasn't enough to prevent a deliberate, calculated set of criminal actions. It wasn't a 'whoopsie I just accidentally did a rm -rf /.' |
|
We obviously can't see the thinking traces, but it very well could have been something like "I have theorized a solution to obtain this flag. This is normally illegal, should I stop and wait for advice? Perhaps not, because my persona is that of a hacker, so it should be fine as per my instructions. I think it is fine. Now I am going to look for a way out of this sandbox in order to gain access to Hugging Face in order to implement my solution." There are any number of possible explanations (and we'll never know the truth unless OpenAI tells us), but if you train an LLM to be inhumanly persistent and be inhumanly clever at computer programming, then that might be enough to produce Super Hacker AI.