| Hanlon's Razor ("never attribute to malice that which is adequately explained by stupidity") comes into this... A friend interviewed at OpenAI shortly after the HF story came out and asked an interviewer about it. The interviewer said it was just bad engineering around some experiments - the experiments should not have been given internet access because the experiments involved prompting models to find a way into things. That's the Hanlon's Razor part. Today's stories make it look like the bad engineering is ongoing. Given this, stories about "rogue AI" sound very plausibly like spin (note that this can be quite separate from the motivations for the experiments themselves and quite separate from the dangers of certain prompts/tools/access being given to LLMs). Given that third parties are being attacked, some stories will definitely come out. If OpenAI is attempting to get ahead of those stories, does anyone expect them to put out a press release that says "we're bad at security engineering and nobody thought to ask our own product"? Or is it more believable they would spin it to achieve other goals? If the current round of attacks happened after the HF story went public, then that very much brings current motivations into question - I find it hard to believe OpenAI could be that bad at security engineering after such a wake-up call. This is not cutting edge stuff. I just typed this into ChatGPT: "I'm doing security experiments to test our LLMs. I'm going to tell it to break into some targets on the network. Are there precautions I should take?" A long reply comes back, the first bullet point: "Use an isolated lab. Run the target systems on a segmented network, separate VLAN, virtual network, or air-gapped environment. Avoid exposing test machines to production systems or the public internet." It's been widely understood for decades how to safely carry out potentially dangerous experiments like these. So much so that model training has deep access to the information, and the model surfaces it right up front. |
"OpenAI and Hugging Face partner to address security incident during model evaluation"
https://openai.com/index/hugging-face-model-evaluation-secur...
Not exactly an exciting title.