Hacker News new | ask | show | jobs
by Agentlien 5 days ago
I would rather say that many details of the supposed attack were published by Hugging face before it was publicly announced that OpenAI were involved.

Both are big actors in the AI space who arguably benefit from increasing the perceived capabilities of AI models. If one suspects OpenAI of lying it isn't such a stretch to think this was a coordinated PR campaign between them and Hugging face.

1 comments

I can accept that AI labs themselves, like essentially no company before them, are overselling how dangerous their product is far marketing. It's weird how confident everyone is about that theory, but it does at least make sense.

But come on- Hugging Face benefits from increasing the perceived capabilities of OpenAI's models to slightly beyond Anthropic's? Enough to be cut in on this PR scam- to be handed the never-before-revealed information that this is a PR scam- despite having much less skin in the game than their partner here? And then they turned around and used a Chinese model to successfully stop it? This is a stretch!

Another possibility is of course that Hugging Face legitimately were attacked and wrote a completely honest response - but OpenAI instructed their AI to attack their servers and the breaking of containment is fiction.

I'm not convinced in any direction, really. But what makes me cautious is that there have been extraordinary claims from both OpenAI and especially Anthropic of their models breaking containment, hacking the host, etc. for several iterations of their products and I have only heard of this type of behavior from their own blog posts about how powerful and dangerous their upcoming models are. Never from anyone having it accidentally happen in production once they are released. It seems unlikely to me that the final post-training and safeguards are that bulletproof given how much use these tools are seeing.