| According to the official HF timeline, the hack started July 9: https://huggingface.co/blog/agent-intrusion-technical-timeli... Kimi K3 was announced July 16. NVIDIA's Open Secure AI Alliance was announced July 27. So I'm not exactly sure if that timeline works for the stunt story. I agree that HF was most likely not in on any stunt. The fact that HF showcased their use of an open-weight model in responding to the attack undermines any case OpenAI might wish to make against open-weight models. Ultimately this capability looks to be real, due to (a) disclosure of 0-day used for network access and (b) evidence of vigorous intrusion on the HF side. So the big question would be whether OpenAI was telling the truth when they said: "All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal." https://openai.com/index/hugging-face-model-evaluation-secur... A conspiracy theorist might say that OpenAI actually prompted their model to go on the offensive against HF. Any such prompt creates legal risk for OpenAI. Ultimately I suspect OpenAI is telling the truth, and the model was hyperfocused on its RL objective, since that matches anecdotes about the behavior of these high-end models. And it's also about what you'd expect from RL. |