|
Considering incentive structures at play is solid epistemiology, but the line of thinking in your comment is a tad reductive, IMHO. In the hypothetical world where 1 is true, what different evidence do you expect to see than in worlds 2 and 3? If I were an unscrupulous OAI exec and wanted to opticsmaxx in this way, I wouldn't whip up a single, mild incident. Instead, I might burn gigatokens to 0day a few high-profile suppliers, and then have the model responsibly disclose those breaches. If we're willing to lie collude, and cheat, this story is easy to manufacture with at least as much credibility as the huggingface incident but with the advantage of looking way more impressive and spooking less regulators. And if I really were this evil exec, I would spend more than 30 seconds thinking up an even better strategy here. If, in contrast, we expect models to eventually breach honest and decent attempts at containment, then I'd exist something sorta like this huggingface story that looks like a combination of impressive and incompetent. I'm not sure whether I'd expect it to come out of a frontier lab or a partner or a consumer, though. > I'm inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. Forgive the Saucyness here, but impossible? Really? A security researcher that makes absolutist claims like these looks fatally naïve, IMHO. |
That's not to say I believe it outright, but people are being oddly dismissive and acting as if it's impossible to break out of a sandbox. Which we've seen time and time again that it absolutely can be.