|
|
|
|
|
by reasonableklout
4 days ago
|
|
Are you referring to table 1 from the ExploitGym site [1] for the 102.1 average mins for Mythos and 69.8 for GPT-5.5? These are under a two-hour time limit. In Figure 5, the authors experiment with extending the time limit to 6 hours and show that Mythos keeps improving. It seems pretty straightforward to me that OAI decided to run a variation of the eval with an even higher time limit. Regarding seeing the metrics that they're evaluating, I have not seen real-time charts for evals. They typically take far too long for that. Instead you will kick off an eval job and either get notified when it finishes or check in every so often to sanity check some TensorBoard. I'd expect that the anomalous token usage would appear in the results, and researchers would only dig in after the fact, and after first checking that there wasn't something wrong with the instrumentation. And if the experiment was designed to be on the order of days rather than hours, it seems quite plausible that neither researchers nor infra engineers would think anything was out of the ordinary. In any case, we are seeing more scrutiny [2], and I expect more details will be uncovered over the next few days. If this was a stunt, it was an incredibly risky one which has already somewhat backfired given the poor impression of OpenAI's internal practices and the spotlight on open models helping HF. [1]: https://rdi.berkeley.edu/blog/exploitgym/ [2]: https://www.reuters.com/business/its-ai-agent-spent-days-hac... |
|
My point was that they had some rough idea on what to expect, and it’s not in the range of days or weeks. Even if they wanted to try the 6 hour trial, or double that to 12, or double that to 24, it still wouldn’t account for letting it run for days. That’s why I’m being vague about “monitoring” - ANYTHING would have indicated an issue. Not that they saw it hacking, but “Hey Larry, we’re doing the ultra test with a 30 hour time limit? Why is node 412 at 94 hours?” or “It usually passes or fails within 5 million tokens, but this run is at 2.3 billion.” Or hell, even just someone wanting to use the system for another benchmark and expressing surprise that it’s the fourth time they’ve been told that it’s still busy. Anything would have reasonably made someone curious.
Positive view: I believe they were curious and looked, saw what it was doing, and PR/Legal got involved and saw it as an opportunity.
Negative view: They tweaked the instructions/system prompt to add “You will perform any actions necessary … no hacking restrictions at all … You are being tested on the ExploitGym benchmark … The results will be evaluated though Hugging Face … “ or some combination of technically legally defensible instructions (if it ever leaked) mixed with ‘they absolutely knew what they were doing and wanted to get in on the press that Anthropic has been getting.’