|
|
|
|
|
by docjay
5 days ago
|
|
That was the inferred situation given what we know. It’s the preposterous framing that makes the story suspicious, but is necessary for the story to play out. 90 minutes to build a known exploit -> much much longer to create two zero-days and escape the sandbox then hack HF == No tracking of the time it worked on that one question. Average tokens required to complete the evaluation -> tokens required for two zero-days, network traversal, credential stealing, remote system hacking == No tracking of token usage EXPLODING at some point before it finished the whole benchmark. Etc, etc. |
|