Hacker News new | ask | show | jobs
by throwa356262 4 days ago
"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration."

I guess there is no way this can happen without benchmark being part of the training data??

4 comments

What seems to implied is that some of his hold out testing suite includes simple/common tests that are out in the wild, and for those opus went straight to a memorised solution .
Simple/common tests is not an explanation for why only now Opus 5 is the only model encoding the answers like this. Something like the holdout test suite being leaked or Anthropic cheating (e.g. 'accidentally' including previous hold out run data in Opus 5 training) makes a much stronger fit.
No, I believe you are misunderstanding the quote. Each “question” in ARC-AGI-3 is a game that has hidden rules that you can understand if you look at the game board long enough. This quote means that Opus 5 is looking at the game board, figuring out the rules, and writing out the rules before it makes a single move. You can do the same thing if you go to the ARC-AGI-3 website and try some of the games.
They state the puzzle is "Witness-like" which I assume means that it follows the rules from the well-known puzzle game "The Witness" which Opus definitely knows.
It may read information about the benchmark, such as on blog post, or Twitter feeds (example OP) without active cheat