|
|
|
|
|
by Tadpole9181
12 days ago
|
|
The point of Arc-AGI-3 is to measure model performance. We already know that models can one-shot and iterate on very rudimentary game implementations. And, naturally, once it effectively has a copy of the source code, it can use that to play the game better. This harness is really moving the goalpost by defeating the entire point of the test. Instead of seeing the strength of a model's world view, its ability to internally derive and intuit rules, and its ability to keep track of game state over time, we're just letting the AI cheat. This is just the LLM equivalent of running a chess engine to the side. And this harness would not work in a remotely complex game and relies on the fact that Arc-AGI-3 is a focused test that only made the games as complicated as they needed to be for current model performance. |
|
The games are designed to allow assessment of a system. Knowing better systems to solve the games is a step forward. If any of the frontier labs could have one-shotted -3 in March with a custom harness, they would have done so.