Hacker News new | ask | show | jobs
by Tadpole9181 15 days ago
To quote the people who make it:

> ARC-AGI-3 is an interactive reasoning benchmark which challenges AI agents to explore novel environments, acquire goals on the fly, build adaptable world models, and learn continuously.

This harness does nothing to actually accomplish those goals.

It's a clever trick, sure, but you aren't allowed to use a calculator on your basic algebra tests in school for a reason.

2 comments

I don't think we got continuous learning here, but we very specifically got interim goal setting and custom world models; the thinking traces demonstrate this round trip of building a world model, mental or coded, then stopping when reality doesn't correlate, then hypothesizing and creating a new model.
And this demonstrates this benchmark does not necessitate achieving those goals to achieve a perfect score. You seem to miss the point that almost all of math, physics, computer science is built on constructing an objective then cheating to attain it, and that demonstrates that there is an equivalency. Maybe the benchmark is flawed, or maybe the goals are not strictly necessary to attain.

For instance, is solving a math proof by enumerating all permutations exhaustively on a computer cheating? Does it matter that it is not a proof by construction? That its not descriptive? Of course not. The proof of the four color theorem is all that’s necessary and sufficient to prove it. Calculator at an algebra exam? Who cares. This isn’t an exam, this is the real world. The fact an AI can use a physics harness to perfectly achieve ARC-AGI-3 without attaining those goals demonstrates the power of the technique and that the goals are not necessary for that class of problems. Then find another benchmark that actually demands the goals be necessary and sufficient to achieve the benchmark goals. But don’t denigrate the fact we have technology today that yesterday was a fantasy.