Hacker News new | ask | show | jobs
by tossandthrow 16 days ago
These types of tests are kind of moot as agentic harnesses are taking over.

IMHO an Ai is the llm plus it's harness.

A good harness would allow the llm to investigate on a map.

Just like the llm can use a python script to figure out how many r's there are in strawberry.

These tests are simply not that predictable of performance of the llm.

1 comments

The test here is not how close the state is to Africa, the test is coming up with a question that is hard for other AIs to answer.
I don't see how that invalidates their point.

Besides, n=1 benchmark seems like more of a coin toss.