Hacker News new | ask | show | jobs
by charcircuit 2 days ago
Is no benchmark properly sandboxed? It feels like every single the logs are provided for a benchmark run the LLM is cheating in a way that should have been clearly blocked by a sandbox or the harness.
1 comments

There's an anecdotal paper 'How We Broke Top AI Agent Benchmarks: And What Comes Next' https://moogician.github.io/blog/2026/trustworthy-benchmarks...