Y
Hacker News
new
|
ask
|
show
|
jobs
by
charcircuit
2 days ago
Is no benchmark properly sandboxed? It feels like every single the logs are provided for a benchmark run the LLM is cheating in a way that should have been clearly blocked by a sandbox or the harness.
1 comments
mpavlov
2 days ago
There's an anecdotal paper 'How We Broke Top AI Agent Benchmarks: And What Comes Next'
https://moogician.github.io/blog/2026/trustworthy-benchmarks...
link