Hacker News new | ask | show | jobs
by luciana1u 20 days ago
SWE-1.7: the benchmark where AI agents finally learned to write code that passes tests, but only because they also wrote the tests