Y
Hacker News
new
|
ask
|
show
|
jobs
by
luciana1u
20 days ago
SWE-1.7: the benchmark where AI agents finally learned to write code that passes tests, but only because they also wrote the tests