Hacker News new | ask | show | jobs
by stared 14 days ago
See the "caveats" section.

It was also our initial assumption, but to our surprise there were no signs of models "knowing" the solutions (we investigated all trajectories). Compare and contrast with SWE-Bench Verified, for which models suddenly generate solutions, or at least, have substantial hints, https://openai.com/index/why-we-no-longer-evaluate-swe-bench....