|
|
|
|
|
by stingraycharles
9 hours ago
|
|
Ehr, the SWE bench examples are particularly horrible as those are just publicly available historical PRs. So if the models are trained on GitHub data, it will be included. So almost by design that particular benchmark is tainted, and benchmarks recall rather than reasoning. |
|