I totally missed that, because in the charts they showcase for coding, the SWEBench score is not present, they only include it at the end of the post in tables. Hmm.
The SWEBench benchmarks are really gamed at this point and should not be trusted period. The solutions are effectively in the training sets and have been for a while.
Then we are left with what? FrontierCode maybe? IIRC that one evaluates not only if tests pass but also code quality - e.g. whether the maintainer would accept the pull request as is.