Hacker News new | ask | show | jobs
by shay_ker 21 days ago
Didn't we all know from the start that all of SWE-Bench was flawed? Even the authors concede the limitations and have long since moved on.
1 comments

SWE-Bench Pro was created to replace SWE-Bench and fix these problems.
SWE-bench Verified was created to fix the problems of SWE-bench.

Then SWE-Bench Pro was created because SWE-bench Verified had flaws.

Now SWE-Bench Pro is shown to have flaws.

Is there a way to benchmark the accuracy, validity improvements in these successive benchmarks?
Bench Bench Pro Maxx Series S 360? The original Bench Bench Pro Maxx Series S had some quality issues, so that's the current followup. We've also released a higher order benchmark developed out of Bench Bench Pro Maxx Series S 360 One King Ranch edition, allowing future benchmark towers to be fully self-contained.
Boo, I thought you were going for Street fighter references at first.
> Many benchmarks include the task of analyzing coding agent aptitude tests. This has led to bench-benchmarks comparing LLM test benchmarking methods, e.g. M Sampson (2025), PL Royle (2024). We performed a bench-bench-benchmark analysis using Mythos Ultra Max 9.6 of these bench-benchmarks.

> Methods: We instructed various LLMs to perform the task "Write a prompt instructing a variety of LLMs to write a benchmark for benchmark analysis, including the instructions 'Create a benchmark for...

Wait where is that from? Googling just takes me back here.
I wrote it myself, adopted from xkcd 1447.
Well, we now have DeepSWE