Hacker News new | ask | show | jobs
by osti 19 days ago
GPT usually performs better on DeepSWE while Claude does better on FrontierCode. These two coding benchmarks are pretty much the only ones right now that's still worth taking a look at imo.