Hacker News new | ask | show | jobs
by bel8 11 days ago
What harness? These can make or break benchmarks because of tool call failures/limitations.

And is the benchmark open source?

1 comments

Pi. The benchmark is local, mostly stuff from my work, I run it everytime a new model comes up. The top model rn is gpt 5.6 sol, followed by fugu ultra, fable, opus 4.8, gpt 5.5 and glm 5.2 (which is the REAL IMPRESSIVE one still). Kimi-k3 is 14th in the list.
Cool, and how is it ranked?