Hacker News new | ask | show | jobs
by brammertottens 36 days ago
It's an interesting post, but i'm a bit skeptical on their decision to report the best run for each agent, and not just the mean over the 5 runs. We have seen this as well in running benchmarks, that variance within one setup can be pretty big.