Hacker News new | ask | show | jobs
Evaluating performance and efficiency of the GitHub Copilot agentic harness (github.blog)
3 points by mariuz 36 days ago
1 comments

It's an interesting post, but i'm a bit skeptical on their decision to report the best run for each agent, and not just the mean over the 5 runs. We have seen this as well in running benchmarks, that variance within one setup can be pretty big.