Hacker News new | ask | show | jobs
by panqueca 780 days ago
HumanEval Benchmark: 95.1 @ GPT-3.5

I wonder if it can be combined with projects like SWE-Agent to build powerful yet opensource coding agents.

- https://paperswithcode.com/sota/code-generation-on-humaneval

- https://github.com/princeton-nlp/SWE-agent