|
|
|
|
|
by dannyw
23 days ago
|
|
Basically like passes@6 or passes@5 if you’re doing a benchmark, except for your real tasks. Pro is quite limited on the web UI I reckon. This approach can be highly effective for reasonably verifiable task, for example, write comprehensive unit tests pointing out a tricky bug, get multiple agents to swarm at it. |
|
https://www.erdosproblems.com/