Hacker News new | ask | show | jobs
Wandr Benchmark: Evaluating Research Agents That Must Search Wide and Deep (research.perplexity.ai)
1 points by tagawa 16 days ago
1 comments

Repo with benchmark tasks, evaluation harness, tech report:

https://github.com/perplexityai/wandr