Y
Hacker News
new
|
ask
|
show
|
jobs
Wandr Benchmark: Evaluating Research Agents That Must Search Wide and Deep
(
research.perplexity.ai
)
1 points
by
tagawa
16 days ago
1 comments
tagawa
16 days ago
Repo with benchmark tasks, evaluation harness, tech report:
https://github.com/perplexityai/wandr
link
https://github.com/perplexityai/wandr