DeepResearchBench
futuresearch.ai- Category
- Developer Tools
- Type
- TOOL
- Date
About
Tests whether agents can research questions on the live web and synthesize correct answers. Results depend on the browsing harness and the date of the run, so compare both alongside scores.
Why it made the leaderboard
Compare the task, benchmark version, harness, and grading method before using model scores to choose a model.
Tags
benchmarkevaluation
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.