We’re open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer. https://t.co/gp2BWjFK4d
Wide-and-deep research requires two capabilities. 1) Agents must search broadly enough to find all qualifying entities 2) Agents must investigate deep enough to support every claim with evidence.

WANDR represents these requirements as hierarchical, independently verifiable records. It consists of 500 research tasks that require 170,495 source-backed records across three tiers of difficulty.

WANDR tests how well agents discover large sets of entities and verify specific facts about each one. It provides a dense, interpretable eval signal that reveals whether an agent fails, and where. The pipeline also doubles as a semi-automated factory for training data.

An openly available research that grades by re-fetching each cited source instead of comparing to a fixed answer key, so tasks can include facts that change over time and failures localise to a specific claim.
Checking sign-in…
Loading comments…