Introducing Web Search Benchmarks 🌐 Rankings of search tools across different models and configurations to help you decide how to ground your agent: https://t.co/GkyEte2YO0

We chose four diverse benchmarks, ran each with four models across four search depths on Exa, Parallel, Perplexity, and the model's native engine. We ranked all the combinations by quality, cost, and speed.
Of all the factors we tested, increasing the search budget had the biggest impact on results. Going from a budget of 1 to 25 turns roughly doubles the scores on the BrowseComp benchmark.

Swapping models had greater effect on score than swapping engines. The average difference in score between frontier and cost-efficient models was ~15 points. Swapping engines while holding the model constant shifted scores by ~10 points on average.
Concrete numbers for tuning search: raising the turn budget helps most, expect failed queries to nearly double search spend, and do not assume a lab's native search beats a third-party engine.
Checking sign-in…
Loading comments…