Vibeleaderboard
← All Intel
Intel / post

OpenAI web search lands mid-pack on the Artificial Analysis Search Index

Source
x.com
Date
A
ArtificialAnlys@ArtificialAnlys

OpenAI Web Search debuts on the Artificial Analysis Search Index at 74, the 5th best provider behind Perplexity, Octen, Parallel and Brave @OpenAI Web Search is the first integrated first-party search tool on our board. Instead of an agent calling a Search API in our Stirrup harness, GPT-5.6 Luna (medium reasoning) calls OpenAI's built-in web_search tool, and OpenAI runs the whole search loop inside one Responses API call. All setups use the same underlying model, GPT-5.6 Luna (medium reasoning). At ~$0.05 per task, OpenAI Web Search costs less than Parallel (advanced, $0.06) and Perplexity (medium, $0.07), but about twice as much as Octen ($0.024), which scores 3 points higher. Key benchmarking results for OpenAI Web Search (medium search context): ➤ 5th best provider on the Artificial Analysis Search Index: OpenAI Web Search scores 74, a 41-point lift over the same model with no search (33). It places 7th of 26 products, level with You . com (highlights), Nimble (standard) and Exa (auto) at 74, behind the three Perplexity variants (77 to 80), Octen (77), Parallel (advanced) and Brave (LLM context) at 75 ➤ OpenAI Web Search performs best on AA-Omniscience, our proprietary…

Read the full post on X
Why it matters

Shows built-in OpenAI web search lifts the same model from 33 to 74 and costs about $0.05 per task, but trails Octen on score at twice the price. Useful when choosing a search backend for agents.

Key takeaways · AI-distilled
  • Unlike the 25 other setups, which call a Search API from Artificial Analysis's Stirrup , OpenAI Web Search runs the whole search loop inside a single Responses API call. GPT-5.6 Luna (medium reasoning) is held fixed across every setup.
  • Results depend on the task: OpenAI Web Search scored 72% on the factual AA-Omniscience , 3rd of 26 and 1 point behind Firecrawl, but 73.5% on multi-hop BrowseComp, 13th of 26, against 85% to 87% for the Perplexity variants and Octen.
  • Artificial Analysis reports OpenAI bills about 40k input per task, search results included, versus 125k for the leanest Search API tested (Perplexity low), putting its model cost near $0.009 per task.
  • Pricing is $10 per 1,000 web_search calls plus search content billed as model input tokens. Artificial Analysis says it blocked known contamination sources through the tool's domain filter.
Terms in this piece · Glossary
  • agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…