We built SciArena to test how well AI models handle scientific literature questions, as judged by researchers. It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes. Here's what they told us. 🧵
What did researchers value most in a model’s answer? 📚 Citation quality—references that are real, relevant, & checkable (23%) 🔬 Depth (19%) 🎯 Directly answering the question (16%) Fluent prose alone didn't win.

o3 finished on top of the SciArena leaderboard – ahead of Claude Opus 4.1, Gemini 3 Pro Preview, & open-weights models like DeepSeek-R1 – with answers researchers called more detailed + to the point. Learn more in our updated blog: https://t.co/DT9LtfSGih
SciArena also allowed us to collect high-quality, expert-annotated ground truth on evaluation data. This is a unique resource & especially important as AI agents are increasingly used to judge the quality of other AI agents, and such evaluations need to be grounded + verified.
Researchers ranked checkable citations above depth and directness when judging model answers on scientific literature, and o3 led the final standings. SciArena retires July 15, leaving its expert-annotated preference data as the durable artifact.
articleLitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review PlatformRuotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
postAnnouncing AA-AnalystAgent, our new agentic benchmark for quantitative analysis…Artificial AnalysisChecking sign-in…
Loading comments…