
Introduces a 380-task requiring agents to combine SQL queries and open web search into one verifiable answer, testing constraint handoff that prior deep-research benchmarks ignored.
“even state-of-the-art models like GLM-5.2, Claude-Sonnet-4.6 and GPT-5 achieve only about 50-54% Pass@8 on the hard subset”
“directional reasoning is substantially more difficult than parallel intersection”
articleIris: Climbing to the Search FrontierZiyuan Liu, Hengqi Liu, Zichuan Wang, Yang Qin, Jiachen Liang, Xu Chu, Shaowei Chen, Yuantao Gu, Mu Chuan
articleBacktrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQsRuoxi Zhao, Maziar RaissiChecking sign-in…
Loading comments…