
SimpleQA is saturated for search-augmented models; this replaces it with 374 realistic multi-hop questions scored against explicit requirement checklists, and current frontier systems only reach 26.7–59.1%.
articleIntroducing Research Eval A Benchmark For Search Augmented LlmsReka AI editorial sitemap
articleDeepsecBenchMalte Ubl
articleSearchAuditor: Auditing and Attributing Failures in Long-Horizon Search AgentsZhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong CaoChecking sign-in…
Loading comments…