
Allocating expensive model reasoning and human review across a candidate pool is the same problem every agentic research pipeline has, framed here as search and recommendation rather than prompting.
articleTen advances in mathematics and theoretical computer scienceopenai.com
articleAutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model ResearchMarjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
articleDiagnostic Foundation for Evaluating LLMs' Research Integrity as Co-ScientistsYash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin LiSign in to comment.
Loading comments…