Releasing a new "Agentic Reviewer" for research papers. I started coding this as a weekend project, and @jyx_su made it much better. I was inspired by a student who had a paper rejected 6 times over 3 years. Their feedback loop -- waiting ~6 months for feedback each time -- was painfully slow. We wanted to see if an agentic workflow can help researchers iterate faster. When we trained the system on ICLR 2025 reviews and measured Spearman correlation (higher is better) on the test set: - Correlation between two human reviewers: 0.41 - Correlation between AI and a human reviewer: 0.42 This suggests agentic reviewing is approaching human-level performance. The agent grounds its feedback by searching arXiv, so it works best in fields like AI where research is freely published there. It’s an experimental tool, but I hope it helps you with your research. Check it out here:

NeurIPS received 21,575 paper submissions this year. Our Agentic Reviewer, released last week, just surpassed this in number of papers submitted and reviewed. It's clear agentic paper reviewing is here to stay and will be impactful! https://t.co/v4ZjnCgskN
An agentic reviewer scoring 0.42 Spearman against human reviewers — matching the 0.41 correlation between two humans — turns a six-month paper feedback loop into a same-day one, and shows agentic critique reaching inter-rater parity.
articleLitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review PlatformRuotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
videoBuilding uReview, Uber’s Multi-Agent Code Review Engine — Will Bond & Ameya Ketkar, UberAI EngineerChecking sign-in…
Loading comments…