RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review
- Source
- arxiv.org
- Author
- Shuyu Guo, Wenxiang Hu, Yuyue Zhao, Yougang Lyu, Xiaohui Yan
- Date
It tackles a specific failure mode in peer review — reviewers being gamed by adversarial prompts hidden in papers — with an explicit rubric-then-judgment pipeline rather than a single opaque judge call.
- LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
“Peer review at major venues is under unprecedented submission pressure, motivating the use of large language models (LLMs) as review assistants.”
“the prevailing paradigms each capture only half of a good review: training-free agents gather broad evidence but produce undirected critiques, while training-based reviewers inherit human discriminative judgement together with its noise and uneven coverage.”
“Experiments on real-world submissions show that RubricReviewer produces reviews that are markedly more comprehensive and more discriminative than prior systems, and exhibits the strongest robustness against adversarial prompt-injection attacks.”
articleLitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review PlatformRuotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
articleLlm As A Judge Evaluate Ai AgentsOpenRouter editorial sitemap
articleEvaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge)Eugene Yan
Checking sign-in…
Loading comments…