Vibeleaderboard
← All Intel
Intel / article

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

Source
arxiv.org
Author
Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
Date
Why it matters

Puts a number on how far literature-review agents still sit behind expert drafts, and shows the agentic scaffold — not the base model — carries most of the gain.

Terms in this piece · Glossary
  • LLM-as-judge — Using one model to score another's output against a rubric, so quality can be measured at a scale human grading cannot reach.
Recommended reads
Comments

Checking sign-in…

Loading comments…