Vibeleaderboard
← All Intel
Intel / article

Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

Source
arxiv.org
Author
Yekun Chai, Qiwei Peng, Haoyi Xiong
Date
Why it matters

The leading model scores 38.7% pass@3 but only 5.9% pass^3. Single-shot or best-of-k scores overstate agent reliability, so evaluate with pass^k and closed-loop outcomes.

Key takeaways · AI-distilled
  • XiangqiBench gives an 119 tactical Chinese chess endgames with forced mates and requires it to actually deliver checkmate against an engine defender, recording 8,568 multi-turn trajectories from 12 frontier LLMs.
  • Conversion gap: models played the stored correct first move in 26.1% of sighted trials, yet only 13.9% of those trials ended in a win, so naming the right move says little about finishing the job.
  • Consistency gap: the leading model hit 38.7% pass@3 but only 5.9% pass^3, winning all three attempts on just 7 of the 46 positions it ever won, which is why the authors want reliability reported next to coverage.
  • Simulation gap: 32.3% of accepted simulation calls stopped on an illegal move, and in 49.3% of comparable cases the real defender replied differently than the agent's simulated line, so self-written rollouts cannot anticipate the opponent.
Terms in this piece · Glossary
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Recommended reads
Comments

Checking sign-in…

Loading comments…