Skip to main content
Back to timeline
arXivSource publication:

XiangqiBench uses 119 endgames and 8,568 trajectories to show that finding the move is not winning the game

Related research and updates

Synopsis

The authors introduce XiangqiBench, an executable Chinese chess benchmark in which, starting from 119 tactical endgames with forced mates, an LLM agent must deliver checkmate against an engine defender, and they record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols, finding that three signals that look like competence each overstate closed-loop success: models play the stored reference first move in 26.1% of Sighted trials but only 13.9% of these trials end in a win, the leading model reaches 38.7% pass@3 but only 5.9% pass^3, and 32.3% of accepted simulation calls stop on an illegal move while in 49.3% of comparable cases the real defender replies differently from the simulated line.

Source-provided article image: Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents
Figure 1 ·

Figure 1 : XiangqiBench at a glance. Left: a task instance with the real trajectory overlaid on the initial board (numbered plies; deep red for the agent, charcoal for the defender; curves separate annotations and are not piece paths), tool availability by setting, and the scale of the evaluation. Right: the second Red decision of Gemini 3.1 Pro in Sighted trial 3 on case SQYQ-251, after real plies 1–2 (Appendix G ). The complete pre-simulation think response is shown; simulate then checks three model-specified plies, including an assumed Black reply, on a copy of the board. The defender’s actual reply informs the next decision, and the terminal Red win determines task success.

arXiv

Interpretation

The paper introduces XiangqiBench, an executable closed-loop Chinese chess benchmark: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. Unlike static evaluations that credit a model for naming the right move, this benchmark requires the agent to carry a plan through to a verified outcome while the opponent responds, and an interactive REPL interface separates real moves, state queries, and forward simulation. Built on 119 endgames, 12 frontier LLMs, and 8,568 multi-turn trajectories recorded under two observation protocols.

The paper reports a Conversion Gap: models play the stored reference first move in 26.1% of Sighted trials, yet only 13.9% of these trials end in a win. This shows that recognizing the correct move and completing checkmate against a responding opponent are different capabilities, and that static evaluation overstates closed-loop success. Based on trial statistics under the Sighted protocol, comparing the rate of playing the reference first move with the win rate of those trials.

The paper reports a Consistency Gap: the leading model reaches 38.7% pass@3 but only 5.9% pass^3, winning all three trials on 7 of the 46 positions it ever wins. This separates coverage from reliability, showing that a single success does not imply stable, repeatable success. Based on the comparison of pass@3 and pass^3 for the leading model across repeated trials, and the count of 7 of 46 ever-won positions won in all three trials.

The paper reports a Simulation Gap: 32.3% of accepted simulation calls stop on an illegal move, and in 49.3% of comparable cases the real defender replies differently from the line the agent simulated. This indicates that self-authored rollouts, even when they check legality, cannot anticipate the opponent's actual replies, so simulation ability alone does not support closed-loop success. Based on the proportion of accepted simulation calls stopping on an illegal move and the proportion of comparable cases where the real reply diverges from the simulated line.

Perspective

The work targets forced-mate scenarios in Chinese chess tactical endgames and applies to agent evaluation where a plan must be carried through to a verified outcome while an opponent responds; its REPL interface separates real moves, state queries, and forward simulation, enabling controlled comparison across observation protocols. For researchers and engineers designing closed-loop evaluations and distinguishing coverage from reliability, this benchmark offers directly reusable task forms and statistical measures.

At the abstract level the presentation is aggregate statistics; readers may still want detail on differences across models, endgame difficulty, and the two observation protocols, as well as the precise relationship between accepted simulation calls and illegal-move determinations, which the figures and trajectory analyses in the body can clarify. In addition, how far the conclusion that closed-loop success and reliability should be scored together extends to other board-game or non-board interactive tasks remains an open question worth watching.

Sources