Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

XiangqiBench uses 119 endgames and 8,568 trajectories to show that finding the move is not winning the game

The authors introduce XiangqiBench, an executable Chinese chess benchmark in which, starting from 119 tactical endgames with forced mates, an LLM agent must deliver checkmate against an engine defender, and they record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols, finding that three signals that look like competence each overstate closed-loop success: models play the stored reference first move in 26.1% of Sighted trials but only 13.9% of these trials end in a win, the leading model reaches 38.7% pass@3 but only 5.9% pass^3, and 32.3% of accepted simulation calls stop on an illegal move while in 49.3% of comparable cases the real defender replies differently from the simulated line.