Public articles linked to the same research event.
arXiv The authors introduce XiangqiBench, an executable Chinese chess benchmark in which, starting from 119 tactical endgames with forced mates, an LLM agent must deliver checkmate against an engine defender, and they record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols, finding that three signals that look like competence each overstate closed-loop success: models play the stored reference first move in 26.1% of Sighted trials but only 13.9% of these trials end in a win, the leading model reaches 38.7% pass@3 but only 5.9% pass^3, and 32.3% of accepted simulation calls stop on an illegal move while in 49.3% of comparable cases the real defender replies differently from the simulated line.
The authors introduce XiangqiBench, an executable Chinese chess benchmark in which, starting from 119 tactical endgames with forced mates, an LLM agent must deliver checkmate against an engine defender, and they record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols, finding that three signals that look like competence each overstate closed-loop success: models play the stored reference first move in 26.1% of Sighted trials but only 13.9% of these trials end in a win, the leading model reaches 38.7% pass@3 but only 5.9% pass^3, and 32.3% of accepted simulation calls stop on an illegal move while in 49.3% of comparable cases the real defender replies differently from the simulated line.
The authors introduce XiangqiBench, an executable Chinese chess benchmark in which, starting from 119 tactical endgames with forced mates, an LLM agent must deliver checkmate against an engine defender, and they record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols, finding that three signals that look like competence each overstate closed-loop success: models play the stored reference first move in 26.1% of Sighted trials but only 13.9% of these trials end in a win, the leading model reaches 38.7% pass@3 but only 5.9% pass^3, and 32.3% of accepted simulation calls stop on an illegal move while in 49.3% of comparable cases the real defender replies differently from the simulated line.
The authors introduce XiangqiBench, an executable Chinese chess benchmark in which, starting from 119 tactical endgames with forced mates, an LLM agent must deliver checkmate against an engine defender, and they record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols, finding that three signals that look like competence each overstate closed-loop success: models play the stored reference first move in 26.1% of Sighted trials but only 13.9% of these trials end in a win, the leading model reaches 38.7% pass@3 but only 5.9% pass^3, and 32.3% of accepted simulation calls stop on an illegal move while in 49.3% of comparable cases the real defender replies differently from the simulated line.