Clueing Up LLMs with Tool-Augmented Deductive Reasoning
Synopsis
This work adapts the board game Clue into a text-based multi-agent environment with six LLM players (three each from GPT-4o-mini and Gemini-2.5-Flash) across 18 games and introduces an external possibility-matrix tool (YES/NO/MAYBE cells plus an accusation_ready signal) that externalizes belief-state tracking, yielding near-perfect accusation accuracy for tool-augmented agents (Gemini-2.5-Flash 1.00, GPT-4o-mini 0.96), raising non-tool agents in mixed games from about 0.81/3 to about 2.71/3, while leaving agents' autonomy over accusation timing unchanged.
Figure 1: Example turn in Clue game environment with the possibility matrix tool highlighted; the matrix records Yes, No, and Maybe values for possible cards in other players’ hands.
arXivInterpretation
It designs and implements a structured possibility-matrix tool for interactive game environments, with three components (envelope_candidates, known_cards_by_player, game_matrix) encoding each card's status per holder (six players plus the envelope) as YES/NO/MAYBE, injected each turn with the prefix "Use this matrix state as the authoritative belief state for this turn". Unlike prior practice of maintaining belief state implicitly through natural-language reasoning logs, the tool offloads extended-memory and deductive constraints from the agent into an explicit external representation, and does not require native tool-calling support in the model. The method is applied consistently across 18 games (six each for Baseline, Matrix, and Mixed) and is illustrated with concrete tool outputs, such as GPT4o_MINI_3's known hand and the game_matrix entry for the Rope card.
Tool augmentation sharply improves accusation accuracy: in Matrix games 94.5% of players identified the correct solution, with normalized accuracy of 1.00 for Gemini-2.5-Flash and 0.96 for GPT-4o-mini, whereas in Baseline games no player reached the final solution, with 47% identifying only one card and 36% identifying none. Prior work by Ansell and Toney (2026) found models still struggled in the text-based Clue setting even with text-based fine-tuning on related logic puzzles; this work shows externalized belief state can yield near-perfect performance. Based on six games per condition and 36 player-game observations per condition, with per-player accuracy shown in Figure 3; the authors note the gain cannot be attributed to increased information exposure, since baseline agents actually approached the knowledge ceiling.
Tool augmentation reduces incorrect deductions: Gemini-2.5-Flash dropped from 2.4 to 0.1 incorrect deductions per game and GPT-4o-mini from 2.39 to 1.67, though GPT-4o-mini retained a similar incorrect-deduction rate across conditions, indicating structured output does not remove unsupported reasoning for every model. It links the reduction of inconsistent reasoning to model-specific behavior, noting that Gemini appears to treat the matrix as a hard constraint. Based on programmatic labeling of player logs, where each inferred card assignment is compared against the ground-truth game state and labeled correct or incorrect, giving log-level, checkable evidence.
Benefits of structured belief state spill over: baseline-prompted players in Mixed games reached a mean accuracy of 2.71/3 versus 0.81/3 in Baseline games, suggesting more informative suggestions raise the quality of shared evidence for all players; yet all Mixed-game wins went to GPT-4o-mini, with Gemini-2.5-Flash winning none in either role. It extends the effect of tool augmentation from the tool user to other agents in the same game, while suggesting win outcomes may depend more on model behavior (such as willingness to accuse) than on matrix use. From a stratified comparison by prompt type across six Mixed games; the sample is small and the authors frame this as a possible game-environment effect.
Perspective
The results apply to deductive reasoning tasks that require maintaining a consistent belief state across many steps, especially interactive settings whose state can be explicitly encoded as a constraint matrix; they are most directly useful to developers using lightweight models who want to improve agent performance without fine-tuning. The authors note the possibility matrix is designed around Clue's card-ownership structure, so the intended setting is deductive environments with clear structure and enumerable constraints rather than open-ended reasoning in general.
The authors list three open questions: the matrix's generality is untested, and whether similar representations extend to inductive or abductive reasoning and to deductive environments with different state structures remains to be examined; the evaluation is limited in scale, with only two model families and 18 games, so larger evaluations across more models (especially stronger reasoning-oriented ones) are needed to tell whether the benefit compensates for lightweight-model limitations or reflects a broader advantage of structured belief tracking; and the baseline compares only against natural-language history reasoning, so it does not isolate whether reduced context length, explicit constraint representation, improved memory retention, or some combination drives the improvement, which would require intermediate baselines such as summarized histories or model-generated tables. In addition, although the tool provides an accusation_ready signal, players still delayed accusations by a mean of 8.2 to 18.0 turns after knowing the solution, leaving how to adjust decision timing through prompt-based intervention an open question.
