Skip to main content
Back to timeline
MIT News - Artificial intelligenceSource publication:

MIT-led team's Ataraxos beats the world's strongest Stratego player 15-1-4, training on under one hundredth of DeepNash's examples

Synopsis

Researchers from MIT, Carnegie Mellon University, New York University, and Stanford University developed an AI system called Ataraxos that combines a self-play reinforcement learning "blueprint strategy" with decision-time planning using a generative model, defeating the strongest human Stratego player by a record 15-1-4 margin and achieving a 39-2 record against top players at the Stratego world championship, while using less than one hundredth of DeepMind's DeepNash training examples and less than one thirtieth of its self-play games, and generalizing to other imperfect-information games such as Barrage Stratego, Hanabi, and Dou dizhu.

AI-generated editorial illustration: This game-playing AI is the new champ at Stratego

Interpretation

Ataraxos reaches superhuman Stratego play, beating the strongest player in the world by a record 15-1-4 margin and posting a 39-2 record against top human players at the Stratego world championship. Prior systems, including DeepMind's, were computationally demanding and costly and were still not strong enough to beat top human Stratego players, so this is the first system reported to defeat top-ranked humans on this benchmark. The text gives concrete match scores (15-1-4 and 39-2) against the strongest player and top world-championship players; Stratego's possible piece configurations exceed 10 to the 66th power, far more than chess, indicating the task's difficulty.

The system substantially improves training efficiency over DeepNash: it reaches strictly higher playing strength while using less than one hundredth of the training examples and less than one thirtieth of the self-play games. Earlier approaches relied on sophisticated, computationally demanding operations, whereas Ataraxos uses efficient algorithms that avoid getting stuck trying to predict every possible move, reducing training cost. The efficiency comparison comes from a direct statement by Farina with quantified ratios (under one hundredth, under one thirtieth), though the text does not list absolute training volumes or hardware details.

The method's key ingredient is combining a self-play reinforcement learning blueprint strategy with decision-time planning, in which a generative model estimates the likely identities of the opponent's hidden pieces by probability before evaluating future choices. The authors describe this innovative use of a generative model for decision-time planning as the missing piece that enabled superhuman performance, letting the system zoom in on the specific board and opponent rather than guessing blindly. The text presents this mechanism as the key to the performance breakthrough through author interview, explaining how it works, but provides no ablation study or quantitative decomposition of component contributions.

The approach generalizes: the researchers adapted Ataraxos to Barrage Stratego, Hanabi, and Dou dizhu, games with different rules and designs, and it achieved superhuman performance in each. This indicates the method is not a Stratego-specific solution but can transfer to cooperative, multiplayer, and differently structured imperfect-information games. The text names three specific games and states superhuman performance in each instance, but does not give specific scores or opponent-level details for those games.

Perspective

The result applies to decision settings with hidden information where enumerating all possibilities is not feasible, and the authors explicitly frame it as an algorithmic foundation that could transfer to real-world problems such as business negotiations, cybersecurity, and even military maneuvers. For researchers, it provides a system whose training cost on a benchmark like Stratego is far below prior methods; for practitioners hoping to apply AI to imperfect-information decisions, it shows that general-purpose algorithms can perform well on such tasks. The authors also note that before adoption can happen, humans need a way to audit the model's decisions, so their next step is to build interpretability measures into Ataraxos so it can explain its decision-making in a way a human could understand.

Readers should still watch for several things: the text does not give absolute sample counts, compute, or hardware for training, presenting efficiency only as relative ratios; the key role of the generative model in decision-time planning comes from author statements without quantitative ablation support; superhuman performance in Barrage Stratego, Hanabi, and Dou dizhu is not accompanied by specific scores or opponent levels; and transfer from board games to real settings such as business negotiations or cybersecurity remains the authors' outlook rather than a demonstrated result. In addition, this is a public-facing account that does not include the paper's figures or experimental details, so the description of the method's internals rests mainly on author interviews.

Sources