ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Ataraxos lands in Nature: MIT team's AI beats top humans at Stratego on 1% of the training

Ataraxos lands in Nature: MIT team's AI beats top humans at Stratego on 1% of the training

AI information • Admin • • 5 views

Ataraxos is the AI system research teams from MIT, Carnegie Mellon, NYU, and Stanford published in Nature on September 30, 2026. It defeated the world's top-ranked human player at "Stratego" by a wide margin — the first time an AI has pulled decisively ahead of elite humans in a game of imperfect information. The name comes from Greek, meaning "unperturbed," which also describes its playing style.

First move: a blueprint strategy forged by self-play

The team trained Ataraxos with self-play reinforcement learning: the model plays vast numbers of games against itself, distilling a blueprint strategy. The algorithms are more efficient: they learn fast and never enumerate every possible move. Playing strength strictly exceeds DeepMind's DeepNash, using less than one hundredth of its training examples and less than one thirtieth of its self-play games — while DeepMind's old approach, despite millions of dollars in training costs, never beat top human players.

Second move: imagining the opponent's hidden pieces before acting

The blueprint strategy is only the starting point. Before each move, Ataraxos refines its choice on the fly with decision-time planning: it calls a generative model to probabilistically infer the identities of the opponent's hidden pieces, reconstructs the most plausible board state, then evaluates future lines and picks the best move. The team considers this innovative use of a generative model the missing piece on the road to superhuman performance.

The record — and it generalizes

Ataraxos beat the world's strongest Stratego player 15-1-4, a record margin, and went 39-2 against top humans at the world championship. It was also ported to Barrage Stratego, the cooperative card game Hanabi, and "two versus one" Dou dizhu — superhuman across the board.

It calculates risk very differently from humans: an exposed trump card makes humans panic and overcorrect, leaking information; Ataraxos stays eerily composed and gives nothing away.

From the board to the negotiating table

Stratego is routinely used to model real-world imperfect-information problems: trading, command, negotiation, and cybersecurity share the same structure. The team's next step is interpretability — teaching Ataraxos to explain its decisions: "Humans must have the final say in whether a recommendation is followed; before deployment, we need to be able to audit the model's decisions."

Reinforcement-learning-driven decision-making has already proven itself in reasoning models — Alibaba's QwQ-32B used reinforcement learning to push 32B reasoning to the level of larger models. Ataraxos proves the approach works just as well when the cards stay hidden.

Recommended Tools

More