Breakthrough in Artificial Intelligence: Ataraxos Beats Human Stratego Player
A team of researchers from top universities has created an AI system that can reliably beat a human player at the strategy game Stratego.

A breakthrough in artificial intelligence has been achieved by a team of researchers from several top universities. They have created an AI system called Ataraxos that can reliably beat human players at the classic strategy game Stratego.
The victory came against Pim Niemeijer, widely regarded as the best Stratego player of all time. The match was not even close, with Ataraxos winning 15 games to one, and drawing four times. This achievement required only a relatively modest amount of computing power, just 16 graphics processing units (GPUs) were used in training.
In Stratego , each player has a set of pieces representing military ranks, from a marshal down to a spy, along with bombs and a flag. The objective is to capture the opponent's flag while hiding your own piece identities behind a mask of uncertainty. This makes Stratego an imperfect-information game, where players must make decisions without complete knowledge of their opponent's moves.
Eugene Vinitsky, a researcher at NYU and co-author of the study, notes that Stratego has a unique feature: it involves a massive amount of hidden information that unfolds over time. This characteristic has made it challenging for computers to crack until now.
The game of Stratego presents a unique challenge for artificial intelligence (AI) researchers due to its complex rules and hidden information. With 40 pieces on the board that could be in any order, there are more than a decillion possible setups, making it difficult for computers to crack.
A key factor contributing to this complexity is the game's length, which can easily last up to 2,000 moves. This is significantly longer than most chess games, which typically consist of around 40 moves. The extended duration allows players to employ various tactics and strategies, including bluffing, which adds another layer of intricacy.
Bluffing involves moving a weak piece as if it were a stronger one in order to deceive the opponent. However, when done too frequently, threats become meaningless, while avoiding bluffs altogether makes the player predictable. This delicate balance has stumped earlier AI attempts, including DeepMind's DeepNash, introduced in 2022.
To overcome this challenge, Ataraxos was trained using a self-play method, where it played against itself over 163 million games. During these sessions, moves that led to wins were reinforced and played more often in future matches, while losses resulted in less frequent play.
The training process involved making significant changes in strategy early on and smaller adjustments later on, addressing the common issue of hidden information causing self-play learning algorithms to loop indefinitely.
The development team made significant progress in creating an AI capable of playing Stratego by introducing a new element: thinking ahead before each move. In most games, AIs refine their strategy through a search just before acting, but this approach was not feasible for Stratego due to its vast search space.
To overcome this limitation, the team trained a second neural network, known as a belief model, which predicted the opponent's hidden pieces based on their movements. This allowed Ataraxos to sample plausible scenarios and evaluate candidate moves in each one, rather than exhaustively exploring every possible arrangement.
The name Ataraxos, derived from ancient Greek, conveys a sense of calmness and composure. According to Farina, the AI's structure and lack of human emotions enable it to maintain a level head even in situations where a human player might panic or make impulsive decisions.
Ataraxos' strategy is notable for its ability to conceal its own weaknesses from the opponent. When the AI determines that its adversary has no reason to suspect a vulnerability, it deliberately avoids drawing attention to that area of the board, even if it appears precarious in hindsight.
By adopting this approach, Ataraxos minimizes the risk of exposing itself and instead focuses on slowly and methodically rebuilding its position in the game.
The team behind Ataraxos acknowledges that knowing sensitive information can be challenging for humans to ignore, but not for machines. This disparity in decision-making allows AI opponents like Ataraxos to make bold moves that a human might only consider while bluffing.
Ataraxos's human opponent, Niemeijer, is a highly skilled player with four world championships and over 600 weeks as the top-ranked player under his belt. He was able to play against Ataraxos in an online setting, winning just once out of twenty games played over three weeks, earning $100 for each victory.
Niemeijer's lone win against Ataraxos may not have been a result of a flaw in the AI's strategy, but rather good luck. As researchers point out, playing Stratego well involves randomizing the placement of pieces, making it impossible to account for every possible outcome. Even with a perfect strategy, there will be instances where the human player loses.
The team notes that Niemeijer got lucky at times, but so did they. This back-and-forth dynamic highlights the unpredictability of the game and the challenges faced by both human and AI players.
The surprise strategies employed by Ataraxos in Stratego have left players bewildered, as evidenced by its tendency to hide its flag behind an unusual combination of two bombs.
Ataraxos's cost-effectiveness was a significant factor in its success. The team trained the AI on 16 GPUs for just over a week, compared to DeepNash's lengthy training period on Google's specialized chips.
The Ataraxos team developed a simulator that enabled them to run millions of moves per second on graphics cards, significantly reducing the computational power required. This innovation was crucial in allowing the team to access and utilize a substantial number of GPUs within their budget.
This efficiency allowed Ataraxos to learn much faster than DeepNash, requiring only about 34 times fewer games to achieve comparable strength.
The versatility of the Ataraxos architecture is noteworthy, as it has successfully applied its approach to other games. The team's algorithm outperformed world champions in Barrage Stratego and mastered Hanabi, a cooperative card game.
The Ataraxos team now aims to tackle more complex real-world problems, which typically involve uncertain outcomes and multiple variables. They believe that the underlying principles of their approach can be adapted to address such challenges by first developing simplified models of these issues.
The researchers behind Ataraxos are now exploring ways to adapt their approach to more complex challenges in game theory and strategy.
They believe that understanding how Ataraxos makes its moves could provide insights into developing stronger but also more explainable AI strategies.
Facts based on reporting originally published by Ars Technica.
You may republish this story, in full or in part, if you credit Noti Group and link to it (licence CC BY 4.0). Photos are not included.












