AI Has Finally Learned To Play Stratego (arstechnica.com) 2
Researchers from Carnegie Mellon, MIT, NYU, and Stanford have built an AI called Ataraxos that finally cracked Stratego, beating four-time world champion Pim Niemeijer 15 games to one with four draws. Its key advantage was a second neural network that estimates the identities of hidden enemy pieces, letting the system search plausible game states rather than brute-force a massive hidden-information space. Ars Technica reports: Just like [DeepMind's DeepNash, introduced in 2022], Ataraxos learned by playing against itself -- 163 million games in total. In these self-play sessions, moves that led to wins were reinforced and played more often in future matches, while moves that led to losses were played less, which was the same simple training idea. The difference was in how much Ataraxos adjusted after each game, because hidden information tends to send self-play learning algorithms around in circles. The team addressed this by making big, bold changes in strategy early in training and small, careful ones later.
The even bigger innovation was something DeepNash never had: thinking ahead before each move. AIs like AlphaGo refine their general strategy with a search just before acting. DeepMind couldn't make that work in Stratego because the search space was too large, leaving it an open question whether it was worth trying. "This is one of the things that we did figure out how to do," Farina said. The solution was a second neural network, a belief model, trained to guess the opponent's hidden pieces based on how they had been moving. This way, instead of iterating through every possible arrangement, Ataraxos samples plausible ones, plays out candidate moves in each, and picks based on how they turned out. And it shows in its playstyle.
[...] DeepNash was trained for two to three months on 1,024 of Google's specialized chips, a run the Ataraxos team estimates would cost $3 million to $4.5 million at 2025 prices. Ataraxos, in contrast, needed 16 GPUs for a week, plus an additional four GPUs for four days to train the belief model. Farina and lead author Samuel Sokota achieved this efficiency by writing a simulator that runs millions of moves per second on graphics cards. "At the scale that we are in academia, we don't really have access to an entire field of GPUs," Farina said. The algorithm also learned far faster -- it played about 34 times fewer games than DeepNash, and still ended up much stronger. The findings have been published in the journal Nature.
The even bigger innovation was something DeepNash never had: thinking ahead before each move. AIs like AlphaGo refine their general strategy with a search just before acting. DeepMind couldn't make that work in Stratego because the search space was too large, leaving it an open question whether it was worth trying. "This is one of the things that we did figure out how to do," Farina said. The solution was a second neural network, a belief model, trained to guess the opponent's hidden pieces based on how they had been moving. This way, instead of iterating through every possible arrangement, Ataraxos samples plausible ones, plays out candidate moves in each, and picks based on how they turned out. And it shows in its playstyle.
[...] DeepNash was trained for two to three months on 1,024 of Google's specialized chips, a run the Ataraxos team estimates would cost $3 million to $4.5 million at 2025 prices. Ataraxos, in contrast, needed 16 GPUs for a week, plus an additional four GPUs for four days to train the belief model. Farina and lead author Samuel Sokota achieved this efficiency by writing a simulator that runs millions of moves per second on graphics cards. "At the scale that we are in academia, we don't really have access to an entire field of GPUs," Farina said. The algorithm also learned far faster -- it played about 34 times fewer games than DeepNash, and still ended up much stronger. The findings have been published in the journal Nature.