{"id":9390,"date":"2026-09-30T15:00:00","date_gmt":"2026-09-30T15:00:00","guid":{"rendered":"https:\/\/www.aiproblog.com\/index.php\/2026\/09\/30\/this-game-playing-ai-is-the-new-champ-at-stratego\/"},"modified":"2026-09-30T15:00:00","modified_gmt":"2026-09-30T15:00:00","slug":"this-game-playing-ai-is-the-new-champ-at-stratego","status":"publish","type":"post","link":"https:\/\/www.aiproblog.com\/index.php\/2026\/09\/30\/this-game-playing-ai-is-the-new-champ-at-stratego\/","title":{"rendered":"This game-playing AI is the new champ at Stratego"},"content":{"rendered":"<p>Author: Adam Zewe | MIT News<\/p>\n<div>\n<p>A new AI system that excels at challenging games with hidden information could someday help human decision-makers select ideal strategies to outfox opponents in complicated situations like military maneuvers.<\/p>\n<p>Using advances in machine-learning, researchers from MIT, Carnegie Mellon University, New York University, and Stanford University developed an AI that defeated top-ranked human players of the board wargame Stratego by a large margin \u2014 something no AI system had been able to achieve.\u00a0<\/p>\n<p>Stratego, a two-player game of imperfect information, in which the opponent\u2019s piece identities remain hidden, is often used as a benchmark to test the strategic thinking abilities of powerful AI models.<\/p>\n<p>To build their model, the researchers combined efficient training algorithms with new techniques tailored for calculated decision-making in hidden information settings.\u00a0<\/p>\n<p>The AI system achieved greater performance at Stratego than the next best models, while being far cheaper and less computationally demanding to train. The system also outperformed top human players in other strategic games with different rules and designs, demonstrating how it can be generalized for a variety of use-cases.<\/p>\n<p>The AI system could be adapted to help humans tackle many real-world problems with hidden information, such as business negotiations or cybersecurity.\u00a0<\/p>\n<p>\u201cIn the kind of imperfect information tasks you would face in reality, you often don\u2019t have the luxury of enumerating through all the possibilities. There are just too many. Having AI algorithms that are general purpose and can provably perform this challenging task so well is a big step forward,\u201d says Gabriele Farina, an assistant professor in the Department of Electrical Engineering and Computer Science (EECS), principal investigator at the Laboratory for Information and Decision Systems (LIDS), and senior author of a paper on this AI system.<\/p>\n<p>He is joined on the paper by lead author Samuel Sokota, a graduate student at Carnegie Mellon;\u00a0Eugene Vinitsky, an assistant professor at NYU; Zico Kolter, a professor at Carnegie Mellon; Hengyuan Hu, a graduate student at Stanford; and Zhiyuan Fan, an EECS graduate student at MIT. The research <a href=\"https:\/\/www.nature.com\/articles\/s41586-026-11036-y\" target=\"_blank\" rel=\"noopener\">appears today in <em>Nature<\/em><\/a>.<\/p>\n<p><strong>Hidden information<\/strong><\/p>\n<p>The world is full of imperfect information problems.\u00a0<\/p>\n<p>In these interactions, some parties possess information others do not. For instance, traders in financial markets may not know the rationale behind the trades of others, while military forces likely don\u2019t have full knowledge of enemy positions.\u00a0<\/p>\n<p>With hidden information, the decisions parties make, as well as the decisions they choose not to make, are intertwined in such a way that it is extremely difficult to determine the best steps to take next.<\/p>\n<p>\u201cThe more you bluff, the more your opponent expects it, and the less each bluff is worth. It\u2019s not obvious how to reason about that,\u201d Sokota explains. \u201cIt\u2019s very different from a setting like chess, where the best move is still the best move no matter how often you\u2019ve played it.\u201d<\/p>\n<p>Stratego is often used to model imperfect information situations. In this board wargame, which resembles military chess, players arrange 40 pieces on their side of a board and then move pieces across the board to capture their opponent\u2019s flag.\u00a0<\/p>\n<p>But the identity of all pieces remains secret until they collide, and then the lower-ranking piece is eliminated.<\/p>\n<p>The possible piece configurations number more than 10 to the 66th power \u2014 an exponentially greater number than in chess \u2014 making Stratego extremely difficult for an AI system to play well.\u00a0<\/p>\n<p>Past efforts, such as Google\u2019s DeepMind, relied on sophisticated operations that were computationally demanding and costly. But even with millions of dollars in training costs, these models were still not strong enough to beat top human Stratego players.<\/p>\n<p>\u201cWith Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting,\u201d Farina says.<\/p>\n<p>The MIT researchers set out to develop a full AI system that could achieve superhuman performance for less cost, which they called Ataraxos (a Greek word used to describe one who is unbothered or free from anxiety).<\/p>\n<p><strong>A two-pronged approach<\/strong><\/p>\n<p>To build Ataraxos, the researchers trained the model using a technique called self-play reinforcement learning. The model plays against itself many times to learn a strong \u201cblueprint strategy\u201d of how to excel at Stratego.\u00a0<\/p>\n<p>They designed especially efficient algorithms, which enabled Ataraxos to learn much faster than prior methods while ensuring it didn\u2019t get stuck trying to predict every possible move. This reduces training costs and boosts performance.\u00a0<\/p>\n<p>\u201cOur system reaches strictly higher playing strength than DeepNash (DeepMind\u2019s system) while using less than one hundredth of the training examples and less than one thirtieth of the self-play games, indicating a massive improvement in efficiency,\u201d says Farina.<\/p>\n<p>During a game, Ataraxos uses the blueprint strategy as a starting point to set up the board and begin thinking about its next moves at each round of play.\u00a0<\/p>\n<p>But before acting, it refines its choices on the fly using a technique called decision-time planning. The system employs a generative model that uses probabilities to estimate the likely identities of the opponent\u2019s hidden pieces, then evaluates future choices before selecting the next move.\u00a0<\/p>\n<p>\u201cRather than just guessing blindly, we use decision-time planning to find the most plausible state of the board. Using this generative model allows us to really zoom in on the specific board and opponent we are facing,\u201d Farina says.<\/p>\n<p>The innovative use of this generative model for decision-time planning was the missing piece that enabled Ataraxos to achieve superhuman performance.<\/p>\n<p>Ataraxos beat the strongest Stratego player in the world by a record margin of 15-1-4 and achieved a 39-2 record against top human players at the Stratego world championship. \u201cAtaraxos is good at calculating risk in a way that humans are not. A human might start freaking out if their most valuable piece is exposed, but the bot can be surprisingly composed. It doesn\u2019t overcorrect and give away its secrets,\u201d Farina says.<\/p>\n<p>The researchers also adapted Ataraxos for other imperfect information games, including Barrage Stratego (a faster-paced variant with fewer pieces), Hanabi (a cooperative card game with many players), and\u00a0Dou dizhu (a game in which two players cooperate against a third).<\/p>\n<p>The system achieved superhuman performance in each instance, demonstrating the generality of this method.<\/p>\n<p>In the future, the researchers want to build interpretability measures into Ataraxos so the system can explain its decision-making in a way that a human could understand.\u00a0<\/p>\n<p>\u201cHumans must have the final say in whether a recommendation is followed, so before adoption can happen, we need a way to audit the model\u2019s decisions. We still have a long way to go, but I hope these algorithms can be the foundation for a lot more work to come,\u201d Farina says.<\/p>\n<p>This research is funded, in part, by the Office of Naval Research, the New York University Department of Civil and Urban Engineering, the C2SMART Center, the National Science Foundation, and a Schmidt Sciences\u00a0AI2050 Early Career Fellowship.\u00a0 \u00a0<\/p>\n<\/div>\n<p><a href=\"https:\/\/news.mit.edu\/2026\/game-playing-ai-stratego-new-champ-0930\">Go to Source<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Author: Adam Zewe | MIT News A new AI system that excels at challenging games with hidden information could someday help human decision-makers select ideal [&hellip;] <span class=\"read-more-link\"><a class=\"read-more\" href=\"https:\/\/www.aiproblog.com\/index.php\/2026\/09\/30\/this-game-playing-ai-is-the-new-champ-at-stratego\/\">Read More<\/a><\/span><\/p>\n","protected":false},"author":1,"featured_media":464,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_bbp_topic_count":0,"_bbp_reply_count":0,"_bbp_total_topic_count":0,"_bbp_total_reply_count":0,"_bbp_voice_count":0,"_bbp_anonymous_reply_count":0,"_bbp_topic_count_hidden":0,"_bbp_reply_count_hidden":0,"_bbp_forum_subforum_count":0,"footnotes":""},"categories":[24],"tags":[],"_links":{"self":[{"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/posts\/9390"}],"collection":[{"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/comments?post=9390"}],"version-history":[{"count":0,"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/posts\/9390\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/media\/469"}],"wp:attachment":[{"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/media?parent=9390"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/categories?post=9390"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.aiproblog.com\/index.php\/wp-json\/wp\/v2\/tags?post=9390"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}