With most information hidden, the game Stratego had stumped AI until now
https://arstechnica.com/science/2026/10/ai-finally-beat-the-best-stratego-player-in-history-and-did-it-on-a-budget/Imo, this is the critical piece and what makes the AI work at all.
With hidden information games, the best move depends on information you don’t have. So a move could be good or bad, it just depends on something that’s impossible to know.
You’d like to search ahead, meaning “if I do this they will do that” but that’s impossible since you don’t even know what the opponent can do because you don’t know their hidden state.
If the possible hidden states are randomly distributed, you are screwed. It’s just like rock paper scissors: there’s no best move if your opponent is unpredictable.
However if you can quickly learn to predict their moves, it becomes possible to make informed decisions about what to do.
I thought this was slightly less crank-coded than trying to prove the Riemann Hypothesis, but maybe these days you just ask Claude to do that and it tells you there's a counterexample at 1 + πi that no one ever noticed before.
Too bad I never played against an AI before they cracked it.
Just 16 GPUs, and a few thousand dollars?
What about “researchers from Carnegie Mellon, MIT, New York University, and Stanford University” this wasn’t just anyone.
Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.
[1] https://www.hasbro.com/common/instruct/Stratego.PDF "When an attack is made, the attacker is the only player who has to declare the number of his or her piece. The defender does not reveal the number of his or her piece, but resolves the attack by removing whatever piece has a lower number from the gameboard. Players keep their own captured pieces. Exception: when a Scout attacks, the defender must reveal the number of his or her piece.
I am particularly disappointed that it has influenced how people play the game.
The joy comes from the journey and the experience.
Look at competitive chess and Go and how they have fundamentally been transformed. It's not better and now the box is opened, it can't be closed.
Bridge is played as a pair vs pair game, with North/South and East/West being the two pairs and seated around the table in these compass directions. A bridge hand consists of two phases: there is first an auction phase, where players go around the table bidding on contracts (agreeing to take a certain number of tricks with a certain trump suit) until a final contract is decided. Then there is the cardplay phase, where the player who won the auction is the declarer, their partner is the dummy, and the other pair are defenders. The dummy's hand is placed face up on the the table and the declarer controls which cards are played from dummy, so the cardplay phase is effectively played by only three players now, with each of the three knowing one common hand (dummy) and one private hand (their own) and not knowing the other two hands.
In both the auction and (for the defense) the cardplay phases, it is important for players to exchange some information about their hand to their partner. However, any information you exchange about your own hands also helps your opponents. You might naturally conclude that you want to come up with some secret scheme to exchange information which your opponents don't know (and it is even possible to exchange encrypted information which your opponents can't know--if the defense is known to hold a certain card, but declarer doesn't know in which hand it is, the defense could say that a signal means one thing if the card is in one defender's hand, but means a different thing if it's in the other defender's hand).
But it turns out that this ends up being very uninteresting to play, so instead, when playing bridge, there is an important rule: all of your partnership agreements must be public. If a certain bid that I make promises that I have at least 5 spades in my hand, it is the opponents' right to know that this is our agreement. You must be able to explain the information which your action provides, and you must be able to use the information that the opponents give you themselves.
This poses several problems for self-play reinforcement learning. First, a naive self-play approach will produce agreements that cannot be explained to a human. What really needs to happen is that your partner, when determining what hands you might have as part of search, must not do so simply by sampling its own system (ie by asking what it itself would have done with hand X or hand Y). The information and possibilities really need to be mediated by some kind of intermediate, rules-based description, which can be provided to the opponents as well.
You also need to be able to encode and ingest the opponents' agreements, and to use this information to inform your own decisions. And you need, in particular, to be able to handle a wide variety of agreements from your opponents; it's not enough to force them to play the same system as you.
You must also account for deceit. If, for example, I have a bid which promises that I have at least 2 cards in every suit, it's perfectly legal for me to lie and make this bid when I only have 1 card in some suit--as long as my partner is in the dark about this just as much as the opponents. So if you make this bid, and your machine opponents assume there is a 0% probability of you having lied about your hand, it is possible that they will make gross errors by not accounting for this possibility (for example, they may be in a position where all of their actions are equivalent if you told the truth, but where one action is clearly better if you didn't--a human player will naturally take this action, but a robot may just select an action randomly).
It's an interesting game and a very interesting AI challenge.