Exactly What a Player Knows


Watching the agent play Jaipur, some of its moves didn’t make sense to me. It would take another card when I thought it should be selling.

Was it making a bad decision, or did it not have enough information to make a good one? Investigating those endgame decisions showed that the relevant information was present. But inspecting the inputs exposed a separate gap.

The agent could count the cards in the opponent’s hand. It couldn’t represent cards it had watched that opponent take from the market.

The target is simple: give the agent everything a player could know, and nothing they couldn’t. That includes what the player remembers, not just what is visible on the table right now.

That view is the player’s observation. The encoder turns it into numbers the network can use. Information can be lost at either step.

Remembering the cards

An earlier encoder change had already replaced bare card counts with counts of each kind. The network could distinguish four diamonds from four leather cards in its own hand. But the opponent’s hand was still reduced to a total, regardless of what had happened in earlier turns.

Suppose your opponent takes a diamond from the market. Their hand is private, but that diamond isn’t a mystery: you saw it go in. The old observation forgot its identity as soon as it entered the hand.

The fix was to keep a record of publicly revealed cards. A card seen in a public container stays known when the player visibly takes it into their hand. Cards dealt unseen stay unknown. When a known card is sold face up, it is no longer reported as being in the hand. A shuffle or round reset clears the relevant record.

The observation can now say “one diamond, four unknown cards” instead of just “five cards.” Seven extra inputs, taking the encoder from 78 numbers to 85, let the network read the remembered card kinds alongside an unknown remainder. Training and play use the same view, so the agent learns with information it will actually have when making a move.

This is a first form of memory across moves, but it lives in the engine. The network in these experiments receives a fresh observation for each decision; it doesn’t carry its own memory from turn to turn. The engine preserves a small part of the history for it. It keeps a record of which cards have been publicly seen, then uses that record to include remembered cards in the agent’s view. It is not a memory of every action, or of an opponent’s habits.

Knowing when to forget

Remembering a card is only useful if the record remains trustworthy. The engine knows where every card really is. The player does not.

Jaipur’s sales and exchanges reveal which goods leave the hand. A face-down discard in another game is harder. If the opponent moves one of five cards into another private container, the diamond might have moved, or it might still be in the hand. The observer cannot tell.

Removing the diamond from the record only when it actually moved would leak the answer. So would continuing to report its exact location through the secret transfer. The honest statement is “the diamond is in one of these two places,” not “the diamond is definitely still in the hand.”

The simple tracker used here does not represent that uncertainty. It follows revealed cards through actual containers, so it is not a general solution for private-to-private moves. Clearing records on a shuffle is conservative, but forgetting only the card that secretly moved is not. The Jaipur result should not be mistaken for proof that every hidden-information game is covered.

Why not give it all the cards?

The simulator has the answers, so it is easy to cross that boundary. An earlier diagnostic did exactly that: it gave an existing searcher the opponent’s true hand instead of having it consider several possible hands.

Against the hand-written heuristic, the searcher scored 52.1% with sampled hidden cards and 46.5% with the true hand, over 600 games per condition. The 95% intervals were wide and overlapping: 48.0 to 56.1% and 42.4 to 50.5%. The extra information did not produce a demonstrated benefit in that test.

That was cheating at decision time, not training a model with privileged information. It doesn’t establish that such training cannot help. It does show why that experiment cannot replace this one: handing an existing searcher secrets is different from training an agent to use knowledge a player could legitimately retain.

What changed in play

Consider a diamond appearing in the market after the opponent has visibly collected another. Taking it could deny them a pair. If they have already collected enough to sell, selling your own diamonds first could secure the better tokens. The moves were always legal. What was missing was the information that could make one more urgent than another.

Those illustrate how the new inputs could inform a decision, not a claim that I isolated either tactic in a replay. The recorded behavioral change is broader: the recall-trained agent took goods on 29% of its moves, versus 25% for the control. That is consistent with racing or denying the opponent’s collection, but the action counts alone don’t tell us why it took those cards.

The win result was clearer. I trained two agents from scratch with the same recipe and training seed, changing only the encoder. Across 1,000 games, split equally between seats, the recall agent won 55.0% of decisive games: 533 wins, 436 losses and 31 draws, with a 95% interval of 51.9 to 58.1%. That is one pair of training runs, not an average across training seeds.

Giving the agent what a human could know is the right target, not a proven performance optimum. Here, preserving one small piece of that knowledge changed its play and improved its results. A diamond could leave the public market without disappearing from the agent’s understanding of the game.