by

Miximax-based betting approach

I’m reminded of a wonderful paper (Davidson A – Opponent Modeling in Poker: Learning and Playing in a Hostile, Uncertain Environment) that I found years ago. At the time, I was involved in coding projects and each time I felt I learned something new, it was like an epiphany. Now I can recall some of the best bits, although the details are somewhat hazy.

Davidson compared the Miximax-based betting strategy with three other programs, FBS-Poki, SBS-Poki, and ArtBot. ArtBot was a rather random player, most effectively passive. It was like playing with someone who is inconsistent almost out of a deliberate range of options. So, winning a hand against ArtBot helped, but the potential pots were small, so there aren’t any hundred thousand dollar pots. ArtBot did outperform FBS-Poki, who was a bit of a pushover, with about +0.35 small bets per hand won.

It is like having a range of different sparring partners to practice your moves with. Davidson was putting the Miximax player into play here, she began from nothing and had the ability to sometimes use a strategy developed during the previous games. This Miximax is not just a min-max player. It was a Miximix – a modified version, which does not always take the highest EV move. You can think of it as these are adding an element of randomness that is necessary so the Miximax player does not become predictable and limited to its choices and ensures it has open options when playing.

So, how did it go? Early on, Miximax struggles a little (as the simplistic new student might struggle with their first few games); but once Miximax understands the opponent-type, then it starts building up chips. Average over the FBS and SBS approach and Miximax is +0.4 small bets to +0.5 small bets per hand. ArtBot is struggling more, getting between +0.1 and +0.2 small bets every hand. The fast changing patterns of ArtBot seems to expose Miximax’s context trees flaws.

Then, Davidson adds a really interesting part which is only scratching the surface. The AI should the learn faster, either by developing further context trees or borrowing ideas from previous opponent models. The most frustrating is developing this into multiplayer games where the game tree is immense. Davidson hints that this may be too challenging to handle with more players unless they make shrink the game tree in a major way. This was difficult when desktop PCs could only run at 1GHz, and 4GB RAM was a lot of processing capacity. But now, not an issue.

I also remember that Davidson has this incredible finish where he quotes Josh Billings. Billings said “Life consists not in holding good cards, but in playing well those you hold”. In poker as in life, making better choices with what you have often overruns pure luck.

Miximax

There’s this figure which illustrates Miximax’s performance versus different opponents. It is really clear that as the AI learns more about its opponents it adjusts and drives up its win rate. Of course it might be erratic at first, but then it settles down once it begins accumulating data, just like the way that we humans all learn by our experiences.

So ultimately Davidson’s work illustrates that building a poker AI is more than just playing the numbers game. Its about operating through a shroud of uncertainty, and generating reasonable estimates based on patterns and probabilities. If poker and AI interest you, or just watching machines operate within games, this is a great read.