A Journey Through Regret Minimization and Poker AI

I remember when I found Richard Gibson’s Ph.D. thesis from the University of Alberta. It was a massive document detailing regret minimization and strategy stitching in extensive-form games, in this case the example of three-player limit Texas hold’em. This one stands out to me because it was created at the most fundamental turn of my learning and discovery to become a poker AI developer.

At the time, I was very engrossed with my studies, and therefore had little appreciation of the nuances of game theory, and how it expresses itself in poker. Gibson’s work was very helpful. His study of Counterfactual Regret Minimization and its application to the game of poker provided me the sound theoretical basis I was in need of, at the time. His research had immense practical applications, and I was already thrilled to begin to use some of these concepts in my projects.

One night, at the end of a long day of programming and reading, I came across Gibson’s Chapter 5. He proposed novel algorithms such as Probing, Average Strategy Sampling, and Pure CFR – they sounded more like practical tools to address computation times and memory costs than theoretical novelties. It was a gold mine for someone in the position I was in, facing such limited computational resources.

His work inspired me to attempt to build some of these algorithms into my own poker bot. I can remember those nights debugging – a cup of coffee in one hand and Gibson’s dissertation in the other. Then, there was one particular night where things just clicked. My poker bot, which had been floundering, not even close to making profitable decisions, began to show signs of improvement. It was as if Gibson himself was jointly guiding my hand through all the complex subtleties of CFR and its potential applications.

It was most rewarding to test my poker bot in a small online poker tournament. I enjoyed it sailing through hands with its new found efficiency and reviewing its strategies. After all the study, programming, and sheer will, it came down to that.

It is pretty incredible how far we have come in poker AI from those times. All of the theories and algorithms that had previously only occupied the realm of academic papers have now made their way to advanced poker bots. And it all began with the inspirational works that pushed researchers like Gibson towards contributing.

Anyone interested in how the poker AI worked or how all the strategies came to be, should really check out the resources, and see what’s out there today. And if you ever get lost in the complexity, just remember, every great endeavor in AI emerges from a single line of code, and a whole bunch of curiosity.

Continue to learn, continue to program, and maybe one day your project will be one of the trend-setting projects in the world of poker AI.

Best regards 😉

485 Words

Poker Bot Philosophy

crazy coder of poker AI bot

So, I am an Artificial Intelligence enthusiast. Of course, developing AI that plays poker means—more than once—you are banking into some of the deeper philosophical questions of our time. Not about the meaning of life or if pineapple belongs on pizza—it does, by the way—but actually something a bit more important: the nature of poker and how to model it for a bot.

First off, let’s break down poker. Obviously, it’s a game. To a bot engineer, however, poker is all about systems—systems of rules and interactions. Poker doesn’t represent cards or chips but the players. And it really doesn’t make much sense to have a poker game without a player, much like having a computer and never putting it on the Internet.

It’s a blend of game situation, history by opponents, own mood—and the phase of the moon. Yes, some players are that superstitious. And that is exactly what makes modeling a poker player interesting and challenging—the mix of logical and illogical.

The comforting thought, in all of this, for engineers is that poker really does represent a finite number of states. Every player begins with some chips; there’s only so many cards are in the deck; at any time during a game, the number of possible moves is limited. It’s a godsend when trying to model a game. The idea of trying to model an infinite—it’s like trying to find the end of the internet.

Now, this is where it gets interesting. Players don’t do the same action in the same situation all of the time. They mix up their choice: make decisions like “I’ll fold 45 percent of the time and call 55 percent of the time.” It’s this randomness that your poker bot has to emulate since in poker, when you become predictable, you’re dead.

But that brings up a difficult-sounding concept: If players can make mixed moves, that must mean I have to model an infinity of possibilities. Thankfully no. While it sounds daunting, it’s really an issue of accuracy. As I take small changes in a player’s mixed strategy, I get small changes in expected profit to that player. It’s not about modeling infinity here, but how much reality one can cope with without going nuts.

Now, let us take a look at some strategies that make things simpler. Whether a player’s head is full of brains, sawdust, or algorithms, their strategy can always be reduced to being representable by a Look-Up Table of decisions. That is, the table says what to do given any situation. Now, when several players with their LUTs sit at the table, this goes down into an increasingly exciting strategic interplay with expected profit.

The temptation of infinity had mixed moves. Let us kill this myth at once. Imagine planning to play an infinite number of hands. It is just impractical. We think in terms of finite sessions and look at the expeeCTed profit over these sessions.

Consider, if you call, profit may be $40, and in case you fold, it could be -$10. Mixing your move leads to profit calculation, like: 0.4 * $40 + 0.6 * -$10 = $10. Small changes in these mixed strategies will alter expected profit only slightly, proving it about precision, but definitely not infinity.

To drive the point home, let’s consider some practical examples. Suppose, for example, that you sit down at the table behind a stack of $665. You are against a superstitious opponent who always folds when he sees the number 666. Suddenly, what had, to this point, seemed to you perhaps like a small difference in your stack size may very well make a huge difference in your expected profit. It’s those kinds of quirky little details that make poker so fascinating and maddening all at once.

There you go. Modeling a poker bot includes finiteness of the game, introducing randomness into mixed moves, and simplification of strategies into something our bots can handle. A mix of philosophy, math, and a dash of humor is added.

663 Words

Poker AI for Texas Hold’em

I’ve looked over some of the past work on poker AI, but after a long day of work and a few beers, I might be a little slow this evening. Nevertheless, I have had some time to read this paper on AI for Texas Hold’em, and I started to think about how we could make machines learn to bluff better than the average person.

To start, poker is more than a card game. Poker is an enigmatic fusion of statistics, psychology, and a dash of magic. Have you ever tried to guess if the guy sitting opposite of you is going to raise or fold? That is comparable to guess what my cat wants for dinner: impossible. So after the dust of chaos has settled, you look to artificial intelligence to help manage the storm.

Building a poker AI is like trying to teach chess to an infant, but with fewer snacks and slightly less chaos (most of the time). You will start with the basics: you first want to teach it to identify strong hands. Of course, there is a charm to poker because it has an element of randomness — you do not get to see the hands of your opponents. You want the AI to have to deal with partial information, to learn to make intelligent predictions, and to not think too poorly of itself when someone puts the hammer down.

For example, you are dealt the hand A♣ Q♥. The flop shows 3♦ 4♠ J♥. So, your AI has to determine, “What are the odds that this hand is any good?” It’s a bit like deciding if the leftover pizza you have in the fridge is still safe to eat a week later. Spoiler: it probably isn’t and your AI has a ~58.5% chance of getting it right. However, the more players involved, the odds decrease like your WiFi does during a critical Zoom call. Your shiny A-Q does not win only ~6.9% of the time against 5 opponents. Ouch.

Then we come to “potential” – as in, your hand is not great, but if you get lucky, could potentially give you a royal flush. It’s like gambling on your startup’s future success with 0% chance of prevailing. Take for example holding 6♦ 7♦ with a flop of 5♦ A♠ 8♦. Not looking good? If you hit the perfect turn and river, you may get a straight flush instead. So now your AI must go from “I’m screwed” to “maybe there’s a chance I will win!”

Before we move on, let’s talk about your opponents, because this is not solitaire! Your AI also has to model opponents in order to infer whether they are tight or loose, aggressive or passive. It’s like trying to figure out if your neighbor will give you your lawnmower back on time. Your AI uses neural networks, Bayes, possibly even something called particle filtering (whatever that is). I guess it’s like waving a magic wand over your opponent to try to guess their next move.

To increase the awe factor even more, your AI creates game trees for every possible outcome.It’s sort of like trying to think through, simultaneously, every possible conversation with your employer – it is helpful but tedious. And these trees help your AI reflect on the best course of action based on that value of every action. Raise, fold, call – all plotted out.

And now the goal is for AI to get smarter and better. Over millions of hands, it learns and becomes a poker master. Isn’t it somewhat similar to your child suddenly improving their score at a video game after 6 hours of virtual practice? Or like, your cat finally perfected the skill of knocking things off the table.

Look at this figure. The chart shows VPIP (Voluntarily Put Money In Pot) values for each player cluster. It is a clever way of saying, “Who’s the sucker who always bets?” It turns out – the lower the stakes, the more players want to see a flop – as if every cat has a dream of hitting that miracle river card.

In sum, developing a poker bot is not an easy task. You might think of it more like running a marathon, with some fresh math problem you’re handed at each mile. But that’s just so cool to see it outsmart humans or execute those clever moves? Priceless… Just make sure you provide your AI with decent data (we have billions of these hm2 files), somewhat like you would a cat – don’t give it too much at once, and don’t give it anything dangerous it can lash back at you!

And just a quick reminder, I’m open to collaborations and interesting projects. You can reach out to me on Linkedin, preferably starting right from the point. I have been involved in many projects, judging a poker AI competition, and developing the architecture of one of the best commercial poker AIs (my humble opinion) for my clients and some private projects. Ciao!

826 Words