In 2014, researchers at the University of Alberta used their CFR+ algorithm to successfully solve this poker variant, which is quite impressive. They analyzed 3.19×10^14 decision nodes! This would be like trying to count all the stars in a galaxy except instead of stars you are counting more complicated poker hands and betting strategies.
The best way to think of CFR+ (Counterfactual Regret Minimization Plus) is that you are playing a game of poker against yourself, but not just for fun – you are trying to be the best poker player possible. That is exactly what the CFR+ algorithm does. It has no idea how to play, and simply makes random moves. At the end of every game, it looks back at its decisions and thinks, “if I had done this instead of that, would I have done better?” and it will change its strategy by favouring better moves more often.
You could think about this, like learning how to ride a bike, but falling the first time and every time until you don’t fall anymore. Every time you fall off, you learn what not to do next time. CFR+ does this over billions of poker games, learning a strategy that minimizes its ” regret” which is the gap between a the move it made and what the best move would have been. The result is a strategy that nobody can beat, just as you eventually learn to stop falling off your bike and can start riding smoothly. By averaging over all of these games, CFR+ discovers a so-called Nash equilibrium in the sense that it constructs a smart enough strategy to never lose over the long run with any adversaries, including the best of the best.
I am especially fascinated by the fact that this strategy supports some long-accepted truisms about poker while rejecting others. For example, it can almost never “limp” (call the first bet) and any respectable player would acknowledge this with a simple nod of their head. It shows that sometimes you can have a better chance at being capped with a pair of twos (by raising the final bet) instead of a pair of aces. I can picture an experienced poker player reading this from home and saying “No Way!” while right at the same time they are reading it on the screen wondering if they are reading this correctly.
This study has provided humans with a great deal more than just a solution; it facilitates decision making and strategies of human players. The study implications go well beyond poker. These types of algorithms have immense applicability in situations to inform uncertain decisions and strategies in many other large decision invoking fields including security and medical. I wonder who would have even guessed that tools developed late at night playing poker could eventually figure into airport security protocols.
Turing once justified his gaming algorithm work because, quite frankly, it was hard to see as anything other than a jolly good time. This study attains that spirit. Solving for poker was more than just a scientific possible; it has been an interesting intellectual project. If only we could all have jobs that comprised solely of playing games all day for the purposes of science.
Lastly, it is clear that humans have been playing poker for thousands of years, and we have barely scratched the surface of the interplay of decision strategies. And who knows. The next big break through will be a study on another board game, like monopoly, where we will end up figuring out the exact strategy to deploy, so that family game nights end without everyone flipping the board up at the same time in frustrated despair…..


