| Poker-AI.org https://poker-ai.org/phpbb/ |
|
| Strategy Purification and Thresholding https://poker-ai.org/phpbb/viewtopic.php?f=25&t=2421 |
Page 1 of 1 |
| Author: | proud2bBot [ Thu Mar 21, 2013 1:12 am ] |
| Post subject: | Strategy Purification and Thresholding |
Strategy Purification and Thresholding: Effective Non-Equilibrium Approaches for Playing Large Games by: Sam Ganzfried, Tuomas Sandholm, and Kevin Waugh Abstract There has been signicant recent interest in computing effective strategies for playing large imperfect-information games. Much prior work involves computing an approximate equilibrium strategy in a smaller abstract game, then playing this strategy in the full game (with the hope that it also well approximates an equilibrium in the full game). In this paper, we present a family of modications to this approach that work by constructing non-equilibrium strategies in the abstract game, which are then played in the full game. Our new procedures, called purication and thresholding, modify the action probabilities of an abstract equilibrium by preferring the higher-probability actions. Using a variety of domains, we show that these approaches lead to significantly stronger play than the standard equilibrium approach. As one example, our program that uses purication came in first place in the two-player no-limit Texas Hold'em total bankroll division of the 2010 Annual Computer Poker Competition. Surprisingly, we also show that purication significantly improves performance (against the full equilibrium strategy) in random 4x4 matrix games using random 3x3 abstractions. We present several additional results (both theoretical and empirical). Overall, one can view these approaches as ways of achieving robustness against overfitting one's strategy to one's lossy abstraction. Perhaps surprisingly, the performance gains do not necessarily come at the expense of worst-case exploitability http://www.cs.cmu.edu/~sandholm/StrategyPurification_AAMAS2012_camera_ready_2.pdf |
|
| Author: | proud2bBot [ Thu Mar 21, 2013 1:13 am ] |
| Post subject: | Re: Strategy Purification and Thresholding |
I wonder if anyone applied these approaches or did some performance tests, in particular for NLH games. I might try it out later this week and will post my results too. |
|
| Author: | cantina [ Thu Mar 21, 2013 1:20 am ] |
| Post subject: | Re: Strategy Purification and Thresholding |
Yep, standard practice. |
|
| Author: | proud2bBot [ Fri Mar 22, 2013 7:05 pm ] |
| Post subject: | Re: Strategy Purification and Thresholding |
I was a bit in doubt of the technique, but my first results are suggesting that it actually works. I tested a NL HU game with 25bb and 169/600/600/300 buckets. I trained the model for 100M iterations (yes, its not yet converged, but already plays kind of solid) and used this as a baseline bot. Then I used the baseline model and applied thresholding with a parameter of 0.05 to create the test bot and let them play for 50M hands against each other. The test bot beats the baseline with 2.1bb/100. I haven't tried different threshold values yet and I'd imagine that the "best" parameter also depends on the level of convergence of the level. However, the results are quite promising (for comparison: using two exact similar abstractions without thresholding, one trained 50M iterations and the other one 100M, the 100M beats the 50M by 3.2bb/100 over 50M test games). |
|
| Author: | cantina [ Wed Mar 27, 2013 12:04 pm ] |
| Post subject: | Re: Strategy Purification and Thresholding |
I'd be interested to see more experimental stuff done with this, like purifying with squares/square roots, or something to that effect. There's a whole host of fun stuff that could be tried. One such idea: Using something like NEAT supplied with a more complex picture of the hand value and board texture to evolve a situationally-aware purification strategy. |
|
| Author: | proud2bBot [ Wed Mar 27, 2013 2:36 pm ] |
| Post subject: | Re: Strategy Purification and Thresholding |
Having now a 25bb strategy that is learned longer (1 billion iterations), i played around with it and comparing threshold-0.05, threshold-0.1 and purify to the baseline. The results are: threshold-0.05: +0.1bb/100 threshold-0.1: +1.29 purify: +4.0bb/100 |
|
| Page 1 of 1 | All times are UTC |
| Powered by phpBB® Forum Software © phpBB Group http://www.phpbb.com/ |
|