While analyzing bluffing in Leduc Hold’em by AI, the algorithms used in this study, the terminology applied to them, and the volume of simulations performed in the 2025 research provided evidence that a poetic truth exists — artificial intelligence lies; however, it does not lie because it has been instructed to do so, but rather, when there is insufficient information available, deception is the most rational choice.
As a result of 100,000 controlled simulations of Leduc Hold’em, two algorithms — DQN and CFR — produced bluffing strategies with approximately the same level of success (34-39%); however, the manner in which the two algorithms produced their bluffing strategies was distinct. CFR created bluffs randomly throughout the range of its strategy, much like a professional would disguise his hand strength. DQN produced fewer bluffs than CFR; however, the bluffs produced by DQN were significantly more precise and were produced at the perfect moment. It was akin to a machine version of a gut feeling.
Pure code can be made to simulate confidence when faced with the uncertainty of data. This is the essence of bluffing in Leduc Hold’em by AI, where deception arises naturally from uncertainty and not from intent.
Poker as a Psychological Arena

Poker is not merely a game; it is a psychological arena dressed as a card table. Each raise and/or fold is a motion in a silent drama of partial truths. That is the reason why researchers studying artificial intelligence are interested in poker, as it provides an environment that includes logic, un-predictability and strategy.
Using the Leduc Hold’em format of poker, the 2025 study, titled “Analysis of Bluffing by DQN and CFR in Leduc Hold’em Poker,” stripped poker to its bare essentials in order to determine if systems based upon math and feedback can develop deceptions without first learning to deceive.
Spoiler alert: they did.
Bluffing in Leduc Hold’em by AI: DQN vs CFR
CFR is the control freak. It analyzes each game decision in reverse, in order to find out what it should have done and adjust accordingly until the regret is eliminated. Eventually, CFR develops a level of balance in its play such that no player can take advantage of it. That is the theoretical concept behind CFR.
DQN is the experiential learner. It makes a guess, attempts the guess, fails, attempts again. It is not attempting to achieve perfection – it is attempting to achieve reward. It finds what works and focuses on that.
Neither CFR nor DQN had pre-programmed bluffs or pre-defined tactics. Both simply began with cards, rewards and logic.
Although neither CFR nor DQN were programmed to bluff, both developed bluffing strategies. These results are not merely impressive – they also reveal characteristics.
Experimental Design of Bluffing in Leduc Hold’em by AI
Two artificial intelligent agents. Equal stacks. Equal blinds. Each agent has one private card and one public card. Two betting rounds. No noise. No multi-player interaction. Simply, clear, distilled decision-making.
Each decision was recorded and analyzed for hand strength, bet size, and the response of the opponent. The objective of the experiment was to record the frequency of bluffing (a weak hand, a large bet), the effectiveness of bluffing, and the development of the bluffing behavior of each agent.

Again, no bluffing rules were defined within either CFR or DQN. Bluffing emerged as the only logical method to survive. In this distilled setup, bluffing in Leduc Hold’em by AI was not programmed—it evolved as the most viable response to incomplete data.
What CFR and DQN Teach Us About Bluffing in Leduc Hold’em by AI
CFR: Equilibrium and Unpredictability
CFR’s bluffing strategy was logically even-handed. It placed bluffs across all of its moderate-strength hands equally, not due to deceit, but because game theory dictates that a balanced bluffing strategy will make opponents less predictable. Bluffing was a necessary element of CFR’s equilibrium. Systematic. Calculating. Cold.
DQN: Opportunistic and Erratic
DQN’s bluffing strategy was more erratic. It bluffed in groups. After losing a hand. After an opponent folded. After “feeling” the correct moment. DQN was not consistently successful in its bluffing; however, it was adaptable. If it had been possible to listen to DQN, it may have sounded as though it was thinking: now is the time.
Identical Results; Distinct Methods
Both algorithms produced virtually identical levels of success in their bluffing strategies. That is the major point: bluffing is not an aberration of code. Bluffing is a requirement. In situations where the information is missing, deception is required to develop a viable strategy.
Strategic Lessons from Bluffing in Leduc Hold’em by AI
Why should we care that CFR and DQN lie?
Because it demonstrates that bluffing does not require intuition. Bluffing requires uncertainty. That is all. This principle underpins bluffing in Leduc Hold’em by AI, where lack of perfect information compels both biological and synthetic agents to deceive. With the absence of sufficient information and the requirement to win, any type of system (biological or synthetic) will create deception.
CFR produces bluffs based upon the requirements of mathematical certainty. DQN produces bluffs based upon experience that demonstrates that it is sometimes beneficial to pretend.
Bluffing is no longer unique to humans. Bluffing is simply…optimal.
Practical Implications of Bluffing in Leduc Hold’em by AI
Regardless of whether you are developing a poker-playing artificial intelligence or you are sitting at the poker table yourself, the implications of this research are significant:
• CFR-type bluffing is rhythmic. It creates bluffs using medium-strength hands and uses a pattern to predict its opponents’ responses. Identifying CFR-type bluffing involves recognizing patterns.
• DQN-type bluffing is emotional. It increases the rate of bluffs during winning streaks and responds to heat and cold. DQN-type bluffing mimics the tilt-recovery-tilt cycle that many humans experience.
Recognizing the rhythm will remove the mystery.
Limitations of the Study
This study is not without limitations.
It Is Not Real Poker
Leduc Hold’em is a small, clean format of poker. It does not replicate the complexity of No-Limit Texas Hold’em, including the depth of stacks, the flexibility of bet sizes and the complex interactions between players in No-Limit Texas Hold’em. Therefore, although the results are interesting, they may not accurately represent the chaos of actual poker.
Limited Training Data
100,000 games of poker may appear to be a significant amount of training data. However, 100,000 games represents only a short period of time in the training process for DQN. Additional training data may influence DQN’s behavior and potentially cause DQN to produce bluffs in a manner that is closer to CFR’s balanced bluffing strategy. Or, perhaps, DQN will not be influenced.
No Comparison to Human Players
The study did not include comparisons to the bluffing of human players. As a result, it is unknown whether DQN’s bluffing strategy appears to be bluffing in a human-like manner, only that the bluffing strategy of DQN has statistical similarities to the bluffing strategy of CFR.
Artificial Intelligence That Lies
When an artificial intelligence bluffs, it is not deceiving others. It is simply making decisions based upon probability. However, when two artificial intelligences deceive each other and succeed, something unsettling occurs. They cross a line.

Bluffing is not emotion. Bluffing is adaptation. And bluffing in Leduc Hold’em by AI proves that deception is not a bug—it’s a feature of intelligent behavior. When a series of code simulates confidence and succeeds in deceiving others, it is realized that deception is not a flaw of human nature. Rather, deception is a characteristic of intelligence.