No-Regret Learning in Extensive-Form Games with Imperfect Recall

Lanctot, Marc; Gibson, Richard G.; Burch, Neil; Bowling, Michael

No-Regret Learning in Extensive-Form Games with Imperfect Recall

Marc Lanctot, Richard G. Gibson, Neil Burch, Michael Bowling

ICML 2012

/icml/2012/lanctot2012icml-regret/

Abstract

Counterfactual Regret Minimization (CFR) is an efficient no-regret learning algorithm for decision problems modeled as extensive games. CFR's regret bounds depend on the requirement of perfect recall: players always remember information that was revealed to them and the order in which it was revealed. In games without perfect recall, however, CFR's guarantees do not apply. In this paper, we present the first regret bound for CFR when applied to a general class of games with imperfect recall. We also show that CFR applied to any abstraction belonging to our class results in a regret bound not just for the abstract game, but for the full game as well. We verify our theory and show how imperfect recall can be used to trade a small increase in regret for a significant reduction in memory in three domains: die-roll poker, phantom tic-tac-toe, and Bluff.

PDF Semantic Scholar

Cite

Text

Lanctot et al. "No-Regret Learning in Extensive-Form Games with Imperfect Recall." International Conference on Machine Learning, 2012.

Markdown

[Lanctot et al. "No-Regret Learning in Extensive-Form Games with Imperfect Recall." International Conference on Machine Learning, 2012.](https://mlanthology.org/icml/2012/lanctot2012icml-regret/)

BibTeX

@inproceedings{lanctot2012icml-regret,
  title     = {{No-Regret Learning in Extensive-Form Games with Imperfect Recall}},
  author    = {Lanctot, Marc and Gibson, Richard G. and Burch, Neil and Bowling, Michael},
  booktitle = {International Conference on Machine Learning},
  year      = {2012},
  url       = {https://mlanthology.org/icml/2012/lanctot2012icml-regret/}
}