CFR variant and iterate reporting in small imperfect-information games
Abstract Empirical comparisons of counterfactual regret minimization (CFR) variants often mix two distinct choices: the update rule used during training and the iterate reported at evaluation time. This note isolates those choices in an exact small-game setting where exploitability is computed by exact best response, so no measurement noise or approximation confounds the comparison. We formalise a shared-table scheduling framework in which fixed CFR variants appear as constant policies and phase schedules appear as non-constant policies; the formal result is a policy-class inclusion showing that, under a common update budget and an exact best-iterate oracle, the class contains every fixed (variant, iterate) baseline. This is deliberately an expressivity statement and carries no convergence guarantee for a trained controller. On Kuhn poker and Leduc hold’em, with every convention stated and a validated exploitability engine, the diagnostics show that the reporting choice does not behave as informal comparisons suggest: for all regret-matching variants (CFR, CFR+, LCFR, DCFR, ECFR) the time-averaged strategy is strictly less exploitable than the current iterate, which oscillates and does not converge; only the predictive method PCFR+ exhibits a converging current iterate, and only on the smaller game. Consequently, reporting the “better” iterate reduces in practice to reporting the average, and the update rule—together with the frequently unstated choice of alternating versus simultaneous updates, which alone shifts exploitability by one to two orders of magnitude—matters far more than the iterate choice. We contribute a controlled diagnostic protocol, a budget-corrected selection definition, explicit accounting of update cost and selection oracles, and a fully reproducible artifact: the complete solver, the exact best-response engine, and the raw logs and figure-generation code are deposited, so every number and figure regenerates from the solver rather than from tabulated values.
Authors
- Behbod Keshavarzi
- Hamidreza Navidi (ORCID: https://orcid.org/0000-0003-1072-8786)
Institutions
- Shahed University (IR)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1038/s41598-026-71749-y
- Primary Topic
- Artificial Intelligence in Games
- Type
- article
- Field-Weighted Citation Impact
- 0.00