CFR variant and iterate reporting in small imperfect-information games

Abstract Empirical comparisons of counterfactual regret minimization (CFR) variants often mix two distinct choices: the update rule used during training and the iterate reported at evaluation time. This note isolates those choices in an exact small-game setting where exploitability is computed by exact best response, so no measurement noise or approximation confounds the comparison. We formalise a shared-table scheduling framework in which fixed CFR variants appear as constant policies and phase schedules appear as non-constant policies; the formal result is a policy-class inclusion showing that, under a common update budget and an exact best-iterate oracle, the class contains every fixed (variant, iterate) baseline. This is deliberately an expressivity statement and carries no convergence guarantee for a trained controller. On Kuhn poker and Leduc hold’em, with every convention stated and a validated exploitability engine, the diagnostics show that the reporting choice does not behave as informal comparisons suggest: for all regret-matching variants (CFR, CFR+, LCFR, DCFR, ECFR) the time-averaged strategy is strictly less exploitable than the current iterate, which oscillates and does not converge; only the predictive method PCFR+ exhibits a converging current iterate, and only on the smaller game. Consequently, reporting the “better” iterate reduces in practice to reporting the average, and the update rule—together with the frequently unstated choice of alternating versus simultaneous updates, which alone shifts exploitability by one to two orders of magnitude—matters far more than the iterate choice. We contribute a controlled diagnostic protocol, a budget-corrected selection definition, explicit accounting of update cost and selection oracles, and a fully reproducible artifact: the complete solver, the exact best-response engine, and the raw logs and figure-generation code are deposited, so every number and figure regenerates from the solver rather than from tabulated values.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-17
DOI
https://doi.org/10.1038/s41598-026-71749-y
Primary Topic
Artificial Intelligence in Games
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CFR variant and iterate reporting in small imperfect-information games

Behbod Keshavarzi, Hamidreza Navidi
Scientific Reports
Artificial Intelligence in Games
article

CFR variant and iterate reporting in small imperfect-information games

Behbod Keshavarzi, Hamidreza Navidi
article en

Abstract

Abstract Empirical comparisons of counterfactual regret minimization (CFR) variants often mix two distinct choices: the update rule used during training and the iterate reported at evaluation time. This note isolates those choices in an exact small-game setting where exploitability is computed by exact best response, so no measurement noise or approximation confounds the comparison. We formalise a shared-table scheduling framework in which fixed CFR variants appear as constant policies and phase schedules appear as non-constant policies; the formal result is a policy-class inclusion showing that, under a common update budget and an exact best-iterate oracle, the class contains every fixed (variant, iterate) baseline. This is deliberately an expressivity statement and carries no convergence guarantee for a trained controller. On Kuhn poker and Leduc hold’em, with every convention stated and a validated exploitability engine, the diagnostics show that the reporting choice does not behave as informal comparisons suggest: for all regret-matching variants (CFR, CFR+, LCFR, DCFR, ECFR) the time-averaged strategy is strictly less exploitable than the current iterate, which oscillates and does not converge; only the predictive method PCFR+ exhibits a converging current iterate, and only on the smaller game. Consequently, reporting the “better” iterate reduces in practice to reporting the average, and the update rule—together with the frequently unstated choice of alternating versus simultaneous updates, which alone shifts exploitability by one to two orders of magnitude—matters far more than the iterate choice. We contribute a controlled diagnostic protocol, a budget-corrected selection definition, explicit accounting of update cost and selection oracles, and a fully reproducible artifact: the complete solver, the exact best-response engine, and the raw logs and figure-generation code are deposited, so every number and figure regenerates from the solver rather than from tabulated values.

Scientific Reports
Shahed University (IR)
Reduced inequalities
Openalex Percentile: Top 9%
Artificial Intelligence in Games
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.