Optimization of coarsened exact matching through random forest classification

Abstract Background Coarsened exact matching (CEM) was developed to resolve the “matching paradox” in propensity score matching (PSM), where narrower caliper widths can paradoxically worsen covariate balance. However, its application to complex medical data faces practical challenges, including the need for subjective decisions when categorizing continuous variables and substantial reductions in sample size when multiple variables are handled. Methods We investigated several data-driven strategies for determining coarsening rules using classification-tree models. Specifically, we explored coarsening informed by classification and regression trees (CART), random forests (RF), and the most representative tree (MRT). The MRT provides a single representative tree that can be used to summarize and interpret the splitting structure learned by RF, thereby addressing the limited interpretability of ensemble partitions. Simulation studies comparing different matching procedures were conducted across several scenarios, with varying degrees of similarity between the propensity score and outcome models, as well as varying complexities in each modeling structure. To demonstrate the practical utility of this framework, we applied it to the Framingham Heart Study dataset. Results For the tree-based CEM methods, covariate imbalance and bias generally decreased as matching became more stringent, reflecting the monotonic behavior expected from CEM-based matching. PSM often achieved the smallest bias when a large proportion of subjects was retained, but stricter propensity-score-based matching did not necessarily improve estimation performance. In the Framingham study, PSM and the CEM-based approaches produced broadly similar effect estimates; tree-based CEM preserved the full sample size, whereas standard CEM incurred sample loss. Conclusions Integrating random forest with CEM may help mitigate key limitations of traditional CEM, particularly those related to subjective coarsening choices and sample loss. By providing an interpretable data-driven coarsening strategy, CEM+MRT may offer a practical option for implementing CEM when transparent coarsening rules are desired.

Authors

Publication Details

Journal
BMC Medical Research Methodology
Published
2026-09-18
DOI
https://doi.org/10.1186/s12874-026-03011-y
Primary Topic
Advanced Causal Inference Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Optimization of coarsened exact matching through random forest classification

Ryo Mishima, Daisuke Koide
BMC Medical Research Methodology
Advanced Causal Inference Techniques
article

Optimization of coarsened exact matching through random forest classification

Ryo Mishima, Daisuke Koide
article en

Abstract

Abstract Background Coarsened exact matching (CEM) was developed to resolve the “matching paradox” in propensity score matching (PSM), where narrower caliper widths can paradoxically worsen covariate balance. However, its application to complex medical data faces practical challenges, including the need for subjective decisions when categorizing continuous variables and substantial reductions in sample size when multiple variables are handled. Methods We investigated several data-driven strategies for determining coarsening rules using classification-tree models. Specifically, we explored coarsening informed by classification and regression trees (CART), random forests (RF), and the most representative tree (MRT). The MRT provides a single representative tree that can be used to summarize and interpret the splitting structure learned by RF, thereby addressing the limited interpretability of ensemble partitions. Simulation studies comparing different matching procedures were conducted across several scenarios, with varying degrees of similarity between the propensity score and outcome models, as well as varying complexities in each modeling structure. To demonstrate the practical utility of this framework, we applied it to the Framingham Heart Study dataset. Results For the tree-based CEM methods, covariate imbalance and bias generally decreased as matching became more stringent, reflecting the monotonic behavior expected from CEM-based matching. PSM often achieved the smallest bias when a large proportion of subjects was retained, but stricter propensity-score-based matching did not necessarily improve estimation performance. In the Framingham study, PSM and the CEM-based approaches produced broadly similar effect estimates; tree-based CEM preserved the full sample size, whereas standard CEM incurred sample loss. Conclusions Integrating random forest with CEM may help mitigate key limitations of traditional CEM, particularly those related to subjective coarsening choices and sample loss. By providing an interpretable data-driven coarsening strategy, CEM+MRT may offer a practical option for implementing CEM when transparent coarsening rules are desired.

BMC Medical Research Methodology
Life in Land
Openalex Percentile: Top 8%
Advanced Causal Inference Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Optimization of coarsened exact matching through random forest classification — Ryo Mishima, Daisuke Koide · BMC Medical Research Methodology (2026) | TGRS Research Map | TGRS