Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

Abstract Persistent low retention and completion rates in medications for opioid use disorder (MOUD) have driven the use of machine learning (ML) models to predict retention and identify patients at risk of premature discontinuation. However, the fairness of these models across patient populations remains largely unexplored, raising concerns about their application in treatment decision support. This study systematically assesses algorithmic fairness in ML models for predicting MOUD retention and premature discontinuation and investigates the effectiveness of bias mitigation techniques. Using the cross-sectional Treatment Episode Data Set–Discharges (TEDS-D), which includes treatment episodes for individuals in the U.S. discharged between 2015 and 2019, we trained four ML models to predict premature treatment discontinuation and retention beyond 180 days among individuals receiving outpatient MOUD. We evaluated overall performance and subgroup-level error rates across patient subgroups defined by race, ethnicity, age, and sex, complemented by model explanation analyses. We further assessed pre-processing, in-processing, and post-processing bias mitigation techniques and their effects on both fairness and predictive performance. The models exhibited substantial performance differences across patient subgroups, including overestimation of the likelihood of premature discontinuation for Black patients and of treatment retention beyond 180 days for older patients. Model explanation analyses further identified race and age as influential predictors, but their impacts on model predictions varied substantially across patient subgroups. Bias mitigation strategies reduced specific fairness gaps but often introduced trade-offs, such as increased error rates for other subgroups or reductions in overall predictive performance. These findings demonstrate that ML models for MOUD outcome prediction can exhibit subgroup-level performance gaps even when overall predictive performance appears acceptable and that bias mitigation can reduce, but not fully eliminate, these gaps without trade-offs. By demonstrating the importance of fairness-aware evaluation and transparent reporting of subgroup performance, this study provides practical insights for the responsible and context-sensitive use of ML models for risk stratification and care prioritization in MOUD treatment settings.

Authors

Institutions

Publication Details

Journal
Journal of Healthcare Informatics Research
Published
2026-08-24
DOI
https://doi.org/10.1007/s41666-026-00251-x
Primary Topic
Opioid Use Disorder Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

Tongnian Wang, Carolina Vivas-Valencia, Yuanxiong Guo, Yanmin Gong et al.
Journal of Healthcare Informatics Research
Opioid Use Disorder Treatment
article

Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

Tongnian Wang, Carolina Vivas-Valencia, Yuanxiong Guo, Yanmin Gong, Cici Bauer, Kim-Kwang Raymond Choo
article en

Abstract

Abstract Persistent low retention and completion rates in medications for opioid use disorder (MOUD) have driven the use of machine learning (ML) models to predict retention and identify patients at risk of premature discontinuation. However, the fairness of these models across patient populations remains largely unexplored, raising concerns about their application in treatment decision support. This study systematically assesses algorithmic fairness in ML models for predicting MOUD retention and premature discontinuation and investigates the effectiveness of bias mitigation techniques. Using the cross-sectional Treatment Episode Data Set–Discharges (TEDS-D), which includes treatment episodes for individuals in the U.S. discharged between 2015 and 2019, we trained four ML models to predict premature treatment discontinuation and retention beyond 180 days among individuals receiving outpatient MOUD. We evaluated overall performance and subgroup-level error rates across patient subgroups defined by race, ethnicity, age, and sex, complemented by model explanation analyses. We further assessed pre-processing, in-processing, and post-processing bias mitigation techniques and their effects on both fairness and predictive performance. The models exhibited substantial performance differences across patient subgroups, including overestimation of the likelihood of premature discontinuation for Black patients and of treatment retention beyond 180 days for older patients. Model explanation analyses further identified race and age as influential predictors, but their impacts on model predictions varied substantially across patient subgroups. Bias mitigation strategies reduced specific fairness gaps but often introduced trade-offs, such as increased error rates for other subgroups or reductions in overall predictive performance. These findings demonstrate that ML models for MOUD outcome prediction can exhibit subgroup-level performance gaps even when overall predictive performance appears acceptable and that bias mitigation can reduce, but not fully eliminate, these gaps without trade-offs. By demonstrating the importance of fairness-aware evaluation and transparent reporting of subgroup performance, this study provides practical insights for the responsible and context-sensitive use of ML models for risk stratification and care prioritization in MOUD treatment settings.

Journal of Healthcare Informatics Research
Texas A&M University – San Antonio (US), University of Tennessee at Chattanooga (US), The University of Texas at San Antonio (US), Texas A&M University (US), The University of Texas Health Science Center at Houston (US)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Opioid Use Disorder Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.