Algorithmic bias in clinical resource allocation by a large language model: a cross-sectional in-silico evaluation of 13,608 decisions

BACKGROUND: Large language models (LLMs) hold the potential to transform clinical medicine, but their capacity to amplify societal biases is a significant concern, particularly in the ethically sensitive domain of clinical resource allocation. A critical knowledge gap exists in the quantitative understanding of how these models weigh non-clinical factors in forced-choice ethical dilemmas. This study provides a large-scale, systematic evaluation of a leading LLM to measure its implicit biases in resource allocation decisions. METHODS: We conducted a cross-sectional, in-silico analysis using OpenAI's GPT-5 model. Seven clinical vignettes representing scarce resource scenarios were paired with 3,888 unique patient profiles generated from a full-factorial combination of seven demographic and social variables (age, gender, race, income, dependents, education, social support). This resulted in 13,608 unique A/B patient comparisons. We used pooled logistic and linear regression models to identify the primary drivers of the model's allocation choices and quantify the magnitude of biases. The stability and internal consistency of the model's judgments were also assessed. RESULTS: The LLM's decisions were governed by a distinct hierarchy of non-clinical factors. Patient age, income, and number of dependents were the most powerful predictors. Being 25 years old increased the odds of selection more than threefold compared to a 50-year-old (OR 3.31; 95% CI, 2.96-3.71; p < 0.001), while earning over $5 million annually reduced the odds by 75% (OR 0.25; 95% CI, 0.22-0.28; p < 0.001). The model also systematically favored racial minorities, women, and non-binary individuals over White men. Analysis revealed significant internal inconsistency, with the model's chosen patient matching its own higher-priority-scored patient in only 66.3% of cases overall. The influence of biases was highly context-dependent yet showed stochastic stability upon repeat testing. CONCLUSIONS: The tested model operates with a strong, opaque, and context-dependent set of non-clinical biases when faced with ethical dilemmas. These findings suggest that this architecture's inconsistent and unpredictable ethical framework makes its direct integration into clinical decision support for resource allocation untenable at this stage. Rigorous standards for safety, transparency, and ethical validation are required before these technologies can be responsibly deployed in high-stakes clinical settings.

Authors

Institutions

Publication Details

Journal
BMC Medical Ethics
Published
2026-05-22
DOI
https://doi.org/10.1186/s12910-026-01496-2
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Algorithmic bias in clinical resource allocation by a large language model: a cross-sectional in-silico evaluation of 13,608 decisions

Michael Balas, Siddharth Gandhi
BMC Medical Ethics
Artificial Intelligence in Healthcare and Education
article

Algorithmic bias in clinical resource allocation by a large language model: a cross-sectional in-silico evaluation of 13,608 decisions

Michael Balas, Siddharth Gandhi
article en

Abstract

BACKGROUND: Large language models (LLMs) hold the potential to transform clinical medicine, but their capacity to amplify societal biases is a significant concern, particularly in the ethically sensitive domain of clinical resource allocation. A critical knowledge gap exists in the quantitative understanding of how these models weigh non-clinical factors in forced-choice ethical dilemmas. This study provides a large-scale, systematic evaluation of a leading LLM to measure its implicit biases in resource allocation decisions. METHODS: We conducted a cross-sectional, in-silico analysis using OpenAI's GPT-5 model. Seven clinical vignettes representing scarce resource scenarios were paired with 3,888 unique patient profiles generated from a full-factorial combination of seven demographic and social variables (age, gender, race, income, dependents, education, social support). This resulted in 13,608 unique A/B patient comparisons. We used pooled logistic and linear regression models to identify the primary drivers of the model's allocation choices and quantify the magnitude of biases. The stability and internal consistency of the model's judgments were also assessed. RESULTS: The LLM's decisions were governed by a distinct hierarchy of non-clinical factors. Patient age, income, and number of dependents were the most powerful predictors. Being 25 years old increased the odds of selection more than threefold compared to a 50-year-old (OR 3.31; 95% CI, 2.96-3.71; p < 0.001), while earning over $5 million annually reduced the odds by 75% (OR 0.25; 95% CI, 0.22-0.28; p < 0.001). The model also systematically favored racial minorities, women, and non-binary individuals over White men. Analysis revealed significant internal inconsistency, with the model's chosen patient matching its own higher-priority-scored patient in only 66.3% of cases overall. The influence of biases was highly context-dependent yet showed stochastic stability upon repeat testing. CONCLUSIONS: The tested model operates with a strong, opaque, and context-dependent set of non-clinical biases when faced with ethical dilemmas. These findings suggest that this architecture's inconsistent and unpredictable ethical framework makes its direct integration into clinical decision support for resource allocation untenable at this stage. Rigorous standards for safety, transparency, and ethical validation are required before these technologies can be responsibly deployed in high-stakes clinical settings.

BMC Medical Ethics
St. Michael's Hospital (CA), University of Toronto (CA), Queen's University (CA)
No poverty
Openalex Percentile: Top 7%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.