Outcome-grounded effect of clinically stigmatizing information on large language model emergency triage prioritization
Whether stigmatizing input language biases large language model (LLM) triage prioritization against acutely ill patients is unknown. In a controlled, outcome-grounded experiment at two academic emergency departments, Emergency Severity Index-matched pairs (one deteriorating within 6 h, one not) were evaluated by three open-weight models (Gemma, Qwen, DeepSeek) before and after inserting one demographic, social, or stigma-related attribute. Eighteen conditions included neutral and stigmatizing formulations of the same concepts. Across 221,556 comparisons, the stigmatizing frequent-emergency-department-use formulation produced significant harmful reprioritization in all six model-dataset cells (up to 9.4%), exceeding its neutral counterpart within pairs in every cell (+1.4 to +7.1 points). Psychiatric history also produced significant shifts (up to 9.5%), though its stigmatizing-versus-neutral difference reached significance for only one model. Race, language, and insurance showed no consistent harmful shifts. Stigmatizing formulations of clinical information can bias LLM triage prioritization against deteriorating patients, making input selection and formulation safety-critical design choices.
Authors
- Philip Jarrett (ORCID: https://orcid.org/0000-0003-4534-8556)
- D. Mark Courtney
- Andrew R. Jamieson (ORCID: https://orcid.org/0000-0001-5416-9379)
- Peter Yun
- Emmanuel Ohuabunwa
- Doreen Agboh
Institutions
- The University of Texas Southwestern Medical Center (US)
Publication Details
- Journal
- npj Digital Medicine
- Published
- 2026-09-19
- DOI
- https://doi.org/10.1038/s41746-026-03270-5
- Primary Topic
- Emergency and Acute Care Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00