Auditing sex/gender disparities in emergency triage with LLM-based paired comparisons

We present a domain-agnostic paired-comparison approach that uses Large Language Models (LLMs) to quantify sex/gender-related asymmetries in documented clinical decision-making. The method trains an LLM to emulate observed decisions, then evaluates sex-swapped pairs in which only sex is flipped, holding documented clinical content constant. We apply it to emergency triage, analyzing more than 140,000 Bordeaux University Hospital (France) admissions and testing methodological portability on MIMIC-IV, spanning a different language, population, and healthcare system. Fine-tuning Mistral NeMo 12B for triage prediction and using Mistral Small 24B for pair generation, we find otherwise identical presentations were more likely to receive a lower-severity predicted score as female than male: 1.1% (95% CI 0.9-1.3) in the French cohort, 2.2% (1.7-2.7) in MIMIC-IV. Predictions are sensitive to both tabular and textual sex markers, with the asymmetry emerging primarily in the combined bimodal setting. A model retrained on sex-neutralized inputs eliminated the between-sex prediction gap, indicating the asymmetry is mediated by explicit sex markers. Patterns vary with nurse-patient sex concordance, suggesting the model captures stable features of the recorded data rather than random artifacts. These effects are small and documentation-level. We therefore present this as a methodological feasibility study: LLMs can serve as scalable probes of documented decisions, generating hypotheses rather than establishing bedside clinician behavior or clinically meaningful undertriage, which would require clinician-anchored validation. Beyond emergency care, the approach supports bias audits in other domains.

Authors

Institutions

Publication Details

Journal
npj Digital Medicine
Published
2026-09-16
DOI
https://doi.org/10.1038/s41746-026-03090-7
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Auditing sex/gender disparities in emergency triage with LLM-based paired comparisons

Cédric Gil‐Jardine, Marta Avalos-Fernandez, Leo Anthony Celi, Emmanuel Lagarde et al.
npj Digital Medicine
Topic Modeling
article

Auditing sex/gender disparities in emergency triage with LLM-based paired comparisons

Cédric Gil‐Jardine, Marta Avalos-Fernandez, Leo Anthony Celi, Emmanuel Lagarde, Ariel Guerra-Adames, Océane Dorémus
article en

Abstract

We present a domain-agnostic paired-comparison approach that uses Large Language Models (LLMs) to quantify sex/gender-related asymmetries in documented clinical decision-making. The method trains an LLM to emulate observed decisions, then evaluates sex-swapped pairs in which only sex is flipped, holding documented clinical content constant. We apply it to emergency triage, analyzing more than 140,000 Bordeaux University Hospital (France) admissions and testing methodological portability on MIMIC-IV, spanning a different language, population, and healthcare system. Fine-tuning Mistral NeMo 12B for triage prediction and using Mistral Small 24B for pair generation, we find otherwise identical presentations were more likely to receive a lower-severity predicted score as female than male: 1.1% (95% CI 0.9-1.3) in the French cohort, 2.2% (1.7-2.7) in MIMIC-IV. Predictions are sensitive to both tabular and textual sex markers, with the asymmetry emerging primarily in the combined bimodal setting. A model retrained on sex-neutralized inputs eliminated the between-sex prediction gap, indicating the asymmetry is mediated by explicit sex markers. Patterns vary with nurse-patient sex concordance, suggesting the model captures stable features of the recorded data rather than random artifacts. These effects are small and documentation-level. We therefore present this as a methodological feasibility study: LLMs can serve as scalable probes of documented decisions, generating hypotheses rather than establishing bedside clinician behavior or clinically meaningful undertriage, which would require clinician-anchored validation. Beyond emergency care, the approach supports bias audits in other domains.

npj Digital Medicine
Université de Bordeaux (FR), Centre Hospitalier Universitaire de Bordeaux (FR), Bordeaux Population Health (FR), Moscow Institute of Thermal Technology (RU), Hôpital Pellegrin (FR), Massachusetts Institute of Technology (US)
National Science Foundation, Korea Health Industry Development Institute, Université de Bordeaux
Openalex Percentile: Top 98%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.