Structured phenotypes from unstructured reports: large language model extraction of chest CT identifies mortality subgroups in emergency department dyspnea

Acute dyspnea is among the most heterogeneous emergency department (ED) presentations, and no validated framework exists for postdiagnostic risk stratification across its underlying causes. Chest CT reports describe the intrathoracic organs well beyond the presenting complaint, yet this information is rarely reused. We aimed to derive and externally validate mortality phenotypes in ED dyspnea by combining routine laboratory data with large language model (LLM) extraction of chest CT reports. In a multicenter retrospective cohort of 51,843 adults (2016–2025) with acute dyspnea undergoing chest CT at three Korean teaching hospitals, claude-sonnet-4-6 extracted 18 imaging findings per report using a prespecified JSON schema. Phenotypes were derived in one hospital ( n = 12,702) by k-means clustering of 12 routine clinical and laboratory variables plus graded imaging features, then applied unchanged to two external validation hospitals ( n = 17,395; n = 21,746) by frozen-centroid nearest-cluster mapping. We also compared graded (0–3) features with binary features derived from the same extraction by thresholding at severity ≥ 1, holding every downstream step identical. Three phenotypes reproduced across all three hospitals: Well (37.9–40.6%), Inflammatory-Parenchymal (21.1–25.3%), and Cardiac-Vascular-Congestive (34.1–40.5%). Compared with Well, adjusted hazard ratios were 2.43–3.90 for Inflammatory-Parenchymal and 2.81–3.44 for Cardiac-Vascular-Congestive (all P < .001); unadjusted ratios were 2.77–5.11 and 3.27–4.70. Phenotype membership was largely unchanged when age was removed from the feature space (adjusted Rand index 0.80–0.82), and the phenotypes were recovered in pre-pandemic, peak-pandemic, and post-pandemic strata. Graded and binary features gave similar incremental discrimination: adding phenotype indicators to an age- and sex-adjusted Cox baseline raised Harrell C-index by 0.041–0.054 with graded and 0.040–0.053 with binary features (paired difference 0.0010–0.0022). Routine chest CT reports, converted to structured data by an LLM and combined with routine laboratory values, identify reproducible mortality phenotypes in ED dyspnea that carry age- and sex-independent prognostic information across hospitals with differing case-mix. Binary features derived from the same extraction performed similarly, so retaining ordinal severity added little in this pipeline.

Authors

Institutions

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-10-09
DOI
https://doi.org/10.1186/s12911-026-03901-7
Primary Topic
Machine Learning in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Structured phenotypes from unstructured reports: large language model extraction of chest CT identifies mortality subgroups in emergency department dyspnea

차경만, Jee Yong Lim, H J Kim, Jaekwang Shin et al.
BMC Medical Informatics and Decision Making
Machine Learning in Healthcare
article

Structured phenotypes from unstructured reports: large language model extraction of chest CT identifies mortality subgroups in emergency department dyspnea

차경만, Jee Yong Lim, H J Kim, Jaekwang Shin, Jae Hun Oh, Sohee Lee
article en

Abstract

Acute dyspnea is among the most heterogeneous emergency department (ED) presentations, and no validated framework exists for postdiagnostic risk stratification across its underlying causes. Chest CT reports describe the intrathoracic organs well beyond the presenting complaint, yet this information is rarely reused. We aimed to derive and externally validate mortality phenotypes in ED dyspnea by combining routine laboratory data with large language model (LLM) extraction of chest CT reports. In a multicenter retrospective cohort of 51,843 adults (2016–2025) with acute dyspnea undergoing chest CT at three Korean teaching hospitals, claude-sonnet-4-6 extracted 18 imaging findings per report using a prespecified JSON schema. Phenotypes were derived in one hospital ( n = 12,702) by k-means clustering of 12 routine clinical and laboratory variables plus graded imaging features, then applied unchanged to two external validation hospitals ( n = 17,395; n = 21,746) by frozen-centroid nearest-cluster mapping. We also compared graded (0–3) features with binary features derived from the same extraction by thresholding at severity ≥ 1, holding every downstream step identical. Three phenotypes reproduced across all three hospitals: Well (37.9–40.6%), Inflammatory-Parenchymal (21.1–25.3%), and Cardiac-Vascular-Congestive (34.1–40.5%). Compared with Well, adjusted hazard ratios were 2.43–3.90 for Inflammatory-Parenchymal and 2.81–3.44 for Cardiac-Vascular-Congestive (all P < .001); unadjusted ratios were 2.77–5.11 and 3.27–4.70. Phenotype membership was largely unchanged when age was removed from the feature space (adjusted Rand index 0.80–0.82), and the phenotypes were recovered in pre-pandemic, peak-pandemic, and post-pandemic strata. Graded and binary features gave similar incremental discrimination: adding phenotype indicators to an age- and sex-adjusted Cox baseline raised Harrell C-index by 0.041–0.054 with graded and 0.040–0.053 with binary features (paired difference 0.0010–0.0022). Routine chest CT reports, converted to structured data by an LLM and combined with routine laboratory values, identify reproducible mortality phenotypes in ED dyspnea that carry age- and sex-independent prognostic information across hospitals with differing case-mix. Binary features derived from the same extraction performed similarly, so retaining ordinal severity added little in this pipeline.

BMC Medical Informatics and Decision Making
St. Mary's Hospital (US), The Catholic University of Korea St. Vincent's Hospital (KR), The Catholic University of Korea Seoul St. Mary's Hospital (KR), Seokyeong University (KR)
Openalex Percentile: Top 12%
Machine Learning in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.