A multi-site benchmarking framework for scalable extraction of geriatric care constructs from electronic health records

Current Natural Language Processing (NLP) algorithms for detecting geriatric conditions are largely limited to domain-specific models that fail to capture the interdependent, multidimensional nature of comprehensive geriatric assessment. This study aimed to develop and evaluate a comprehensive, scalable, and robust information extraction framework to identify Comprehensive Geriatric Assessment (CGA) and Age-Friendly Health Systems (AFHS) 4Ms-related data elements from unstructured electronic health record (EHR) text across multiple health systems. Using a team science approach grounded in the TRUST framework, we annotated pooled clinical notes from four health systems to produce a gold-standard dataset of 41 CGA- and 4Ms-related geriatric care data elements. Three information extraction approaches were implemented and evaluated: an in-context learning generative large language model (GPT-4o), a hybrid heuristic-LLM model (MedAgingIE), and an instruction-tuned open-source lightweight model (Qwen2-7B-Instruct). Performance was assessed on a blinded test set using macro- and micro-averaged metrics. GPT-4o achieved a macro F1-score of 0.56 and micro F1-score of 0.87; MedAgingIE achieved 0.55 and 0.92; and Qwen2-7B-Instruct achieved 0.30 and 0.81, respectively. MedAgingIE demonstrated the strongest consistency between precision and recall, while GPT-4o showed superior sensitivity for diverse, context-rich geriatric concepts. These findings highlight key trade-offs among symbolic, generative, and instruction-tuned approaches for CGA and 4Ms phenotyping, suggesting that hybrid heuristic-LLM methods offer interpretability and stability, whereas large language models provide greater adaptability for complex clinical narratives.

Authors

Institutions

Publication Details

Journal
npj Health Systems
Published
2026-09-16
DOI
https://doi.org/10.1038/s44401-026-00114-y
Primary Topic
Machine Learning in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A multi-site benchmarking framework for scalable extraction of geriatric care constructs from electronic health records

Sunyang Fu, Huiwen Xu, Min Ji Kwak, Keziah M. Thomas et al.
npj Health Systems
Machine Learning in Healthcare
article

A multi-site benchmarking framework for scalable extraction of geriatric care constructs from electronic health records

Sunyang Fu, Huiwen Xu, Min Ji Kwak, Keziah M. Thomas, Xiaoyang Ruan, Zhiyi Yue, Shreyans Sanghvi, Grace Giles, Alexa Cumming, Jaerong Ahn, Ming Huang, Nan Wang, Yanshan Wang, Erin Hommel, Jennifer St. Sauver, Nahid Rianon, Jiang Jun, Jeffrey S. Wefel, Dae Hyun Kim, Chan Mi Park, Qiuhao Lu, Liwei Wang, Lichao Sun, Huipeng Liu, Andrew Wen, Hongfang Liu
article en

Abstract

Current Natural Language Processing (NLP) algorithms for detecting geriatric conditions are largely limited to domain-specific models that fail to capture the interdependent, multidimensional nature of comprehensive geriatric assessment. This study aimed to develop and evaluate a comprehensive, scalable, and robust information extraction framework to identify Comprehensive Geriatric Assessment (CGA) and Age-Friendly Health Systems (AFHS) 4Ms-related data elements from unstructured electronic health record (EHR) text across multiple health systems. Using a team science approach grounded in the TRUST framework, we annotated pooled clinical notes from four health systems to produce a gold-standard dataset of 41 CGA- and 4Ms-related geriatric care data elements. Three information extraction approaches were implemented and evaluated: an in-context learning generative large language model (GPT-4o), a hybrid heuristic-LLM model (MedAgingIE), and an instruction-tuned open-source lightweight model (Qwen2-7B-Instruct). Performance was assessed on a blinded test set using macro- and micro-averaged metrics. GPT-4o achieved a macro F1-score of 0.56 and micro F1-score of 0.87; MedAgingIE achieved 0.55 and 0.92; and Qwen2-7B-Instruct achieved 0.30 and 0.81, respectively. MedAgingIE demonstrated the strongest consistency between precision and recall, while GPT-4o showed superior sensitivity for diverse, context-rich geriatric concepts. These findings highlight key trade-offs among symbolic, generative, and instruction-tuned approaches for CGA and 4Ms phenotyping, suggesting that hybrid heuristic-LLM methods offer interpretability and stability, whereas large language models provide greater adaptability for complex clinical narratives.

npj Health SystemsVol. 3(1)
Memorial Hermann (US), Beth Israel Deaconess Medical Center (US), Kaiser Permanente (US), The University of Texas MD Anderson Cancer Center (US), Emory University (US), University of Pittsburgh (US), Lehigh University (US), Hebrew SeniorLife (US), Mayo Clinic in Arizona (US), The University of Texas Health Science Center (US), The University of Texas Medical Branch at Galveston (US), The University of Texas at Austin (US), The University of Texas Health Science Center at Houston (US)
Quality Education
Openalex Percentile: Top 8%
Machine Learning in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.