A multi-site benchmarking framework for scalable extraction of geriatric care constructs from electronic health records
Current Natural Language Processing (NLP) algorithms for detecting geriatric conditions are largely limited to domain-specific models that fail to capture the interdependent, multidimensional nature of comprehensive geriatric assessment. This study aimed to develop and evaluate a comprehensive, scalable, and robust information extraction framework to identify Comprehensive Geriatric Assessment (CGA) and Age-Friendly Health Systems (AFHS) 4Ms-related data elements from unstructured electronic health record (EHR) text across multiple health systems. Using a team science approach grounded in the TRUST framework, we annotated pooled clinical notes from four health systems to produce a gold-standard dataset of 41 CGA- and 4Ms-related geriatric care data elements. Three information extraction approaches were implemented and evaluated: an in-context learning generative large language model (GPT-4o), a hybrid heuristic-LLM model (MedAgingIE), and an instruction-tuned open-source lightweight model (Qwen2-7B-Instruct). Performance was assessed on a blinded test set using macro- and micro-averaged metrics. GPT-4o achieved a macro F1-score of 0.56 and micro F1-score of 0.87; MedAgingIE achieved 0.55 and 0.92; and Qwen2-7B-Instruct achieved 0.30 and 0.81, respectively. MedAgingIE demonstrated the strongest consistency between precision and recall, while GPT-4o showed superior sensitivity for diverse, context-rich geriatric concepts. These findings highlight key trade-offs among symbolic, generative, and instruction-tuned approaches for CGA and 4Ms phenotyping, suggesting that hybrid heuristic-LLM methods offer interpretability and stability, whereas large language models provide greater adaptability for complex clinical narratives.
Authors
- Sunyang Fu (ORCID: https://orcid.org/0000-0003-1691-5179)
- Huiwen Xu (ORCID: https://orcid.org/0000-0002-4033-0659)
- Min Ji Kwak (ORCID: https://orcid.org/0000-0003-2778-3984)
- Keziah M. Thomas
- Xiaoyang Ruan (ORCID: https://orcid.org/0009-0006-7085-744X)
- Zhiyi Yue (ORCID: https://orcid.org/0009-0002-3440-0902)
- Shreyans Sanghvi
- Grace Giles
- Alexa Cumming
- Jaerong Ahn
- Ming Huang
- Nan Wang
- Yanshan Wang
- Erin Hommel
- Jennifer St. Sauver
- Nahid Rianon
- Jiang Jun
- Jeffrey S. Wefel
- Dae Hyun Kim
- Chan Mi Park
- Qiuhao Lu
- Liwei Wang
- Lichao Sun
- Huipeng Liu
- Andrew Wen
- Hongfang Liu
Institutions
- Memorial Hermann (US)
- Beth Israel Deaconess Medical Center (US)
- Kaiser Permanente (US)
- The University of Texas MD Anderson Cancer Center (US)
- Emory University (US)
- University of Pittsburgh (US)
- Lehigh University (US)
- Hebrew SeniorLife (US)
- Mayo Clinic in Arizona (US)
- The University of Texas Health Science Center (US)
- The University of Texas Medical Branch at Galveston (US)
- The University of Texas at Austin (US)
- The University of Texas Health Science Center at Houston (US)
Publication Details
- Journal
- npj Health Systems
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1038/s44401-026-00114-y
- Primary Topic
- Machine Learning in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00