A Scalable Method for Validated Data Extraction from Electronic Health Records with Large Language Models
Purpose/Background: Health care organizations increasingly require structured, patient-level clinical variables for treatment decisions, operational workflows, quality measurement, and clinical trial screening. Relevant information is often fragmented across heterogeneous electronic health record (EHR) systems, unstructured formats, and scanned documents. Large language models (LLMs) offer an opportunity to enhance medical record extraction, particularly in oncology where point-of-care structured coding rarely captures a patient’s full longitudinal clinical course. Methods: Two complementary LLM-based approaches were developed. The first performed schema-based Named Entity Recognition and Relation Extraction from unstructured documents, normalizing outputs to Fast Healthcare Interoperability Resources and Observational Health Data Sciences and Informatics vocabularies. The second used a retrieval-augmented checklist framework that queried embedded document text alongside structured data, returning custom outputs with evidence-based justifications and source citations. Performance was evaluated through human validation, automated consistency checks, and iterative error analysis. Results: Schema-based LLM extractors for medications, radiation, and surgical procedures achieved ∼95% accuracy, precision, recall, and F1. Deployed across 3,493 patients, LLM extraction yielded a 207% increase in distinct oncology drug ingredients over structured EHR sources. Oncology therapies were captured for 65% of patients versus 40% in structured data, with dramatically improved attribute coverage: indication for prescription (71% vs. 5%) and reason for discontinuation (10% vs. 0%). For radiation and oncology surgical procedures, LLM extraction yielded more than triple the patient coverage versus structured sources. The checklist framework achieved strong F1 scores in extracting cancer diagnosis variables (99.0%) and lines of therapy (97.6%) across 4,802 validated elements. Cancer diagnoses and dates were identified for 93.5% of patients versus 69.8% in structured data; stage and grade were extracted for 64% and 62% versus near-zero structured availability. A lines-of-therapy checklist generated 4,218 treatment lines across 2,320 patients, capturing regimen dates, best response, discontinuation reasons, and progression dates. Among patients with response data, objective response rate declined from 65% at first line to 22% at third line, with first-line response rates ranging from 43% in colorectal to 85% in esophageal cancer. Conclusions: Two complementary LLM-based strategies substantially enhanced the completeness and utility of structured clinical data from heterogeneous medical records, supporting scalable generation of interoperable patient-level datasets for clinical analytics, research, and operational workflows.
Authors
- John M. Furgason
- Hiba Kouser (ORCID: https://orcid.org/0000-0002-3052-6113)
- Mika E. Newton (ORCID: https://orcid.org/0009-0009-4055-655X)
- Mark A. Shapiro (ORCID: https://orcid.org/0000-0002-7611-5432)
- Glenn A. Kramer
- Jeff Rapp
- Timothy Joseph Stuhlmiller (ORCID: https://orcid.org/0009-0005-0946-2257)
- Frank J. Scarpa
- Santosh Kesari (ORCID: https://orcid.org/0000-0003-3772-6000)
- Kristi Lui
- Hugh Salamon (ORCID: https://orcid.org/0000-0003-1472-5785)
- Madhuri Paul (ORCID: https://orcid.org/0009-0008-3421-0897)
- Kenny K. Wong (ORCID: https://orcid.org/0009-0009-7405-3934)
- Alaa Awawda
- Donald Chuyka
- William Mahoney
- AJ Rabe
Institutions
- NeoGenomics (United States) (US)
Publication Details
- Journal
- AI in Precision Oncology
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1177/2993091x261492372
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00