A Scalable Method for Validated Data Extraction from Electronic Health Records with Large Language Models

Purpose/Background: Health care organizations increasingly require structured, patient-level clinical variables for treatment decisions, operational workflows, quality measurement, and clinical trial screening. Relevant information is often fragmented across heterogeneous electronic health record (EHR) systems, unstructured formats, and scanned documents. Large language models (LLMs) offer an opportunity to enhance medical record extraction, particularly in oncology where point-of-care structured coding rarely captures a patient’s full longitudinal clinical course. Methods: Two complementary LLM-based approaches were developed. The first performed schema-based Named Entity Recognition and Relation Extraction from unstructured documents, normalizing outputs to Fast Healthcare Interoperability Resources and Observational Health Data Sciences and Informatics vocabularies. The second used a retrieval-augmented checklist framework that queried embedded document text alongside structured data, returning custom outputs with evidence-based justifications and source citations. Performance was evaluated through human validation, automated consistency checks, and iterative error analysis. Results: Schema-based LLM extractors for medications, radiation, and surgical procedures achieved ∼95% accuracy, precision, recall, and F1. Deployed across 3,493 patients, LLM extraction yielded a 207% increase in distinct oncology drug ingredients over structured EHR sources. Oncology therapies were captured for 65% of patients versus 40% in structured data, with dramatically improved attribute coverage: indication for prescription (71% vs. 5%) and reason for discontinuation (10% vs. 0%). For radiation and oncology surgical procedures, LLM extraction yielded more than triple the patient coverage versus structured sources. The checklist framework achieved strong F1 scores in extracting cancer diagnosis variables (99.0%) and lines of therapy (97.6%) across 4,802 validated elements. Cancer diagnoses and dates were identified for 93.5% of patients versus 69.8% in structured data; stage and grade were extracted for 64% and 62% versus near-zero structured availability. A lines-of-therapy checklist generated 4,218 treatment lines across 2,320 patients, capturing regimen dates, best response, discontinuation reasons, and progression dates. Among patients with response data, objective response rate declined from 65% at first line to 22% at third line, with first-line response rates ranging from 43% in colorectal to 85% in esophageal cancer. Conclusions: Two complementary LLM-based strategies substantially enhanced the completeness and utility of structured clinical data from heterogeneous medical records, supporting scalable generation of interoperable patient-level datasets for clinical analytics, research, and operational workflows.

Authors

Institutions

Publication Details

Journal
AI in Precision Oncology
Published
2026-09-29
DOI
https://doi.org/10.1177/2993091x261492372
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Scalable Method for Validated Data Extraction from Electronic Health Records with Large Language Models

John M. Furgason, Hiba Kouser, Mika E. Newton, Mark A. Shapiro et al.
AI in Precision Oncology
Topic Modeling
article

A Scalable Method for Validated Data Extraction from Electronic Health Records with Large Language Models

John M. Furgason, Hiba Kouser, Mika E. Newton, Mark A. Shapiro, Glenn A. Kramer, Jeff Rapp, Timothy Joseph Stuhlmiller, Frank J. Scarpa, Santosh Kesari, Kristi Lui, Hugh Salamon, Madhuri Paul, Kenny K. Wong, Alaa Awawda, Donald Chuyka, William Mahoney, AJ Rabe
article en

Abstract

Purpose/Background: Health care organizations increasingly require structured, patient-level clinical variables for treatment decisions, operational workflows, quality measurement, and clinical trial screening. Relevant information is often fragmented across heterogeneous electronic health record (EHR) systems, unstructured formats, and scanned documents. Large language models (LLMs) offer an opportunity to enhance medical record extraction, particularly in oncology where point-of-care structured coding rarely captures a patient’s full longitudinal clinical course. Methods: Two complementary LLM-based approaches were developed. The first performed schema-based Named Entity Recognition and Relation Extraction from unstructured documents, normalizing outputs to Fast Healthcare Interoperability Resources and Observational Health Data Sciences and Informatics vocabularies. The second used a retrieval-augmented checklist framework that queried embedded document text alongside structured data, returning custom outputs with evidence-based justifications and source citations. Performance was evaluated through human validation, automated consistency checks, and iterative error analysis. Results: Schema-based LLM extractors for medications, radiation, and surgical procedures achieved ∼95% accuracy, precision, recall, and F1. Deployed across 3,493 patients, LLM extraction yielded a 207% increase in distinct oncology drug ingredients over structured EHR sources. Oncology therapies were captured for 65% of patients versus 40% in structured data, with dramatically improved attribute coverage: indication for prescription (71% vs. 5%) and reason for discontinuation (10% vs. 0%). For radiation and oncology surgical procedures, LLM extraction yielded more than triple the patient coverage versus structured sources. The checklist framework achieved strong F1 scores in extracting cancer diagnosis variables (99.0%) and lines of therapy (97.6%) across 4,802 validated elements. Cancer diagnoses and dates were identified for 93.5% of patients versus 69.8% in structured data; stage and grade were extracted for 64% and 62% versus near-zero structured availability. A lines-of-therapy checklist generated 4,218 treatment lines across 2,320 patients, capturing regimen dates, best response, discontinuation reasons, and progression dates. Among patients with response data, objective response rate declined from 65% at first line to 22% at third line, with first-line response rates ranging from 43% in colorectal to 85% in esophageal cancer. Conclusions: Two complementary LLM-based strategies substantially enhanced the completeness and utility of structured clinical data from heterogeneous medical records, supporting scalable generation of interoperable patient-level datasets for clinical analytics, research, and operational workflows.

AI in Precision Oncology
NeoGenomics (United States) (US)
Quality Education
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.