Large Language Model Data Abstraction Demonstrates Accuracy and Reliability for NSQIP

BACKGROUND: National Surgical Quality Improvement Program (NSQIP) data collection depends on labor-intensive manual chart abstraction, limiting efficiency, increasing cost, and necessitating patient sampling. This study evaluated whether a large language model (LLM) could accurately abstract unstructured NSQIP breast reconstruction variables compared with conventional human abstraction. STUDY DESIGN: Clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024-February 28, 2025) were manually de-identified and processed using a customized ChatGPT 4.1 workflow targeting individual variables. A faculty plastic surgeon established the reference standard. Overall accuracy of LLM and human abstraction was compared using McNemar's and Chi-square tests. RESULTS: Among 105 patients (73 bilateral, 32 unilateral), 9,048 data points were evaluated. Overall abstraction accuracy was 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction (McNemar p<0.001; Chi-square p<0.001). LLM performance exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables. The most frequent LLM error involved prior breast surgical history (29/61 errors), followed by prepectoral versus subpectoral implant or expander placement, a variable frequently requiring inference from documentation. CONCLUSIONS: In this proof-of-concept validation study, a customized LLM achieved significantly higher abstraction accuracy than conventional human review for general and breast reconstruction NSQIP variables. These findings support LLM-assisted abstraction as a promising approach to improve efficiency, reduce resource requirements, and facilitate broader implementation and expansion of NSQIP, although multicenter validation remains necessary.

Authors

Institutions

Publication Details

Journal
Journal of the American College of Surgeons
Published
2026-09-18
DOI
https://doi.org/10.1097/xcs.0000000000002211
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Large Language Model Data Abstraction Demonstrates Accuracy and Reliability for NSQIP

Alaina Matthews, Michael R. DeLong, Tahera Alnaseri, Mehrnaz Siavoshi et al.
Journal of the American College of Surgeons
Artificial Intelligence in Healthcare and Education
article

Large Language Model Data Abstraction Demonstrates Accuracy and Reliability for NSQIP

Alaina Matthews, Michael R. DeLong, Tahera Alnaseri, Mehrnaz Siavoshi, Ansgar Grunseid, Yasmine Ibrahim, Clifford Y Ko
article en

Abstract

BACKGROUND: National Surgical Quality Improvement Program (NSQIP) data collection depends on labor-intensive manual chart abstraction, limiting efficiency, increasing cost, and necessitating patient sampling. This study evaluated whether a large language model (LLM) could accurately abstract unstructured NSQIP breast reconstruction variables compared with conventional human abstraction. STUDY DESIGN: Clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024-February 28, 2025) were manually de-identified and processed using a customized ChatGPT 4.1 workflow targeting individual variables. A faculty plastic surgeon established the reference standard. Overall accuracy of LLM and human abstraction was compared using McNemar's and Chi-square tests. RESULTS: Among 105 patients (73 bilateral, 32 unilateral), 9,048 data points were evaluated. Overall abstraction accuracy was 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction (McNemar p<0.001; Chi-square p<0.001). LLM performance exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables. The most frequent LLM error involved prior breast surgical history (29/61 errors), followed by prepectoral versus subpectoral implant or expander placement, a variable frequently requiring inference from documentation. CONCLUSIONS: In this proof-of-concept validation study, a customized LLM achieved significantly higher abstraction accuracy than conventional human review for general and breast reconstruction NSQIP variables. These findings support LLM-assisted abstraction as a promising approach to improve efficiency, reduce resource requirements, and facilitate broader implementation and expansion of NSQIP, although multicenter validation remains necessary.

Journal of the American College of Surgeons
University of California, Los Angeles (US), Oldham Council (GB), American College of Surgeons (US)
Decent work and economic growth
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.