Development and validation of a pragmatic pipeline for clinical free-text annotation using locally deployed open-weight large language models

Abstract Clinical information required for surgical data science (SDS) is frequently embedded in unstructured text. We developed and evaluated a reproducible pipeline for selecting locally deployed open-weight large language models (LLMs) for binary symptom annotation. In this retrospective single-center study, 1,100 German emergency-department reports were manually annotated for nausea, vomiting, diarrhea, and dysuria. After reserving 100 reports for prompt formulation and temperature testing, nine LLMs were screened on symptom-specific stratified development sets ( N = 250). Selected models were compared with a negation-aware rule-based baseline in independent validation sets ( N = 750) using F 1 -score and patient-level bootstrap confidence intervals. Temperature 0.0 provided the greatest overall stability. Validation F 1 -scores were 0.985 for vomiting, 0.979 for nausea, 0.824 for dysuria, and 0.814 for diarrhea. Corresponding baseline F 1 -scores were 0.913, 0.724, 0.705, and 0.853, respectively. Paired comparisons favored LLMs for nausea and vomiting; confidence intervals included zero for diarrhea and dysuria. Median inference times ranged from 0.298 to 1.653 s per report. Discrepancies reflected operational criteria, temporal variation, inconsistent documentation, missed mentions, and five reference errors. Pragmatic model screening can identify suitable local LLMs for clinical free-text annotation. The pipeline is reproducible and adaptable but requires context-specific configuration and validation.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-29
DOI
https://doi.org/10.1038/s41598-026-73738-7
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Development and validation of a pragmatic pipeline for clinical free-text annotation using locally deployed open-weight large language models

Philipp Feodorovici, Jan Arensmeyer, Hanno Matthaei, Ingo Gräff et al.
Scientific Reports
Topic Modeling
article

Development and validation of a pragmatic pipeline for clinical free-text annotation using locally deployed open-weight large language models

Philipp Feodorovici, Jan Arensmeyer, Hanno Matthaei, Ingo Gräff, Benjamin Wulff, Jonas Henn, Jörg C. Kalff, Alisa Stoll, Johannes Röttgen, D Subramani
article en

Abstract

Abstract Clinical information required for surgical data science (SDS) is frequently embedded in unstructured text. We developed and evaluated a reproducible pipeline for selecting locally deployed open-weight large language models (LLMs) for binary symptom annotation. In this retrospective single-center study, 1,100 German emergency-department reports were manually annotated for nausea, vomiting, diarrhea, and dysuria. After reserving 100 reports for prompt formulation and temperature testing, nine LLMs were screened on symptom-specific stratified development sets ( N = 250). Selected models were compared with a negation-aware rule-based baseline in independent validation sets ( N = 750) using F 1 -score and patient-level bootstrap confidence intervals. Temperature 0.0 provided the greatest overall stability. Validation F 1 -scores were 0.985 for vomiting, 0.979 for nausea, 0.824 for dysuria, and 0.814 for diarrhea. Corresponding baseline F 1 -scores were 0.913, 0.724, 0.705, and 0.853, respectively. Paired comparisons favored LLMs for nausea and vomiting; confidence intervals included zero for diarrhea and dysuria. Median inference times ranged from 0.298 to 1.653 s per report. Discrepancies reflected operational criteria, temporal variation, inconsistent documentation, missed mentions, and five reference errors. Pragmatic model screening can identify suitable local LLMs for clinical free-text annotation. The pipeline is reproducible and adaptable but requires context-specific configuration and validation.

Scientific ReportsVol. 16(1)
University Hospital Bonn (DE)
Quality Education
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.