How green are large language models for radiology report labelling? Comparing human, rule-based and hybrid workflows

Abstract Objectives To address limited quantitative data on sustainable use of large language models (LLMs) in radiology, we quantified the resource footprint of LLMs for labelling CT pulmonary embolism reports and assessed how a hybrid rule-based–LLM workflow changes time, cost and carbon emissions compared with manual labelling. Materials and methods In this single-centre retrospective study, 2923 structured CT reports were labelled using four workflows: a rule-based extractor (RBE), an LLM-only pipeline using 18 open-weight and four proprietary models, a hybrid RBE–LLM pipeline that routed RBE failures to an LLM, and full manual labelling by radiologists. Ground truth was based on radiologist adjudication. For each LLM, we measured per-report latency, estimated CO 2 emissions and cost. Radiologists recorded the labelling time per report. Results Manual labelling required 32.8 h for 2923 reports (40.4 s/report; €0.42/report) with 95.0% accuracy (95% CI: 93.7–96.2). LLM-only pipelines were less accurate (85.1%; 95% CI: 84.9–85.5) but reduced labelling time to 12.4 h and cost to €2.60 (both p < 0.001). Hybrid RBE–LLM workflows yielded the highest accuracy (98.5%) and lowest resource use: across 22 models, switching from LLM-only to hybrid reduced time (6.7 to 0.97 h), cost (€1.19 to €0.17), and CO 2 (0.82 to 0.12 kg; all p < 0.001). Conclusion LLM-only labelling reduced labour time and direct costs compared with manual annotation. A hybrid RBE–LLM pipeline that forwards rule-based failures to an LLM concentrated compute where needed and markedly decreased time, cost and emissions, supporting targeted deployment of LLMs for sustainable data-annotation workflows in radiology. Critical relevance By quantifying time, cost and carbon emissions of manual, rule-based, LLM and hybrid report labelling, this study identifies sustainable workflows for deploying LLMs in routine radiology reporting. Key Points Manual expert labelling of CT pulmonary embolism reports is time-intensive and costly. Mid-sized LLM configurations provide favourable trade-offs between performance and resource use. Hybrid rule-based–LLM workflows sustain accuracy while reducing resource demands.

Authors

Institutions

Publication Details

Journal
Insights into Imaging
Published
2026-05-27
DOI
https://doi.org/10.1186/s13244-026-02289-2
Citations
1
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
8.20
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

How green are large language models for radiology report labelling? Comparing human, rule-based and hybrid workflows

Jonas Kroschke, Edem Atsiatorme, Martin Moll, Veronika Riebl et al.
1 citations
Insights into Imaging
Artificial Intelligence in Healthcare and Education
8.20
article

How green are large language models for radiology report labelling? Comparing human, rule-based and hybrid workflows

Jonas Kroschke, Edem Atsiatorme, Martin Moll, Veronika Riebl, Alexander Krémer, Arved Bischoff, Patrick Stein, K Schlamp, Timo Leichenich, M Fink, Hans-Ulrich Kauczor
article en
1 citations

Abstract

Abstract Objectives To address limited quantitative data on sustainable use of large language models (LLMs) in radiology, we quantified the resource footprint of LLMs for labelling CT pulmonary embolism reports and assessed how a hybrid rule-based–LLM workflow changes time, cost and carbon emissions compared with manual labelling. Materials and methods In this single-centre retrospective study, 2923 structured CT reports were labelled using four workflows: a rule-based extractor (RBE), an LLM-only pipeline using 18 open-weight and four proprietary models, a hybrid RBE–LLM pipeline that routed RBE failures to an LLM, and full manual labelling by radiologists. Ground truth was based on radiologist adjudication. For each LLM, we measured per-report latency, estimated CO 2 emissions and cost. Radiologists recorded the labelling time per report. Results Manual labelling required 32.8 h for 2923 reports (40.4 s/report; €0.42/report) with 95.0% accuracy (95% CI: 93.7–96.2). LLM-only pipelines were less accurate (85.1%; 95% CI: 84.9–85.5) but reduced labelling time to 12.4 h and cost to €2.60 (both p < 0.001). Hybrid RBE–LLM workflows yielded the highest accuracy (98.5%) and lowest resource use: across 22 models, switching from LLM-only to hybrid reduced time (6.7 to 0.97 h), cost (€1.19 to €0.17), and CO 2 (0.82 to 0.12 kg; all p < 0.001). Conclusion LLM-only labelling reduced labour time and direct costs compared with manual annotation. A hybrid RBE–LLM pipeline that forwards rule-based failures to an LLM concentrated compute where needed and markedly decreased time, cost and emissions, supporting targeted deployment of LLMs for sustainable data-annotation workflows in radiology. Critical relevance By quantifying time, cost and carbon emissions of manual, rule-based, LLM and hybrid report labelling, this study identifies sustainable workflows for deploying LLMs in routine radiology reporting. Key Points Manual expert labelling of CT pulmonary embolism reports is time-intensive and costly. Mid-sized LLM configurations provide favourable trade-offs between performance and resource use. Hybrid rule-based–LLM workflows sustain accuracy while reducing resource demands.

Insights into ImagingVol. 17(1)
Heidelberg University (DE), University Hospital Heidelberg (DE), German Center for Lung Research (DE)
Openalex Percentile: Top 2%
Artificial Intelligence in Healthcare and Education
8.20
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.