Evaluating large language models for lay summaries of radiology reports using tailored prompting strategies and mixed-method assessment

{"Radiology":[0],"reports":[1,53],"are":[2,18],"often":[3],"filled":[4],"with":[5,287],"medical":[6],"jargon":[7],"that":[8,183],"limits":[9],"patient":[10],"understanding.":[11],"Lay":[12],"summaries":[13,47,64,98,182,252],"can":[14],"improve":[15],"understanding":[16],"but":[17],"time-consuming":[19],"for":[20,38,178],"healthcare":[21],"providers":[22],"to":[23,31],"create.":[24],"The":[25,97],"objective":[26],"of":[27,35,144,294],"this":[28,283],"study":[29,284],"is":[30],"explore":[32],"the":[33,55,147,218,249,292],"use":[34],"tailored":[36],"prompts":[37,290],"five":[39],"Large":[40,118],"Language":[41],"Models":[42,120],"(LLMs)":[43],"in":[44,241,297],"generating":[45,179],"lay":[46,63,181,300],"from":[48,54],"radiology":[49,115],"reports.":[50],"Using":[51,141],"100":[52],"publicly":[56],"available":[57],"\\"BioNLP":[58],"2023":[59],"report":[60],"summarization\\"":[61],"dataset,":[62],"were":[65,99,151,159,170],"generated":[66,288],"by":[67,93,109,173],"each":[68],"LLM,":[69],"under":[70],"select":[71],"prompting":[72],"styles":[73],"[Few-Shot":[74],"(GPT-4),":[75],"Generated":[76],"Knowledge":[77],"(GPT-4o":[78],"mini,":[79],"Gemini":[80,84,161,224,232,244],"1.5":[81,85,162,225,233,245],"-":[82,86,124,163,166,226,234,246],"Pro,":[83],"Flash),":[87],"and":[88,117,131,137,153,155,165,176,231,271,276,291],"Zero-Shot":[89],"(Llama":[90],"3.1)]":[91],"informed":[92],"a":[94,102],"pilot":[95],"work.":[96],"evaluated":[100],"using":[101],"mixed-method":[103],"framework:":[104],"subjective":[105],"assessment":[106],"(Likert":[107],"statements)":[108],"blinded":[110],"experts":[111,175,270],"(n":[112],"=":[113],"2":[114],"fellows)":[116],"Reasoning":[119],"(LRMs)":[121],"[(Gemini":[122],"2.5":[123],"Pro":[125,167,235,247],"(LRM":[126,129,206,213,228,236,274,278],"1);":[127,193],"GPT-oss-120b":[128],"2)],":[130],"readability":[132],"metrics":[133],"(Flesch-Kincaid":[134,253],"Grade":[135,254],"Level":[136],"Flesch":[138,259],"Reading":[139,260],"Ease).":[140],"percentage":[142],"agreement":[143,266],"Likert":[145],"statements,":[146],"LLM-prompt":[148],"combinations'":[149],"performances":[150],"ranked,":[152],"Friedman":[154],"post-hoc":[156],"Nemenyi":[157],"tests":[158],"conducted.":[160],"Flash":[164,227],"(generated":[168],"knowledge)":[169],"rated":[171],"highest":[172,219],"human":[174],"LRMs":[177,272],"actionable":[180],"require":[184],"minimal":[185],"supervision":[186],"[P":[187],"<":[188,195,202,209],"4.97":[189],"×":[190,197,204,211],"10-2":[191],"(Rater":[192,199],"P":[194,201,208],"9.03":[196],"10-21":[198],"2),":[200],"6.90":[203],"10-15":[205],"1),":[207],"2.760":[210],"10-5":[212],"2).":[214],"GPT-4":[215],"(few-shot)":[216],"achieved":[217],"human-rated":[220],"accuracy":[221],"(98%),":[222],"while":[223],"1-rated:":[229],"95%)":[230],"2-rated:":[237],"91%)":[238],"ranked":[239],"first":[240],"LRM-rated":[242],"accuracy.":[243],"produced":[248],"most":[250],"accessible":[251],"Level:":[255],"7.55":[256],"±":[257,263],"1.38,":[258],"Ease:":[261],"67.84":[262],"7.78).":[264],"Strong":[265],"was":[267],"observed":[268],"between":[269],"[0.96%":[273],"1)":[275],"3.4%":[277],"2)":[279],"complete":[280],"disagreement].":[281],"Overall,":[282],"highlights":[285],"Gemini-models":[286],"knowledge":[289],"potential":[293],"LRM":[295],"evaluators":[296],"assessing":[298],"LLM-generated":[299],"summaries.":[301]}

Authors

Institutions

Publication Details

Journal
PLOS Digital Health
Published
2026-09-17
DOI
https://doi.org/10.1371/journal.pdig.0001672
Primary Topic
Radiology practices and education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating large language models for lay summaries of radiology reports using tailored prompting strategies and mixed-method assessment

Cynthia Lokker, Nitin Juggath, Ashirbani Saha, Christian B. van der Pol et al.
PLOS Digital Health
Radiology practices and education
article

Evaluating large language models for lay summaries of radiology reports using tailored prompting strategies and mixed-method assessment

Cynthia Lokker, Nitin Juggath, Ashirbani Saha, Christian B. van der Pol, Ambreen Zahoor, Nanziba Tasneem, Kyle McGowan
article en

Abstract

Radiology reports are often filled with medical jargon that limits patient understanding. Lay summaries can improve understanding but are time-consuming for healthcare providers to create. The objective of this study is to explore the use of tailored prompts for five Large Language Models (LLMs) in generating lay summaries from radiology reports. Using 100 reports from the publicly available "BioNLP 2023 report summarization" dataset, lay summaries were generated by each LLM, under select prompting styles [Few-Shot (GPT-4), Generated Knowledge (GPT-4o mini, Gemini 1.5 - Pro, Gemini 1.5 - Flash), and Zero-Shot (Llama 3.1)] informed by a pilot work. The summaries were evaluated using a mixed-method framework: subjective assessment (Likert statements) by blinded experts (n = 2 radiology fellows) and Large Reasoning Models (LRMs) [(Gemini 2.5 - Pro (LRM 1); GPT-oss-120b (LRM 2)], and readability metrics (Flesch-Kincaid Grade Level and Flesch Reading Ease). Using percentage agreement of Likert statements, the LLM-prompt combinations' performances were ranked, and Friedman and post-hoc Nemenyi tests were conducted. Gemini 1.5 - Flash and - Pro (generated knowledge) were rated highest by human experts and LRMs for generating actionable lay summaries that require minimal supervision [P < 4.97 × 10-2 (Rater 1); P < 9.03 × 10-21 (Rater 2), P < 6.90 × 10-15 (LRM 1), P < 2.760 × 10-5 (LRM 2). GPT-4 (few-shot) achieved the highest human-rated accuracy (98%), while Gemini 1.5 - Flash (LRM 1-rated: 95%) and Gemini 1.5 - Pro (LRM 2-rated: 91%) ranked first in LRM-rated accuracy. Gemini 1.5 - Pro produced the most accessible summaries (Flesch-Kincaid Grade Level: 7.55 ± 1.38, Flesch Reading Ease: 67.84 ± 7.78). Strong agreement was observed between experts and LRMs [0.96% (LRM 1) and 3.4% (LRM 2) complete disagreement]. Overall, this study highlights Gemini-models with generated knowledge prompts and the potential of LRM evaluators in assessing LLM-generated lay summaries.

PLOS Digital HealthVol. 5(9)
Juravinski Cancer Centre (CA), Juravinski Hospital (CA), Population Health Research Institute (CA), Hamilton Health Sciences (CA), McMaster University (CA)
Quality Education
Openalex Percentile: Top 12%
Radiology practices and education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.