LLM-Generated Visual Summaries of Cochrane Plain-Language Summaries: A Pilot Expert-Rated Study and Implications for Research Communication

Background: Plain-language summaries (PLSs) improve accessibility of medical research for patients but remain predominantly text-based. Large language models (LLMs) can now generate images from text. We carried out the present exploratory study to evaluate whether LLMs can generate images that demonstrate technical accuracy and theoretical visual usability when derived from PLS content. Methods: In this cross-sectional pilot study, two PLS were randomly selected from each of 37 Cochrane Library themes. Three LLMs, ChatGPT-5.2, Google Gemini 3 Pro, and Google Notebook, generated one image per PLS, yielding 222 images. Two blinded assessors evaluated images using a preliminary, internally expert-validated tool that measured technical accuracy and completeness, visual usability, and hallucination presence. Inter-LLM comparisons were assessed using linear mixed effects model. Results: Gemini outperformed ChatGPT and Notebook across all domains (p < 0.001). Hallucinations occurred exclusively in ChatGPT-generated images (29.73%). Gemini demonstrated the least intra-thematic variability, whereas ChatGPT showed the highest. No significant interaction was found between LLM type and Cochrane theme. Sensitivity analyses, including alternative weighting schemes and leave-one-theme-out analyses, confirmed robust model rankings. Conclusions: In this single-prompt expert-rated pilot study, Google Gemini 3 Pro reliably generated accurate, hallucination-free visual summaries from PLS. These exploratory findings support further patient-centered validation of LLM-generated images as complements to text-based patient education materials.

Authors

Institutions

Publication Details

Journal
Publications
Published
2026-09-16
DOI
https://doi.org/10.3390/publications14030059
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

LLM-Generated Visual Summaries of Cochrane Plain-Language Summaries: A Pilot Expert-Rated Study and Implications for Research Communication

Kannan Sridharan, Gowri Sivaramakrishnan, Jerome Gnanaraj, Juliet Rebecca Joseph
Publications
Artificial Intelligence in Healthcare and Education
article

LLM-Generated Visual Summaries of Cochrane Plain-Language Summaries: A Pilot Expert-Rated Study and Implications for Research Communication

Kannan Sridharan, Gowri Sivaramakrishnan, Jerome Gnanaraj, Juliet Rebecca Joseph
article en

Abstract

Background: Plain-language summaries (PLSs) improve accessibility of medical research for patients but remain predominantly text-based. Large language models (LLMs) can now generate images from text. We carried out the present exploratory study to evaluate whether LLMs can generate images that demonstrate technical accuracy and theoretical visual usability when derived from PLS content. Methods: In this cross-sectional pilot study, two PLS were randomly selected from each of 37 Cochrane Library themes. Three LLMs, ChatGPT-5.2, Google Gemini 3 Pro, and Google Notebook, generated one image per PLS, yielding 222 images. Two blinded assessors evaluated images using a preliminary, internally expert-validated tool that measured technical accuracy and completeness, visual usability, and hallucination presence. Inter-LLM comparisons were assessed using linear mixed effects model. Results: Gemini outperformed ChatGPT and Notebook across all domains (p < 0.001). Hallucinations occurred exclusively in ChatGPT-generated images (29.73%). Gemini demonstrated the least intra-thematic variability, whereas ChatGPT showed the highest. No significant interaction was found between LLM type and Cochrane theme. Sensitivity analyses, including alternative weighting schemes and leave-one-theme-out analyses, confirmed robust model rankings. Conclusions: In this single-prompt expert-rated pilot study, Google Gemini 3 Pro reliably generated accurate, hallucination-free visual summaries from PLS. These exploratory findings support further patient-centered validation of LLM-generated images as complements to text-based patient education materials.

PublicationsVol. 14(3)
Northeastern University (US), Johns Hopkins University (US), Johns Hopkins Bayview Medical Center (US), Arabian Gulf University (BH)
Quality Education
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.