ChatGPT-4-Based Automated Preliminary Clinical Reporting in Myocardial Perfusion Imaging: A Pilot Evaluation

Background/Objectives: Myocardial perfusion scintigraphy (MPS) is a cornerstone non-invasive imaging modality for assessing myocardial ischemia and infarction. This pilot study aimed to evaluate the feasibility of using ChatGPT-4 for automated preliminary clinical reporting in MPS and to identify specific scenarios in which large language model (LLM) performance is robust versus inadequate.Methods: A comparative analysis was conducted using 30 consecutive de-identified MPS cases spanning a broad spectrum of clinical scenarios. Structured clinical data were input into ChatGPT-4 to generate AI-based preliminary reports, which were then compared with reports prepared by two experienced nuclear medicine physicians. Reports were independently evaluated using four criteria—clinical accuracy, report structure, terminological appropriateness, and overall comprehensibility—scored on a 5-point Likert scale. Inter-observer agreement was assessed using Cohen’s kappa.Results: ChatGPT-4 demonstrated strong performance in report structure (median 5), terminological appropriateness (median 4), and overall comprehensibility (median 5), consistently producing well-organized and coherent reports. However, clinical accuracy was significantly lower compared with physician reports (median 4 vs. 5; p = 0.002; effect size r = 0.52), particularly in complex cases such as multivessel ischemia and mixed pathologies, where outputs were occasionally superficial or lacked specificity. Inter-observer agreement between evaluating physicians was substantial (Cohen’s κ = 0.78; 95% CI: 0.64–0.92).Conclusions: ChatGPT-4 shows promise as a supportive tool for preliminary MPS reporting and medical education, but its limitations in higher-order clinical reasoning necessitate careful human oversight. These findings are preliminary and require validation in adequately powered, multicenter studies before any clinical implementation.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-09-24
DOI
https://doi.org/10.3390/diagnostics16193110
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

ChatGPT-4-Based Automated Preliminary Clinical Reporting in Myocardial Perfusion Imaging: A Pilot Evaluation

Mutlay KESKİN, Ece Oğuz
Diagnostics
Artificial Intelligence in Healthcare and Education
article

ChatGPT-4-Based Automated Preliminary Clinical Reporting in Myocardial Perfusion Imaging: A Pilot Evaluation

Mutlay KESKİN, Ece Oğuz
article en

Abstract

Background/Objectives: Myocardial perfusion scintigraphy (MPS) is a cornerstone non-invasive imaging modality for assessing myocardial ischemia and infarction. This pilot study aimed to evaluate the feasibility of using ChatGPT-4 for automated preliminary clinical reporting in MPS and to identify specific scenarios in which large language model (LLM) performance is robust versus inadequate.Methods: A comparative analysis was conducted using 30 consecutive de-identified MPS cases spanning a broad spectrum of clinical scenarios. Structured clinical data were input into ChatGPT-4 to generate AI-based preliminary reports, which were then compared with reports prepared by two experienced nuclear medicine physicians. Reports were independently evaluated using four criteria—clinical accuracy, report structure, terminological appropriateness, and overall comprehensibility—scored on a 5-point Likert scale. Inter-observer agreement was assessed using Cohen’s kappa.Results: ChatGPT-4 demonstrated strong performance in report structure (median 5), terminological appropriateness (median 4), and overall comprehensibility (median 5), consistently producing well-organized and coherent reports. However, clinical accuracy was significantly lower compared with physician reports (median 4 vs. 5; p = 0.002; effect size r = 0.52), particularly in complex cases such as multivessel ischemia and mixed pathologies, where outputs were occasionally superficial or lacked specificity. Inter-observer agreement between evaluating physicians was substantial (Cohen’s κ = 0.78; 95% CI: 0.64–0.92).Conclusions: ChatGPT-4 shows promise as a supportive tool for preliminary MPS reporting and medical education, but its limitations in higher-order clinical reasoning necessitate careful human oversight. These findings are preliminary and require validation in adequately powered, multicenter studies before any clinical implementation.

DiagnosticsVol. 16(19)
Mersin Şehir Eğitim ve Araştırma Hastanesi (TR), Mersin Üniversitesi (TR)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.