Diagnostic Performance of Different AI Tools for Chest Radiography: A CT-Anchored Per-Abnormality Analysis in a Tertiary Referral Center

Background/Objectives: Artificial intelligence (AI) increasingly supports chest radiograph interpretation, but per-abnormality comparisons of commercial tools using CT-anchored reference standards remain limited. We compared two commercial AI tools using a CT-anchored, CXR-targeted reference framework. Methods: We retrospectively reviewed chest radiographs obtained between 30 October 2023 and 10 October 2024. Each radiograph was paired with CT performed within 14 days or within 2 days for rapidly progressive findings. Seven abnormalities across five analysis categories were evaluated: fracture, nodule, pleural effusion, pneumothorax, and airspace disease (edema, atelectasis, opacity/consolidation). Radiologists first confirmed abnormalities on CT then assessed their visibility on the paired radiograph. Outputs from vendors A and B were recorded. Sensitivity was calculated for CT-confirmed, radiographically visible abnormalities; specificity, precision, and accuracy were calculated within the same detectability framework. Vendors were compared using McNemar’s test. Results: The cohort included 1707 radiographs from 1680 patients (mean age, 60.3 ± 16.5 years; 50.5% male). Of 3573 CT-confirmed abnormalities, 1981 (55.4%) were radiographically visible and formed the primary reference-positive set. Vendor A had higher sensitivity for fractures, nodules, and pleural effusions (all p < 0.001), whereas vendor B had higher specificity for these findings (all p < 0.05). For pneumothorax, vendor B showed numerically higher sensitivity but lower specificity and accuracy; case numbers were limited, and precision was low for both tools. No significant differences were observed for harmonized airspace disease. Overall specificity was high, while sensitivity and precision varied across abnormalities. Conclusions: Both tools showed abnormality-specific strengths and limitations, supporting tailored implementation and real-world clinical validation.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-09-27
DOI
https://doi.org/10.3390/diagnostics16193139
Primary Topic
COVID-19 diagnosis using AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Diagnostic Performance of Different AI Tools for Chest Radiography: A CT-Anchored Per-Abnormality Analysis in a Tertiary Referral Center

Selin Ardalı Düzgün, Mustafa Ege Şeker, Deniz Köksal, Gamze Durhan et al.
Diagnostics
COVID-19 diagnosis using AI
article

Diagnostic Performance of Different AI Tools for Chest Radiography: A CT-Anchored Per-Abnormality Analysis in a Tertiary Referral Center

Selin Ardalı Düzgün, Mustafa Ege Şeker, Deniz Köksal, Gamze Durhan, Nuri Saraç, Meltem Gülsün Akpınar, Sevinç Sarınç, İlke Taşçı, Melih Karadag, Atahan Tamturk, Figen Başaran Demirkazık
article en

Abstract

Background/Objectives: Artificial intelligence (AI) increasingly supports chest radiograph interpretation, but per-abnormality comparisons of commercial tools using CT-anchored reference standards remain limited. We compared two commercial AI tools using a CT-anchored, CXR-targeted reference framework. Methods: We retrospectively reviewed chest radiographs obtained between 30 October 2023 and 10 October 2024. Each radiograph was paired with CT performed within 14 days or within 2 days for rapidly progressive findings. Seven abnormalities across five analysis categories were evaluated: fracture, nodule, pleural effusion, pneumothorax, and airspace disease (edema, atelectasis, opacity/consolidation). Radiologists first confirmed abnormalities on CT then assessed their visibility on the paired radiograph. Outputs from vendors A and B were recorded. Sensitivity was calculated for CT-confirmed, radiographically visible abnormalities; specificity, precision, and accuracy were calculated within the same detectability framework. Vendors were compared using McNemar’s test. Results: The cohort included 1707 radiographs from 1680 patients (mean age, 60.3 ± 16.5 years; 50.5% male). Of 3573 CT-confirmed abnormalities, 1981 (55.4%) were radiographically visible and formed the primary reference-positive set. Vendor A had higher sensitivity for fractures, nodules, and pleural effusions (all p < 0.001), whereas vendor B had higher specificity for these findings (all p < 0.05). For pneumothorax, vendor B showed numerically higher sensitivity but lower specificity and accuracy; case numbers were limited, and precision was low for both tools. No significant differences were observed for harmonized airspace disease. Overall specificity was high, while sensitivity and precision varied across abnormalities. Conclusions: Both tools showed abnormality-specific strengths and limitations, supporting tailored implementation and real-world clinical validation.

DiagnosticsVol. 16(19)
University of Wisconsin–Madison (US), Hacettepe University (TR)
Openalex Percentile: Top 12%
COVID-19 diagnosis using AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.