Artificial Intelligence Systems vs. Clinicians in Fetal Echocardiography: A Structured Knowledge‐Based Comparative Assessment
ABSTRACT Purpose Fetal echocardiography requires systematic anatomical assessment, segmental interpretation, and high‐level clinical reasoning. This study aimed to compare the performance of contemporary artificial intelligence (AI) systems and clinicians with different levels of expertise in a standardized fetal echocardiography knowledge‐based assessment. Methods This cross‐sectional comparative assessment study used a 60‐item, single‐best‐answer multiple‐choice examination covering eight fetal echocardiography subdomains: general principles of fetal echocardiography, general fetal cardiac anatomy, left‐sided cardiac pathologies, conotruncal anomalies, isomerism and heterotaxy syndromes, fetal arrhythmias, anomalies of venous return, and pulmonary malformations/right‐sided cardiac pathologies. The same examination was administered to seven publicly accessible AI systems and three clinicians: one senior perinatologist, one junior perinatologist, and one junior obstetrician. All AI systems were accessed during the same predefined evaluation period in January 2026. The primary outcome was response correctness. Secondary analyses included total accuracy, domain‐specific performance, and AI–clinician group comparisons. Results A total of 600 group–question response entries were analyzed, comprising 420 AI‐generated responses and 180 clinician responses. The senior perinatologist achieved the highest total score among all participants, with 53 correct responses out of 60, followed by GPT‐5 with 52 correct responses. Overall accuracy was 78.8% for AI systems and 73.3% for clinicians. The mean domain‐level score was 5.91 for AI systems and 5.50 for clinicians, without a statistically significant difference using either unequal‐variance t ‐test or Mann–Whitney U ‐test. AI systems showed higher mean performance in general fetal echocardiography principles and general fetal cardiac anatomy, whereas the senior perinatologist performed best in complex integrative domains, particularly conotruncal anomalies and isomerism/heterotaxy syndromes. Conclusions In a structured text‐based fetal echocardiography knowledge assessment, AI systems achieved overall performance comparable to clinicians, particularly in standardized knowledge domains. However, the senior perinatologist retained the strongest performance in complex fetal cardiac categories requiring integrative clinical reasoning. These findings support the potential role of AI as an adjunctive educational and structured decision‐support tool in fetal echocardiography, while emphasizing that performance in a multiple‐choice format should not be interpreted as equivalent to autonomous real‐world fetal cardiac diagnosis.
Authors
- Eda Özden Tokalıoğlu (ORCID: https://orcid.org/0000-0003-4901-0544)
- Şükrü Bakırcı (ORCID: https://orcid.org/0000-0002-6007-965X)
- Fatma Doğa Öcal (ORCID: https://orcid.org/0000-0003-4727-7982)
- Dilek Şahın (ORCID: https://orcid.org/0000-0001-8567-9048)
- Özgür Kara (ORCID: https://orcid.org/0000-0002-4204-0014)
- Uğurcan Zorlu (ORCID: https://orcid.org/0000-0002-8912-0812)
- Mehmet Utku Başarır (ORCID: https://orcid.org/0009-0000-5741-3875)
- H. İbrahim Altınsoy (ORCID: https://orcid.org/0000-0001-9616-8139)
Institutions
- Sağlık Bilimleri Üniversitesi (TR)
- Ankara Bilkent City Hospital (TR)
Publication Details
- Journal
- Journal of Clinical Ultrasound
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1002/jcu.70430
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00