Large Language Models Meet Gynecologic Ultrasound: Advancing the Characterization of ADNEXal Masses

Ovarian cancer (OC) is the second most common gynecological malignancy and remains one of the leading causes of gynecological cancer-related mortality worldwide. A major clinical challenge is the lack of an accurate and widely applicable strategy for identifying patients at high risk of malignancy at an early stage. In this context, artificial intelligence (AI) has emerged as a promising tool to improve diagnostic performance. Among AI technologies, large language models (LLMs) have recently shown considerable potential in healthcare applications. In this study, we evaluated the diagnostic performance of ChatGPT (GPT-5) in classifying 300 adnexal masses as benign or malignant and compared its performance with that of the IOTA Simple Rules, the ADNEX model, and expert subjective assessment. We also assessed ChatGPT’s ability to predict the most likely histological diagnosis for each lesion. All adnexal masses were described using the International Ovarian Tumor Analysis (IOTA) terminology, and histopathological examination served as the reference standard. Our findings showed that expert subjective assessment achieved the highest overall diagnostic performance for both benign/malignant classification (accuracy 87.3%; 95% CI, 83.0–90.9%) and prediction of the presumed histological diagnosis. ChatGPT A and ChatGPT B reached a sensitivity of 72.3% and 73.5%, a specificity of 74.5% and 75.9%, a positive predictive value of 75.2% and 76.5%, and a negative predictive value of 71.5% and 72.8%, respectively (inconclusive responses counted as misclassifications), with an overall accuracy of 73.3% and 74.7%. After adequate validation, large language models might complement existing decision-support tools for less experienced examiners, without replacing expert evaluation. Their ease of use and reliance on standardized ultrasound descriptors make them accessible to ultrasonographers with varying levels of expertise.

Authors

Institutions

Publication Details

Journal
Journal of Imaging
Published
2026-09-21
DOI
https://doi.org/10.3390/jimaging12090462
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Large Language Models Meet Gynecologic Ultrasound: Advancing the Characterization of ADNEXal Masses

Vera Loizzi, Paolo Trerotoli, Francesca Arezzo, Daniele La Forgia et al.
Journal of Imaging
Artificial Intelligence in Healthcare and Education
article

Large Language Models Meet Gynecologic Ultrasound: Advancing the Characterization of ADNEXal Masses

Vera Loizzi, Paolo Trerotoli, Francesca Arezzo, Daniele La Forgia, Stefania Di Napoli, Gennaro Cormio, Giulia Soccio, Laura Grazia Zompì, Giuseppe Colonna
article en

Abstract

Ovarian cancer (OC) is the second most common gynecological malignancy and remains one of the leading causes of gynecological cancer-related mortality worldwide. A major clinical challenge is the lack of an accurate and widely applicable strategy for identifying patients at high risk of malignancy at an early stage. In this context, artificial intelligence (AI) has emerged as a promising tool to improve diagnostic performance. Among AI technologies, large language models (LLMs) have recently shown considerable potential in healthcare applications. In this study, we evaluated the diagnostic performance of ChatGPT (GPT-5) in classifying 300 adnexal masses as benign or malignant and compared its performance with that of the IOTA Simple Rules, the ADNEX model, and expert subjective assessment. We also assessed ChatGPT’s ability to predict the most likely histological diagnosis for each lesion. All adnexal masses were described using the International Ovarian Tumor Analysis (IOTA) terminology, and histopathological examination served as the reference standard. Our findings showed that expert subjective assessment achieved the highest overall diagnostic performance for both benign/malignant classification (accuracy 87.3%; 95% CI, 83.0–90.9%) and prediction of the presumed histological diagnosis. ChatGPT A and ChatGPT B reached a sensitivity of 72.3% and 73.5%, a specificity of 74.5% and 75.9%, a positive predictive value of 75.2% and 76.5%, and a negative predictive value of 71.5% and 72.8%, respectively (inconclusive responses counted as misclassifications), with an overall accuracy of 73.3% and 74.7%. After adequate validation, large language models might complement existing decision-support tools for less experienced examiners, without replacing expert evaluation. Their ease of use and reliance on standardized ultrasound descriptors make them accessible to ultrasonographers with varying levels of expertise.

Journal of ImagingVol. 12(9)
Istituto Tumori Bari (IT), Istituti di Ricovero e Cura a Carattere Scientifico (IT), University of Bari Aldo Moro (IT)
Good health and well-being
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.