Comparison and Combination of Radiomics, Deep Learning, and Subjective Assessment for Differentiating Immature and Mature Ovarian Teratomas on Abdominal CT

Objective: Preoperatively, ovarian teratomas are often misclassified into immature and mature types, potentially resulting in adverse clinical outcomes. This study compared the diagnostic accuracy of radiomics-based machine learning (ML), deep learning (DL), and subjective radiologic assessment for differentiating immature from mature ovarian teratomas on abdominal computed tomography (CT) and evaluated whether combining these approaches improves discrimination. Methods: We retrospectively reviewed imaging data of 31 cases with pathologically confirmed immature teratomas and 61 age- and size-matched cases of mature teratomas (mean age=18.1±9.2 y vs. 20.9±10.0 y). Two radiologists independently annotated predefined CT features to construct a multivariable logistic regression model. Radiomics features were extracted from precontrast CT with 3 masks: whole tumor, calcification, and fat. Radiomics models were developed using 4 ML pipelines (logistic regression, random forest, support vector machine, and neural network). DL models were constructed with a 3D-ResNet18 on full-volume and tumor-segmented CT, with and without transfer learning (TL). A combined model was constructed by averaging their calibrated predicted probabilities. The diagnostic performance was assessed by the mean area under the receiver-operating characteristic curve (AUC). Results: In the full cohort, the 4 individual approaches performed similarly, with overlapping CIs: wavelet-transformed precontrast radiomics 0.820 (95% CI: 0.798-0.841), precontrast radiomics 0.817, subjective assessment 0.814, and DL 0.810. Combining them outperformed every individual approach (0.865; ΔAUC: 0.045, 95% CI: 0.032-0.060; Holm-adjusted P <0.001). Among the 81 tumors containing calcification, where calcification-mask radiomics could also be evaluated, all approaches performed better, and the combined model again outperformed each of them (0.891; ΔAUC: 0.035), as did the combined model among the 90 tumors containing fat (0.872; ΔAUC: 0.047; both P <0.001). Conclusion: Radiomics shows moderate performance in CT-based discrimination between immature and mature ovarian teratomas, with subjective radiologic assessment and DL performing at a comparable level. Combining the 3 consistently outperformed anyone alone, indicating that quantitative image analysis complements rather than replaces expert visual interpretation and may help reduce preoperative misclassification.

Authors

Publication Details

Journal
Journal of Computer Assisted Tomography
Published
2026-10-06
DOI
https://doi.org/10.1097/rct.0000000000001927
Primary Topic
Radiomics and Machine Learning in Medical Imaging
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Comparison and Combination of Radiomics, Deep Learning, and Subjective Assessment for Differentiating Immature and Mature Ovarian Teratomas on Abdominal CT

Taek Min Kim, Sang Youn Kim, Jeong Yeon Cho, June Young Seo et al.
Journal of Computer Assisted Tomography
Radiomics and Machine Learning in Medical Imaging
article

Comparison and Combination of Radiomics, Deep Learning, and Subjective Assessment for Differentiating Immature and Mature Ovarian Teratomas on Abdominal CT

Taek Min Kim, Sang Youn Kim, Jeong Yeon Cho, June Young Seo, Jongwoo Park
article en

Abstract

Objective: Preoperatively, ovarian teratomas are often misclassified into immature and mature types, potentially resulting in adverse clinical outcomes. This study compared the diagnostic accuracy of radiomics-based machine learning (ML), deep learning (DL), and subjective radiologic assessment for differentiating immature from mature ovarian teratomas on abdominal computed tomography (CT) and evaluated whether combining these approaches improves discrimination. Methods: We retrospectively reviewed imaging data of 31 cases with pathologically confirmed immature teratomas and 61 age- and size-matched cases of mature teratomas (mean age=18.1±9.2 y vs. 20.9±10.0 y). Two radiologists independently annotated predefined CT features to construct a multivariable logistic regression model. Radiomics features were extracted from precontrast CT with 3 masks: whole tumor, calcification, and fat. Radiomics models were developed using 4 ML pipelines (logistic regression, random forest, support vector machine, and neural network). DL models were constructed with a 3D-ResNet18 on full-volume and tumor-segmented CT, with and without transfer learning (TL). A combined model was constructed by averaging their calibrated predicted probabilities. The diagnostic performance was assessed by the mean area under the receiver-operating characteristic curve (AUC). Results: In the full cohort, the 4 individual approaches performed similarly, with overlapping CIs: wavelet-transformed precontrast radiomics 0.820 (95% CI: 0.798-0.841), precontrast radiomics 0.817, subjective assessment 0.814, and DL 0.810. Combining them outperformed every individual approach (0.865; ΔAUC: 0.045, 95% CI: 0.032-0.060; Holm-adjusted P <0.001). Among the 81 tumors containing calcification, where calcification-mask radiomics could also be evaluated, all approaches performed better, and the combined model again outperformed each of them (0.891; ΔAUC: 0.035), as did the combined model among the 90 tumors containing fat (0.872; ΔAUC: 0.047; both P <0.001). Conclusion: Radiomics shows moderate performance in CT-based discrimination between immature and mature ovarian teratomas, with subjective radiologic assessment and DL performing at a comparable level. Combining the 3 consistently outperformed anyone alone, indicating that quantitative image analysis complements rather than replaces expert visual interpretation and may help reduce preoperative misclassification.

Journal of Computer Assisted Tomography
Openalex Percentile: Top 12%
Radiomics and Machine Learning in Medical Imaging
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.