Clinical evaluation of novel deep learning‐based auto‐segmentation software: Utility and potential pitfalls

BACKGROUND: Accurate contouring of target volumes and organs at risk is critical in radiotherapy. While deep learning (DL) models offer automated contouring, their clinical applicability to real-world cases containing anatomical variations and artifacts requires rigorous validation. PURPOSE: To evaluate the clinical accuracy and potential vulnerabilities of RatoGuide, novel DL-based auto-segmentation software, using a dataset including atypical cases derived from routine clinical practice. METHODS: This single-center retrospective study included 69 thoracic and male pelvic cases. The cohort was intentionally selected to encompass diverse anatomies and artifacts (e.g., pacemakers, SpaceOAR implants, artificial femoral head replacements, and unilateral atelectasis). Auto-contours generated by RatoGuide were compared with expert-approved manual contours. Performance was evaluated quantitatively using the Dice Similarity Coefficient (DSC) and 95th percentile Hausdorff Distance (HD95), and qualitatively via a 5-point visual assessment scale by four independent reviewers. Statistical comparisons between cohorts were performed using the Mann-Whitney U test. Additionally, a dosimetric evaluation was conducted for male pelvic cases to assess clinical impact. RESULTS: In typical cases, the software maintained high segmentation accuracy (thorax: mean DSC 0.856, mean HD95 6.89 mm; male pelvis: mean DSC 0.874, mean HD95 4.11 mm). However, performance declined in atypical cohorts (thorax: mean DSC 0.808, p = 0.0457, mean HD95 12.22 mm, p = 0.0002; male pelvis: mean DSC 0.828, p = 0.1620, mean HD95 6.16 mm, p = 0.0075). Notable decreases in accuracy were observed in challenging scenarios, such as artificial femoral head replacements (DSC: 0.754) and unilateral atelectasis (DSC: 0.784). Qualitative assessment revealed that errors were primarily due to anatomical factors and artifacts. Furthermore, the dosimetric evaluation identified one critical false-negative error where a dose constraint violation was overlooked when the DL contour was used. CONCLUSIONS: RatoGuide demonstrated favorable performance in typical cases, but accuracy declined in atypical cases with artifacts or altered anatomy. For clinical implementation, rigorous visual verification and manual review by experts are essential, particularly for atypical cases and organs in high-dose gradient regions.

Authors

Institutions

Publication Details

Journal
Journal of Applied Clinical Medical Physics
Published
2026-08-25
DOI
https://doi.org/10.1002/acm2.70757
Citations
1
Primary Topic
Advanced Radiotherapy Techniques
Type
article
Field-Weighted Citation Impact
6.42
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Clinical evaluation of novel deep learning‐based auto‐segmentation software: Utility and potential pitfalls

Zhe Chen, Masahide Saito, Noriyuki Kadoya, Masaki Matsuda et al.
1 citations
Journal of Applied Clinical Medical Physics
Advanced Radiotherapy Techniques
6.42
article

Clinical evaluation of novel deep learning‐based auto‐segmentation software: Utility and potential pitfalls

Zhe Chen, Masahide Saito, Noriyuki Kadoya, Masaki Matsuda, Ryota Tozuka, T. Akita, Hikaru Nemoto, Takafumi Komiyama, Keiichi Jingu, Hiroshi Onishi, Kazuma Mochizuki
article en
1 citations

Abstract

BACKGROUND: Accurate contouring of target volumes and organs at risk is critical in radiotherapy. While deep learning (DL) models offer automated contouring, their clinical applicability to real-world cases containing anatomical variations and artifacts requires rigorous validation. PURPOSE: To evaluate the clinical accuracy and potential vulnerabilities of RatoGuide, novel DL-based auto-segmentation software, using a dataset including atypical cases derived from routine clinical practice. METHODS: This single-center retrospective study included 69 thoracic and male pelvic cases. The cohort was intentionally selected to encompass diverse anatomies and artifacts (e.g., pacemakers, SpaceOAR implants, artificial femoral head replacements, and unilateral atelectasis). Auto-contours generated by RatoGuide were compared with expert-approved manual contours. Performance was evaluated quantitatively using the Dice Similarity Coefficient (DSC) and 95th percentile Hausdorff Distance (HD95), and qualitatively via a 5-point visual assessment scale by four independent reviewers. Statistical comparisons between cohorts were performed using the Mann-Whitney U test. Additionally, a dosimetric evaluation was conducted for male pelvic cases to assess clinical impact. RESULTS: In typical cases, the software maintained high segmentation accuracy (thorax: mean DSC 0.856, mean HD95 6.89 mm; male pelvis: mean DSC 0.874, mean HD95 4.11 mm). However, performance declined in atypical cohorts (thorax: mean DSC 0.808, p = 0.0457, mean HD95 12.22 mm, p = 0.0002; male pelvis: mean DSC 0.828, p = 0.1620, mean HD95 6.16 mm, p = 0.0075). Notable decreases in accuracy were observed in challenging scenarios, such as artificial femoral head replacements (DSC: 0.754) and unilateral atelectasis (DSC: 0.784). Qualitative assessment revealed that errors were primarily due to anatomical factors and artifacts. Furthermore, the dosimetric evaluation identified one critical false-negative error where a dose constraint violation was overlooked when the DL contour was used. CONCLUSIONS: RatoGuide demonstrated favorable performance in typical cases, but accuracy declined in atypical cases with artifacts or altered anatomy. For clinical implementation, rigorous visual verification and manual review by experts are essential, particularly for atypical cases and organs in high-dose gradient regions.

Journal of Applied Clinical Medical PhysicsVol. 27(9)
Tohoku University (JP), University of Yamanashi Hospital (JP), University of Yamanashi (JP)
Openalex Percentile: Top 10%
Advanced Radiotherapy Techniques
6.42
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.