An ultrasound-based end-to-end multi-task deep learning model to segment and classify parotid gland tumors: A retrospective multi-center study
Abstract Background This study aimed to develop and validate an end-to-end multi-task deep learning model for simultaneous segmentation and four-class classification of parotid gland tumors (PGTs) based on ultrasound (US) images, with the goal of improving preoperative diagnostic accuracy to further assist with timely clinical decision-making. Methods In this multicenter retrospective study, a total of 1,666 patients with 1,744 PGTs from five medical centers were analyzed. The dataset was divided into an internal dataset ( n = 1,029) and two external test datasets ( n = 229 and n = 486). The proposed GobletNet adopts a multi-task learning framework with a shared encoder to perform tumor segmentation and four-class classification, including malignant tumors, pleomorphic adenoma, Warthin tumor, and other benign lesions. Segmentation performance was evaluated using the Dice similarity coefficient and Intersection over Union, while classification performance was assessed using accuracy, macro AUC, F1-score, and Cohen’s kappa coefficient. Subgroup analyses stratified by age, gender, and tumor diameter were conducted. The proposed model was also compared with three single-task models. In addition, a reader study was performed to evaluate the impact of model assistance on the diagnostic performance of radiologists with different levels of experience. Results GobletNet demonstrated excellent performance in tumor segmentation, achieving Dice coefficients of 0.956, 0.959, and 0.965 for the internal and two external test datasets. For the classification task, the model achieved accuracies of 0.883, 0.841, and 0.843 for the internal and external test datasets, with corresponding one-vs-rest (OVR) macro-averaged AUCs of 0.980, 0.961, and 0.965, outperforming the comparison models. Subgroup analyses showed consistently stable performance across different subgroups (all AUCs > 0.950). In the reader study, GobletNet achieved an OVR macro-averaged AUC of 0.962, which was higher than all radiologists. Model assistance significantly improved radiologists’ diagnostic performance, with accuracy increased from 0.405 – 0.560 to 0.610–0.765. Conclusions GobletNet enables accurate segmentation and four-class classification of PGTs and demonstrates strong stability and generalizability across multicenter datasets. The model significantly improves the diagnostic performance of radiologists with different experience levels and may serve as an effective tool for the preoperative evaluation of PGTs.
Authors
- Lin Sui (ORCID: https://orcid.org/0000-0002-6565-0245)
- Dong Xu (ORCID: https://orcid.org/0000-0002-0583-240X)
- Qianmeng Pan
- Yujie Cai
- Xi Zhu
- Tian Jiang
- Bin Li
- Kai Wang
- Chen Chen
- Qian Li
- Na Feng
- Wei Wei
- Hui Wang
- Mei Song
- Lingyan Zhou
Institutions
- Zhejiang Chinese Medical University (CN)
- Wenzhou Medical University (CN)
- Zhejiang Cancer Hospital (CN)
- Zhejiang Lab (CN)
- Dongyang People's Hospital (CN)
- Zhejiang Taizhou Hospital (CN)
- First Affiliated Hospital of Wannan Medical College (CN)
Publication Details
- Journal
- BMC Medicine
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1186/s12916-026-05215-x
- Primary Topic
- Salivary Gland Tumors Diagnosis and Treatment
- Type
- article
- Field-Weighted Citation Impact
- 0.00