Multicenter Validation of Foundation Model Adaptation for Automated Pancreatic Tumor Delineation on CT Scans

Background: Accurate pancreatic tumor segmentation on contrast-enhanced computed tomography (CECT) is important for staging, treatment planning, and response assessment in pancreatic ductal adenocarcinoma (PDAC). Although foundation models have shown promise for medical image segmentation, their effectiveness for disease-specific tumor delineation remains uncertain. This study evaluated whether fine-tuning a Segment Anything Model (SAM)-based framework on the target cohorts improves pancreatic tumor segmentation compared with directly applied foundation models. Methods: In this retrospective multicenter study, CECT examinations from patients with pathologically confirmed PDAC acquired between 2015 and 2025 were included. Two foundation model baselines, nnInteractive and MedSAM2, were compared with three TAGS-based configurations representing increasing levels of adaptation: TAGS (Zero-Shot), SAM-TAGS (fine-tuned from SAM weights), and MSD-TAGS (fine-tuned from a pancreas-specific checkpoint). Performance was assessed using patient-level five-fold cross-validation with the Dice Similarity Coefficient (DSC) and Normalized Surface Dice (NSD). A leave-one-center-out analysis was additionally performed to evaluate generalization to institutions absent from training. Results: A total of 500 patients with PDAC from four independent centers were included. MSD-TAGS achieved the highest overall segmentation performance, with a mean DSC of 0.703 and a mean NSD of 0.835. SAM-TAGS achieved a mean DSC of 0.691 and a mean NSD of 0.822. In comparison, nnInteractive achieved a mean DSC of 0.632 and a mean NSD of 0.782; MedSAM2, 0.585 and 0.792; and TAGS (Zero-Shot), 0.496 and 0.639, respectively. Compared with each competing method, MSD-TAGS achieved significantly higher DSC and NSD values (all Holm-adjusted p < 0.001). It also achieved the highest DSC across all four centers and the highest NSD in three of the four centers. Under leave-one-center-out evaluation, MSD-TAGS declined by 0.025 DSC and SAM-TAGS by 0.089. Conclusions: Task-specific adaptation improved pancreatic tumor segmentation with foundation models, and fine-tuned TAGS outperformed two publicly released promptable models applied without adaptation. Models initialized from a pancreas-specific checkpoint showed greater robustness under center-held-out evaluation. These findings support the importance of organ-specific initialization and target-cohort fine-tuning for disease-specific segmentation and warrant further validation before clinical translation. The multicenter pancreatic tumor CT segmentation dataset, including de-identified images and expert segmentation annotations, is publicly available to facilitate future research.

Authors

Institutions

Publication Details

Journal
Cancers
Published
2026-09-01
DOI
https://doi.org/10.3390/cancers18172836
Primary Topic
Pancreatic and Hepatic Oncology Research
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multicenter Validation of Foundation Model Adaptation for Automated Pancreatic Tumor Delineation on CT Scans

Murat Iren, Hongyi Pan, Görkem Durak, Linkai Peng et al.
Cancers
Pancreatic and Hepatic Oncology Research
article

Multicenter Validation of Foundation Model Adaptation for Automated Pancreatic Tumor Delineation on CT Scans

Murat Iren, Hongyi Pan, Görkem Durak, Linkai Peng, Kadir Atakır, Şükrü Mehmet Ertürk, Emre Uysal, Sıtkı Safa Taflan, Wanying Dou, Mucahit Ekici, Oyku Ikizgul, Elif Keleş, Mustafa Orhan Nalbant, Zongwei Zhou, Andrea Bejar, Okan Çetin, Yavuz Taktak, Nurullah Kaya, Alpay Medetalibeyoğlu, Burak Gültekin, Halil Ertugrul Aktas, Gulbiz Dagoglu Kartal, Fergan Bol, Maide Müreva, Frank H. Miller, Muhammed Enes Tasci, Alper Akin, Eminenur Sen Tasci, Baver Tutun, Berna Akkus Yildirim, Ulas Bagci
article en

Abstract

Background: Accurate pancreatic tumor segmentation on contrast-enhanced computed tomography (CECT) is important for staging, treatment planning, and response assessment in pancreatic ductal adenocarcinoma (PDAC). Although foundation models have shown promise for medical image segmentation, their effectiveness for disease-specific tumor delineation remains uncertain. This study evaluated whether fine-tuning a Segment Anything Model (SAM)-based framework on the target cohorts improves pancreatic tumor segmentation compared with directly applied foundation models. Methods: In this retrospective multicenter study, CECT examinations from patients with pathologically confirmed PDAC acquired between 2015 and 2025 were included. Two foundation model baselines, nnInteractive and MedSAM2, were compared with three TAGS-based configurations representing increasing levels of adaptation: TAGS (Zero-Shot), SAM-TAGS (fine-tuned from SAM weights), and MSD-TAGS (fine-tuned from a pancreas-specific checkpoint). Performance was assessed using patient-level five-fold cross-validation with the Dice Similarity Coefficient (DSC) and Normalized Surface Dice (NSD). A leave-one-center-out analysis was additionally performed to evaluate generalization to institutions absent from training. Results: A total of 500 patients with PDAC from four independent centers were included. MSD-TAGS achieved the highest overall segmentation performance, with a mean DSC of 0.703 and a mean NSD of 0.835. SAM-TAGS achieved a mean DSC of 0.691 and a mean NSD of 0.822. In comparison, nnInteractive achieved a mean DSC of 0.632 and a mean NSD of 0.782; MedSAM2, 0.585 and 0.792; and TAGS (Zero-Shot), 0.496 and 0.639, respectively. Compared with each competing method, MSD-TAGS achieved significantly higher DSC and NSD values (all Holm-adjusted p < 0.001). It also achieved the highest DSC across all four centers and the highest NSD in three of the four centers. Under leave-one-center-out evaluation, MSD-TAGS declined by 0.025 DSC and SAM-TAGS by 0.089. Conclusions: Task-specific adaptation improved pancreatic tumor segmentation with foundation models, and fine-tuned TAGS outperformed two publicly released promptable models applied without adaptation. Models initialized from a pancreas-specific checkpoint showed greater robustness under center-held-out evaluation. These findings support the importance of organ-specific initialization and target-cohort fine-tuning for disease-specific segmentation and warrant further validation before clinical translation. The multicenter pancreatic tumor CT segmentation dataset, including de-identified images and expert segmentation annotations, is publicly available to facilitate future research.

CancersVol. 18(17)
Northwestern University (US), Üsküdar University (TR), Boston Medical Center (US), Intel (United States) (US), Johns Hopkins University (US), University of Chicago (US), Bakırköy Dr.Sadi Konuk Eğitim ve Araştırma Hastanesi (TR), Sağlık Bilimleri Üniversitesi (TR), Istanbul University (TR), Harran University (TR), University of Health Sciences Antigua (AG)
National Institutes of Health
Good health and well-being
Openalex Percentile: Top 14%
Pancreatic and Hepatic Oncology Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.