Benchmarking foundation models for tumor segmentation across multiple cancer types

Abstract Tumor segmentation is a core task in medical image analysis, with direct implications for diagnosis, treatment planning, and disease monitoring. Whether promptable foundation models are mature enough for heterogeneous oncological scenarios remains an open question. We present a multi-cancer benchmark spanning lung, liver, kidney, brain, and breast tumor settings. Conventional supervised models (U-Net, DeepLabV3, Swin UNETR, nnU-Net) are compared against SAM-based foundation models (MedSAM and Medical SAM 2) under a common evaluation protocol. Prompt robustness is assessed by perturbing input bounding boxes through isotropic scaling and spatial shifting at inference time. Fine-tuned Medical SAM 2 with bounding-box prompting achieves the strongest benchmark-level profile, with the best results on Lung1, HCC, and KiTS23, while its zero-shot bounding-box configuration performs best on ATLAS. Swin UNETR ranks first on all three BraTS targets, and nnU-Net 3D full resolution performs best on ISPY1. Bounding-box prompting outperforms point-based guidance throughout the SAM-based family, and fine-tuning has a strong effect on performance. The robustness analysis reveals a trade-off: MedSAM tolerates prompt perturbations, while Medical SAM 2 achieves higher accuracy but degrades under box tightening and spatial shifts. These results support the use of SAM-based models for multi-cancer segmentation, while showing that reliability depends on prompt quality and anatomical context. Code is available at https://github.com/arco-group/tumor_benchmarking .

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-19
DOI
https://doi.org/10.1038/s41598-026-71014-2
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking foundation models for tumor segmentation across multiple cancer types

Paolo Soda, Matteo Tortora, Valerio Guarrasi, Filippo Ruffini et al.
Scientific Reports
Advanced Neural Network Applications
article

Benchmarking foundation models for tumor segmentation across multiple cancer types

Paolo Soda, Matteo Tortora, Valerio Guarrasi, Filippo Ruffini, Elena Mulero Ayllón
article en

Abstract

Abstract Tumor segmentation is a core task in medical image analysis, with direct implications for diagnosis, treatment planning, and disease monitoring. Whether promptable foundation models are mature enough for heterogeneous oncological scenarios remains an open question. We present a multi-cancer benchmark spanning lung, liver, kidney, brain, and breast tumor settings. Conventional supervised models (U-Net, DeepLabV3, Swin UNETR, nnU-Net) are compared against SAM-based foundation models (MedSAM and Medical SAM 2) under a common evaluation protocol. Prompt robustness is assessed by perturbing input bounding boxes through isotropic scaling and spatial shifting at inference time. Fine-tuned Medical SAM 2 with bounding-box prompting achieves the strongest benchmark-level profile, with the best results on Lung1, HCC, and KiTS23, while its zero-shot bounding-box configuration performs best on ATLAS. Swin UNETR ranks first on all three BraTS targets, and nnU-Net 3D full resolution performs best on ISPY1. Bounding-box prompting outperforms point-based guidance throughout the SAM-based family, and fine-tuning has a strong effect on performance. The robustness analysis reveals a trade-off: MedSAM tolerates prompt perturbations, while Medical SAM 2 achieves higher accuracy but degrades under box tightening and spatial shifts. These results support the use of SAM-based models for multi-cancer segmentation, while showing that reliability depends on prompt quality and anatomical context. Code is available at https://github.com/arco-group/tumor_benchmarking .

Scientific Reports
Università Campus Bio-Medico (IT), University of Genoa (IT), Umeå University (SE)
Good health and well-being, Partnerships for the goals
Openalex Percentile: Top 13%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Benchmarking foundation models for tumor segmentation across multiple cancer types — Paolo Soda, Matteo Tortora, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS