
Matteo Tortora, Elena Mulero Ayllón, Filippo Ruffini, Valerio Guarrasi, Paolo Soda · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-71014-2
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Tumor segmentation is a core task in medical image analysis, with direct implications for diagnosis, treatment planning, and disease monitoring. Whether promptable foundation models are mature enough for heterogeneous oncological scenarios remains an open question. We present a multi-cancer benchmark spanning lung, liver, kidney, brain, and breast tumor settings. Conventional supervised models (U-Net, DeepLabV3, Swin UNETR, nnU-Net) are compared against SAM-based foundation models (MedSAM and Medical SAM 2) under a common evaluation protocol. Prompt robustness is assessed by perturbing input bounding boxes through isotropic scaling and spatial shifting at inference time. Fine-tuned Medical SAM 2 with bounding-box prompting achieves the strongest benchmark-level profile, with the best results on Lung1, HCC, and KiTS23, while its zero-shot bounding-box configuration performs best on ATLAS. Swin UNETR ranks first on all three BraTS targets, and nnU-Net 3D full resolution performs best on ISPY1. Bounding-box prompting outperforms point-based guidance throughout the SAM-based family, and fine-tuning has a strong effect on performance. The robustness analysis reveals a trade-off: MedSAM tolerates prompt perturbations, while Medical SAM 2 achieves higher accuracy but degrades under box tightening and spatial shifts. These results support the use of SAM-based models for multi-cancer segmentation, while showing that reliability depends on prompt quality and anatomical context. Code is available at https://github.com/arco-group/tumor_benchmarking .
No comments yet — start the discussion below.