PDSVF: Prompt-driven SAM-Med2D with parallel ViT and state-space fusion for skin lesion segmentation

Accurate segmentation of skin lesions in dermoscopic images underpins computer-aided diagnosis of skin cancer. The task is difficult: lesions show blurred, low-contrast boundaries, vary widely in scale, shape, and color, and are often obscured by hair and other imaging artifacts. Convolutional networks have limited receptive fields, whereas Transformers model global context at quadratic cost yet tend to smooth fine boundaries. Foundation models such as SAM-Med2D offer transferable representations, but adapting them while preserving global semantics and boundary fidelity remains unresolved. We therefore propose PDSVF, a prompt-driven adaptation of SAM-Med2D with parallel ViT and state-space fusion for skin lesion segmentation. Within the image encoder, a ViT-FusionMamba (VTFM) block runs parallel self-attention and FusionMamba branches. The FusionMamba branch couples four-directional 2D selective-scan state-space modeling with a learnable dynamic convolution for texture and boundary cues. A gated fusion mechanism adaptively balances both branches, and an end-to-end dual-layer residual connection stabilizes optimization. Adapter modules keep the pretrained self-attention weights of the backbone frozen, and a prompt encoder injects a weak spatial prompt. Across four public benchmarks (ISIC 2017, ISIC 2018, and two cross-dataset test sets, PH2 and the DermIS and DermQuest collection), PDSVF attains Dice scores of 92.92%, 94.20%, 95.67%, and 93.49%, and IoU scores of 87.15%, 89.24%, 91.83%, and 87.93%, respectively. Every reported value is the mean over three random seeds. Among methods evaluated under the same protocol, PDSVF achieves the highest Dice and IoU. PDSVF provides an effective route to adapt vision foundation models for accurate skin lesion segmentation https://github.com/CCEB1996/PDSVF .

Authors

Institutions

Publication Details

Journal
Biomedical Signal Processing and Control
Published
2026-09-28
DOI
https://doi.org/10.1016/j.bspc.2026.111557
Primary Topic
Cutaneous Melanoma Detection and Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

PDSVF: Prompt-driven SAM-Med2D with parallel ViT and state-space fusion for skin lesion segmentation

Shuang Hu, Tongtong Huo, Zhiwei Wang, Zhewei Ye et al.
Biomedical Signal Processing and Control
Cutaneous Melanoma Detection and Management
article

PDSVF: Prompt-driven SAM-Med2D with parallel ViT and state-space fusion for skin lesion segmentation

Shuang Hu, Tongtong Huo, Zhiwei Wang, Zhewei Ye, Jin Liu, Qifeng Hou, Wei Wu
article en

Abstract

Accurate segmentation of skin lesions in dermoscopic images underpins computer-aided diagnosis of skin cancer. The task is difficult: lesions show blurred, low-contrast boundaries, vary widely in scale, shape, and color, and are often obscured by hair and other imaging artifacts. Convolutional networks have limited receptive fields, whereas Transformers model global context at quadratic cost yet tend to smooth fine boundaries. Foundation models such as SAM-Med2D offer transferable representations, but adapting them while preserving global semantics and boundary fidelity remains unresolved. We therefore propose PDSVF, a prompt-driven adaptation of SAM-Med2D with parallel ViT and state-space fusion for skin lesion segmentation. Within the image encoder, a ViT-FusionMamba (VTFM) block runs parallel self-attention and FusionMamba branches. The FusionMamba branch couples four-directional 2D selective-scan state-space modeling with a learnable dynamic convolution for texture and boundary cues. A gated fusion mechanism adaptively balances both branches, and an end-to-end dual-layer residual connection stabilizes optimization. Adapter modules keep the pretrained self-attention weights of the backbone frozen, and a prompt encoder injects a weak spatial prompt. Across four public benchmarks (ISIC 2017, ISIC 2018, and two cross-dataset test sets, PH2 and the DermIS and DermQuest collection), PDSVF attains Dice scores of 92.92%, 94.20%, 95.67%, and 93.49%, and IoU scores of 87.15%, 89.24%, 91.83%, and 87.93%, respectively. Every reported value is the mean over three random seeds. Among methods evaluated under the same protocol, PDSVF achieves the highest Dice and IoU. PDSVF provides an effective route to adapt vision foundation models for accurate skin lesion segmentation https://github.com/CCEB1996/PDSVF .

Biomedical Signal Processing and ControlVol. 130
Rensselaer Polytechnic Institute (US), Wuhan National Laboratory for Optoelectronics (CN), Wuhan University of Science and Technology (CN), Huazhong University of Science and Technology (CN), Beihang University (CN)
Openalex Percentile: Top 15%
Cutaneous Melanoma Detection and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.