DE-SwinJSCC: dual-enhanced swin transformer for wireless image semantic transmission
With the successful adoption of Transformer architectures in visual tasks, end-to-end joint source channel coding (JSCC) has emerged as a promising paradigm for high-resolution wireless semantic image transmission. However, existing Transformer-based approaches, particularly SwinJSCC, still suffer from several limitations in high-resolution scenarios, including insufficient spatial detail modeling capability, limited channel-aware semantic discrimination, and an unfavorable trade-off between reconstruction performance and computational complexity under the considered AWGN channel. To address these challenges, this paper proposes a novel end-to-end JSCC framework termed DE-SwinJSCC. Built upon the hierarchical Swin Transformer architecture, we propose a dual-enhanced DE-Swin block for joint source–channel coding. The proposed block preserves the original window-based self-attention mechanism while enhancing the feed-forward pathway through two complementary components: a spatial enhancement module (Mona) and a semantic–channel joint modeling module (SCA-MLP). Mona is a multi-scale spatial enhancement module that injects explicit structural priors into feature representations, improving local geometry and texture modeling under AWGN channel noise. SCA-MLP is a semantic-aware channel modulation module that incorporates channel-wise feature recalibration into the feed-forward process, enhancing semantic discrimination and reconstruction robustness under different signal-to-noise ratio (SNR) conditions in the AWGN channel. This design enables simultaneous reinforcement of local structural representation and cross-channel semantic discriminability. Extensive experimental results demonstrate that, under different AWGN channel SNRs and bandwidth constraints, DE-SwinJSCC consistently outperforms existing methods in terms of reconstruction quality and structural fidelity. In particular, on the high-resolution Kodak dataset, the proposed method achieves up to approximately 2.26 dB improvement in PSNR and about $$0.37\\%$$ gain in MS-SSIM compared with SwinJSCC, validating the effectiveness of the proposed dual-enhancement strategy for high-resolution semantic image transmission.
Authors
- Yunyi Liu (ORCID: https://orcid.org/0000-0002-5138-9461)
- Haiqiang Chen (ORCID: https://orcid.org/0000-0002-0694-1595)
- Youming Sun (ORCID: https://orcid.org/0000-0001-9963-3876)
- Jibin Lu
- Xiangcheng Li
- Dongri Ban
Institutions
- Guangxi University (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1038/s41598-026-66105-z
- Primary Topic
- Wireless Signal Modulation Classification
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Natural Science Foundation of China