DAV‐Diff: A Diffusion Model for High‐Quality Remote Sensing Image Reconstruction Based on Multi‐Level Attention and v‐Prediction

ABSTRACT With the advancement of deep learning, diffusion models have gained prominence as a powerful method for high‐fidelity image generation by progressively removing noise to restore structural and textural details. In remote sensing image super‐resolution (RSISR), which typically utilizes paired high‐resolution (HR) and low‐resolution (LR) images for supervised training, applying diffusion models directly faces three persistent challenges. First, common preprocessing techniques such as bicubic interpolation lose substantial prior information, resulting in a lack of high‐frequency details and weakened spatial structural features. Second, existing denoising networks often exhibit insufficient global modeling capability, performing poorly in capturing long‐range spatial relationships and evaluating channel‐wise feature importance. Third, traditional noise prediction (‐prediction) struggles to restore fine details under high‐noise conditions, frequently failing to preserve the complex structures of ground objects in remote sensing scenes. To address these issues, this paper proposes DAV‐Diff, a fast‐sampling conditional diffusion model for RSISR. DAV‐Diff coordinates three task‐oriented components to improve conditional prior representation, long‐range dependency modeling, and high‐noise detail recovery in RSISR. The first is a multi‐dilated residual attention block (MD‐RAB), which enhances the extraction of prior information from LR images and compensates for detail loss caused by interpolation. The second component is a dual‐attention mechanism embedded within the denoising network to strengthen long‐range spatial dependency modeling and adaptive channel feature weighting. The third is a residual‐domain ‐prediction strategy that replaces traditional noise prediction, mitigating detail degradation in high‐noise environments while preserving structural features of complex ground objects. Experimental results show the effectiveness of DAV‐Diff against state‐of‐the‐art methods; ablation studies, computational complexity analysis, and sampling‐step sensitivity experiments further quantify the contribution and efficiency of each component.

Authors

Institutions

Publication Details

Journal
Concurrency and Computation Practice and Experience
Published
2026-08-26
DOI
https://doi.org/10.1002/cpe.70913
Primary Topic
Advanced Image Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

DAV‐Diff: A Diffusion Model for High‐Quality Remote Sensing Image Reconstruction Based on Multi‐Level Attention and v‐Prediction

Chen Li, Shanxu Wu, Lihua Tian
Concurrency and Computation Practice and Experience
Advanced Image Processing Techniques
article

DAV‐Diff: A Diffusion Model for High‐Quality Remote Sensing Image Reconstruction Based on Multi‐Level Attention and v‐Prediction

Chen Li, Shanxu Wu, Lihua Tian
article en

Abstract

ABSTRACT With the advancement of deep learning, diffusion models have gained prominence as a powerful method for high‐fidelity image generation by progressively removing noise to restore structural and textural details. In remote sensing image super‐resolution (RSISR), which typically utilizes paired high‐resolution (HR) and low‐resolution (LR) images for supervised training, applying diffusion models directly faces three persistent challenges. First, common preprocessing techniques such as bicubic interpolation lose substantial prior information, resulting in a lack of high‐frequency details and weakened spatial structural features. Second, existing denoising networks often exhibit insufficient global modeling capability, performing poorly in capturing long‐range spatial relationships and evaluating channel‐wise feature importance. Third, traditional noise prediction (‐prediction) struggles to restore fine details under high‐noise conditions, frequently failing to preserve the complex structures of ground objects in remote sensing scenes. To address these issues, this paper proposes DAV‐Diff, a fast‐sampling conditional diffusion model for RSISR. DAV‐Diff coordinates three task‐oriented components to improve conditional prior representation, long‐range dependency modeling, and high‐noise detail recovery in RSISR. The first is a multi‐dilated residual attention block (MD‐RAB), which enhances the extraction of prior information from LR images and compensates for detail loss caused by interpolation. The second component is a dual‐attention mechanism embedded within the denoising network to strengthen long‐range spatial dependency modeling and adaptive channel feature weighting. The third is a residual‐domain ‐prediction strategy that replaces traditional noise prediction, mitigating detail degradation in high‐noise environments while preserving structural features of complex ground objects. Experimental results show the effectiveness of DAV‐Diff against state‐of‐the‐art methods; ablation studies, computational complexity analysis, and sampling‐step sensitivity experiments further quantify the contribution and efficiency of each component.

Concurrency and Computation Practice and ExperienceVol. 38(17)
Xi'an Jiaotong University (CN)
Openalex Percentile: Top 12%
Advanced Image Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.