Holistic Satellite Image Super Resolution Using Large Diffuse Generative Models

Satellite image super-resolution (SR) is a critical task in machine vision, supporting applications such as urban planning, agricultural monitoring, and disaster response. The objective of SR is to enhance the spatial resolution of low-resolution (LR) satellite imagery, thereby extracting finer detail. Over the past two decades, numerous SR methods have been developed, primarily leveraging deep learning to achieve significant resolution enhancements. However, prevailing approaches exhibit two principal limitations. First, many methods operate via analytic mathematical techniques or rely on learned priors from training data, without explicitly incorporating structured scene knowledge (e.g., road networks, building footprints). Second, existing methods offer minimal control over the aesthetic and structural characteristics of the output during the SR process, which is a drawback for applications like digital map updating where consistency with existing geospatial databases is essential. In scenarios where rich prior semantic information about a scene is available—such as from OpenStreetMap or other GIS sources—it is natural to harness this data to guide and improve SR reconstruction.This paper introduces a semantic super-resolution (SSR) model designed to address these gaps by integrating semantic priors directly into the SR pipeline. The approach builds upon the Flux Kontext diffusion-based generative framework, incorporating two key modifications. Firstly, an additional textual encoder is trained to convert available semantic information—such as vector data describing roads, buildings, and water bodies—into a dense textual prior that is fed into the network. Secondly, the loss function is refined via a Low-Rank Adaptation (LoRA) adapter, emphasizing semantic consistency of major features in the generated highresolution (HR) image.The proposed SSR model was evaluated against three state-of-the-art deep learning SR baselines: EDSR, SRGAN, and ResShift, using the high-resolution aerial imagery dataset ( HRAID ) comprising 236 images. Qualitative assessment reveals enhanced consistency in the reconstruction of structured features like roads and buildings, attributable to the injected textual prior. Quantitative evaluation demonstrates that our model outperforms the best-performing baseline by 5% in PSNR and by 10% in FID. An ablation study systematically removing components of the SSR model confirms the necessity of both the textual encoder and the semantic-consistency-focused LoRA adapter for achieving these gains.

Authors

Institutions

Publication Details

Journal
˜The œinternational archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences
Published
2026-09-29
DOI
https://doi.org/10.5194/isprs-archives-l-4-w3-2026-75-2026
Primary Topic
Advanced Image Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Holistic Satellite Image Super Resolution Using Large Diffuse Generative Models

Vladimir Alexandrovich Knyaz, Vladimir Vladimirovich Kniaz, Petr V. Moskantsev, Artem N. Bordodymov et al.
˜The œinternational archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences
Advanced Image Processing Techniques
article

Holistic Satellite Image Super Resolution Using Large Diffuse Generative Models

Vladimir Alexandrovich Knyaz, Vladimir Vladimirovich Kniaz, Petr V. Moskantsev, Artem N. Bordodymov, Victor S. Aleksandrov, Egor R. Smirnov
article en

Abstract

Satellite image super-resolution (SR) is a critical task in machine vision, supporting applications such as urban planning, agricultural monitoring, and disaster response. The objective of SR is to enhance the spatial resolution of low-resolution (LR) satellite imagery, thereby extracting finer detail. Over the past two decades, numerous SR methods have been developed, primarily leveraging deep learning to achieve significant resolution enhancements. However, prevailing approaches exhibit two principal limitations. First, many methods operate via analytic mathematical techniques or rely on learned priors from training data, without explicitly incorporating structured scene knowledge (e.g., road networks, building footprints). Second, existing methods offer minimal control over the aesthetic and structural characteristics of the output during the SR process, which is a drawback for applications like digital map updating where consistency with existing geospatial databases is essential. In scenarios where rich prior semantic information about a scene is available—such as from OpenStreetMap or other GIS sources—it is natural to harness this data to guide and improve SR reconstruction.This paper introduces a semantic super-resolution (SSR) model designed to address these gaps by integrating semantic priors directly into the SR pipeline. The approach builds upon the Flux Kontext diffusion-based generative framework, incorporating two key modifications. Firstly, an additional textual encoder is trained to convert available semantic information—such as vector data describing roads, buildings, and water bodies—into a dense textual prior that is fed into the network. Secondly, the loss function is refined via a Low-Rank Adaptation (LoRA) adapter, emphasizing semantic consistency of major features in the generated highresolution (HR) image.The proposed SSR model was evaluated against three state-of-the-art deep learning SR baselines: EDSR, SRGAN, and ResShift, using the high-resolution aerial imagery dataset ( HRAID ) comprising 236 images. Qualitative assessment reveals enhanced consistency in the reconstruction of structured features like roads and buildings, attributable to the injected textual prior. Quantitative evaluation demonstrates that our model outperforms the best-performing baseline by 5% in PSNR and by 10% in FID. An ablation study systematically removing components of the SSR model confirms the necessity of both the textual encoder and the semantic-consistency-focused LoRA adapter for achieving these gains.

˜The œinternational archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciencesVol. L-4/W3-2026(0)
Moscow Institute of Physics and Technology (RU), State Scientific Research Institute of Aviation Systems (RU)
Sustainable cities and communities
Openalex Percentile: Top 14%
Advanced Image Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.