A unified unsupervised framework of multi-modal image super resolution and fusion via implicit neural representation

With the rise of dual-mode therapy and imaging, it is still challenging to accurately and precisely monitor structural organs during therapy. We present a self-supervised implicit neural representation (INR) framework to solve multi-modal image super resolution (MIR) and image fusion (MIF) for image guidance in therapy. In the purpose of preserving features and removing redundancy from the previously scanned images and therapeutic response, the common and unique features are represented respectively by means of segmented-based hash-grid sparse coding. The advantage is that only useful information is preserved while interfering information is discarded. Furthermore, the self-supervised optimization allows effectively training the multi-modal sparse coding INR network in the absence of ground truth. Extensive experiments are carried out on ultrasound (US), computed tomography (CT) and magnetic resonance imaging (MRI) data. In MIR task, our method achieves at least 4–7% structural similarity index measure (SSIM) improvement and 1–2% peak signal-to-noise ratio (PSNR) gain across all datasets, while learned perceptual image patch similarity (LPIPS) decreases by 2–14% and canny-based edge measure (Edge-F1) increases by 3–13%, compare to INR-based methods. In MIF task, our method achieves at least 2% improvement in SSIM, 6% enhancement in quality mutual information (QMI), and 6% increase in mutual information after fusion ( M I abf ) compared to existing INR-based approaches. Succinctly, the qualitative and quantitative results demonstrate the superior performance over the existing state-of-the-art methods in both tasks. The unified model eliminates the reliance on large-scale paired dataset for patient-specific image guidance.

Authors

Institutions

Publication Details

Journal
Biomedical Signal Processing and Control
Published
2026-09-18
DOI
https://doi.org/10.1016/j.bspc.2026.111486
Primary Topic
Advanced Image Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A unified unsupervised framework of multi-modal image super resolution and fusion via implicit neural representation

You Zheng, Ye Jiang, Xu Guo, Jintao Ni et al.
Biomedical Signal Processing and Control
Advanced Image Processing Techniques
article

A unified unsupervised framework of multi-modal image super resolution and fusion via implicit neural representation

You Zheng, Ye Jiang, Xu Guo, Jintao Ni, Bo Ma
article en

Abstract

With the rise of dual-mode therapy and imaging, it is still challenging to accurately and precisely monitor structural organs during therapy. We present a self-supervised implicit neural representation (INR) framework to solve multi-modal image super resolution (MIR) and image fusion (MIF) for image guidance in therapy. In the purpose of preserving features and removing redundancy from the previously scanned images and therapeutic response, the common and unique features are represented respectively by means of segmented-based hash-grid sparse coding. The advantage is that only useful information is preserved while interfering information is discarded. Furthermore, the self-supervised optimization allows effectively training the multi-modal sparse coding INR network in the absence of ground truth. Extensive experiments are carried out on ultrasound (US), computed tomography (CT) and magnetic resonance imaging (MRI) data. In MIR task, our method achieves at least 4–7% structural similarity index measure (SSIM) improvement and 1–2% peak signal-to-noise ratio (PSNR) gain across all datasets, while learned perceptual image patch similarity (LPIPS) decreases by 2–14% and canny-based edge measure (Edge-F1) increases by 3–13%, compare to INR-based methods. In MIF task, our method achieves at least 2% improvement in SSIM, 6% enhancement in quality mutual information (QMI), and 6% increase in mutual information after fusion ( M I abf ) compared to existing INR-based approaches. Succinctly, the qualitative and quantitative results demonstrate the superior performance over the existing state-of-the-art methods in both tasks. The unified model eliminates the reliance on large-scale paired dataset for patient-specific image guidance.

Biomedical Signal Processing and ControlVol. 129
Huazhong University of Science and Technology (CN)
National Key Research and Development Program of China Stem Cell and Translational Research
Openalex Percentile: Top 13%
Advanced Image Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.