Statistical Model-Driven Similarity Hashing for Unsupervised Multimedia Retrieval.

Unsupervised deep cross-modal hash retrieval aims to map multi-modal features into binary hash codes without labels, which is of interest due to the storage efficiency, query speed, and convenient applications. However, existing approaches suffer from two main limitations: (1) Insufficient consideration of text instance similarity, along with independent or redundant fusion to learn multi-modal similarity information. (2) Ignoring the noisy adjacent correlations between multi-modal instances leads to a lack of discriminative capacity in the generated hash codes. To address the challenges, we propose a novel approach called Statistical Model-driven Similarity Hashing. Specifically, Jaccard similarity is introduced to construct the text similarity matrix, which reduces the similarity error between text instances while better considering the asymmetry of elements in text features. After that, original similarity information between various modalities is integrated to construct a unified similarity matrix, where the modality gaps are effectively bridged while reducing the redundant information. In addition, we present a Statistical Model-driven Similarity Enhancement approach, which reduces the noise of similarity relations between multi-modal instances by utilizing Gaussian Mixture Models to keep instances with lower semantic similarity as far away from each other as possible. Moreover, we introduce an enhanced version, which integrates a structure-aware self-paced contrastive learning framework that progressively schedules reliable and fuzzy-positive pairs, jointly optimizing statistical reconstruction and discriminative alignment, and further extends to noisy environments. Extensive experiments on three benchmark datasets demonstrate the excellent performance of the proposed method.

Authors

Publication Details

Journal
PubMed
Published
2026-09-11
DOI
https://doi.org/10.1109/tip.2026.3730836
Primary Topic
Advanced Image and Video Retrieval Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Statistical Model-Driven Similarity Hashing for Unsupervised Multimedia Retrieval.

Binghong Chen, Longzhi Sun, Zhan Yang, Mingjin Kuai et al.
PubMed
Advanced Image and Video Retrieval Techniques
article

Statistical Model-Driven Similarity Hashing for Unsupervised Multimedia Retrieval.

Binghong Chen, Longzhi Sun, Zhan Yang, Mingjin Kuai, Yinan Li
article en

Abstract

Unsupervised deep cross-modal hash retrieval aims to map multi-modal features into binary hash codes without labels, which is of interest due to the storage efficiency, query speed, and convenient applications. However, existing approaches suffer from two main limitations: (1) Insufficient consideration of text instance similarity, along with independent or redundant fusion to learn multi-modal similarity information. (2) Ignoring the noisy adjacent correlations between multi-modal instances leads to a lack of discriminative capacity in the generated hash codes. To address the challenges, we propose a novel approach called Statistical Model-driven Similarity Hashing. Specifically, Jaccard similarity is introduced to construct the text similarity matrix, which reduces the similarity error between text instances while better considering the asymmetry of elements in text features. After that, original similarity information between various modalities is integrated to construct a unified similarity matrix, where the modality gaps are effectively bridged while reducing the redundant information. In addition, we present a Statistical Model-driven Similarity Enhancement approach, which reduces the noise of similarity relations between multi-modal instances by utilizing Gaussian Mixture Models to keep instances with lower semantic similarity as far away from each other as possible. Moreover, we introduce an enhanced version, which integrates a structure-aware self-paced contrastive learning framework that progressively schedules reliable and fuzzy-positive pairs, jointly optimizing statistical reconstruction and discriminative alignment, and further extends to noisy environments. Extensive experiments on three benchmark datasets demonstrate the excellent performance of the proposed method.

PubMedVol. PP
Reduced inequalities
Openalex Percentile: Top 13%
Advanced Image and Video Retrieval Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Statistical Model-Driven Similarity Hashing for Unsupervised Multimedia Retrieval. — Binghong Chen, Longzhi Sun, et al. · PubMed (2026) | TGRS Research Map | TGRS