Attention-enhanced vision transformer hashing for hybrid image retrieval

Large-scale image retrieval requires compact representations without substantially sacrificing retrieval accuracy. However, Vision Transformer Hashing (VTS) concatenates all output tokens before hash projection, resulting in a high-dimensional hashing head with considerable model and memory overhead. We replace this token-concatenation operation with learnable-query multi-head attention pooling, which aggregates ViT tokens into a compact, image-adaptive representation for supervised hash-code learning. We further evaluate the proposed model using a standard two-stage retrieval procedure. In Stage 1, the attention-pooled representation is mapped to a binary code for efficient Hamming-space candidate selection. In Stage 2, the final-layer CLS descriptor from the same shared ViT-B/16 is used to re-rank the shortlisted candidates by cosine similarity. Controlled experiments on MS-COCO and NUS-WIDE compare attention pooling with GeM, average, CLS, and VTS-style concatenation. The proposed hybrid configuration achieves 91.13% and 88.63% mAP@5000 on MS-COCO and NUS-WIDE, respectively, demonstrating performance comparable to the VTS-style Concatenation baseline under the same controlled configuration, with marginal gains of 0.15 and 0.31 percentage points. It also exceeds separately trained 32-bit attention hash-only models by 2.23 and 1.30 points. Relative to concatenation, attention pooling reduces model footprint and batch-1 GPU memory by 63.1% and 62.0%, respectively, while increasing throughput by 21.7%. These results show that attention pooling substantially improves the efficiency of VTS while preserving retrieval effectiveness, whereas continuous re-ranking provides most of the gain over Hamming-only retrieval.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-17
DOI
https://doi.org/10.1007/s44163-026-02163-6
Primary Topic
Advanced Image and Video Retrieval Techniques
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Attention-enhanced vision transformer hashing for hybrid image retrieval

Quynh Dao Thi Thuy, Hoai Le Ba, Uyen Nguyen
Discover Artificial Intelligence
Advanced Image and Video Retrieval Techniques
article

Attention-enhanced vision transformer hashing for hybrid image retrieval

Quynh Dao Thi Thuy, Hoai Le Ba, Uyen Nguyen
article en

Abstract

Large-scale image retrieval requires compact representations without substantially sacrificing retrieval accuracy. However, Vision Transformer Hashing (VTS) concatenates all output tokens before hash projection, resulting in a high-dimensional hashing head with considerable model and memory overhead. We replace this token-concatenation operation with learnable-query multi-head attention pooling, which aggregates ViT tokens into a compact, image-adaptive representation for supervised hash-code learning. We further evaluate the proposed model using a standard two-stage retrieval procedure. In Stage 1, the attention-pooled representation is mapped to a binary code for efficient Hamming-space candidate selection. In Stage 2, the final-layer CLS descriptor from the same shared ViT-B/16 is used to re-rank the shortlisted candidates by cosine similarity. Controlled experiments on MS-COCO and NUS-WIDE compare attention pooling with GeM, average, CLS, and VTS-style concatenation. The proposed hybrid configuration achieves 91.13% and 88.63% mAP@5000 on MS-COCO and NUS-WIDE, respectively, demonstrating performance comparable to the VTS-style Concatenation baseline under the same controlled configuration, with marginal gains of 0.15 and 0.31 percentage points. It also exceeds separately trained 32-bit attention hash-only models by 2.23 and 1.30 points. Relative to concatenation, attention pooling reduces model footprint and batch-1 GPU memory by 63.1% and 62.0%, respectively, while increasing throughput by 21.7%. These results show that attention pooling substantially improves the efficiency of VTS while preserving retrieval effectiveness, whereas continuous re-ranking provides most of the gain over Hamming-only retrieval.

Discover Artificial IntelligenceVol. 6(1)
Research Institute of Posts and Telecommunications (SK), Posts and Telecommunications Institute of Technology, Hanoi University of Science and Technology (VN)
Posts and Telecommunications Institute of Technology, Instituto de Telecomunicações
Openalex Percentile: Top 14%
Advanced Image and Video Retrieval Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.