Cross-Modal Comprehensive Multi-Level Granularity Alignment for Vision-Language Pre-Training

Vision-Language Pre-training (VLP) aims to convert vision-language tasks into image-text matching challenges, necessitating robust interaction between visual and textual modalities. However, current research often faces a trade-off: methods pursuing fine-grained alignment typically rely on computationally expensive object detectors and pre-defined annotations, while coarse-grained alignment methods overlook subtle visual details. To alleviate these issues, we present OMEGA, a C O mprehensive Multi-l E vel G ranularity A lignment framework. OMEGA achieves sufficient cross-modal alignment without relying on external object detectors or annotations. Specifically, we first design a Text-Aware Patch Selection (TAPS) module, which dynamically constructs patch-level regions for fine-grained excavation based on informative token selection, effectively bypassing the need for bounding boxes. Secondly, to facilitate multi-grained information fusion, we introduce a Cross-Grained Alignment (CGA) module to learn modality-shared features across global and object levels. Experimental results demonstrate that OMEGA surpasses existing approaches in five pivotal vision-language tasks—image-text retrieval, visual question answering, visual reasoning, visual entailment, and image captioning, while maintaining superior inference efficiency compared to detector-based methods. We believe this study offers a promising object-free paradigm for scalable vision-language pre-training.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Multimedia Computing Communications and Applications
Published
2026-09-24
DOI
https://doi.org/10.1145/3802546
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Cross-Modal Comprehensive Multi-Level Granularity Alignment for Vision-Language Pre-Training

Jia Wu, Haoran Xie, Darko B. Vukovic, Youquan Wang et al.
ACM Transactions on Multimedia Computing Communications and Applications
Multimodal Machine Learning Applications
article

Cross-Modal Comprehensive Multi-Level Granularity Alignment for Vision-Language Pre-Training

Jia Wu, Haoran Xie, Darko B. Vukovic, Youquan Wang, Jie Cao, ju jiang
article en

Abstract

Vision-Language Pre-training (VLP) aims to convert vision-language tasks into image-text matching challenges, necessitating robust interaction between visual and textual modalities. However, current research often faces a trade-off: methods pursuing fine-grained alignment typically rely on computationally expensive object detectors and pre-defined annotations, while coarse-grained alignment methods overlook subtle visual details. To alleviate these issues, we present OMEGA, a C O mprehensive Multi-l E vel G ranularity A lignment framework. OMEGA achieves sufficient cross-modal alignment without relying on external object detectors or annotations. Specifically, we first design a Text-Aware Patch Selection (TAPS) module, which dynamically constructs patch-level regions for fine-grained excavation based on informative token selection, effectively bypassing the need for bounding boxes. Secondly, to facilitate multi-grained information fusion, we introduce a Cross-Grained Alignment (CGA) module to learn modality-shared features across global and object levels. Experimental results demonstrate that OMEGA surpasses existing approaches in five pivotal vision-language tasks—image-text retrieval, visual question answering, visual reasoning, visual entailment, and image captioning, while maintaining superior inference efficiency compared to detector-based methods. We believe this study offers a promising object-free paradigm for scalable vision-language pre-training.

ACM Transactions on Multimedia Computing Communications and Applications
Nanjing University of Finance and Economics (CN), Hefei University of Technology (CN), St Petersburg University (RU), Nanjing University of Science and Technology (CN), Macquarie University (AU)
Quality Education
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.