Cross-Modal Comprehensive Multi-Level Granularity Alignment for Vision-Language Pre-Training
Vision-Language Pre-training (VLP) aims to convert vision-language tasks into image-text matching challenges, necessitating robust interaction between visual and textual modalities. However, current research often faces a trade-off: methods pursuing fine-grained alignment typically rely on computationally expensive object detectors and pre-defined annotations, while coarse-grained alignment methods overlook subtle visual details. To alleviate these issues, we present OMEGA, a C O mprehensive Multi-l E vel G ranularity A lignment framework. OMEGA achieves sufficient cross-modal alignment without relying on external object detectors or annotations. Specifically, we first design a Text-Aware Patch Selection (TAPS) module, which dynamically constructs patch-level regions for fine-grained excavation based on informative token selection, effectively bypassing the need for bounding boxes. Secondly, to facilitate multi-grained information fusion, we introduce a Cross-Grained Alignment (CGA) module to learn modality-shared features across global and object levels. Experimental results demonstrate that OMEGA surpasses existing approaches in five pivotal vision-language tasks—image-text retrieval, visual question answering, visual reasoning, visual entailment, and image captioning, while maintaining superior inference efficiency compared to detector-based methods. We believe this study offers a promising object-free paradigm for scalable vision-language pre-training.
Authors
- Jia Wu (ORCID: https://orcid.org/0000-0002-1371-5801)
- Haoran Xie (ORCID: https://orcid.org/0000-0003-0965-3617)
- Darko B. Vukovic (ORCID: https://orcid.org/0000-0002-1165-489X)
- Youquan Wang (ORCID: https://orcid.org/0000-0003-4726-7493)
- Jie Cao (ORCID: https://orcid.org/0000-0002-7049-5614)
- ju jiang
Institutions
- Nanjing University of Finance and Economics (CN)
- Hefei University of Technology (CN)
- St Petersburg University (RU)
- Nanjing University of Science and Technology (CN)
- Macquarie University (AU)
Publication Details
- Journal
- ACM Transactions on Multimedia Computing Communications and Applications
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1145/3802546
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00