Linking Vision and Meaning: Text-Image Correspondence in Cultural Heritage Documentation
Cultural heritage (CH) reports contain rich descriptive content, but their rigid linear narrative structure necessitates the integration of disjointed parts of the text and auxiliary visual content (e.g., photographs) by the reader. To address this challenge, we explore the impact of the textual cues on cross-modal alignment performance. Recently, visual-language models (VLM) have been proposed for the automated establishment of text-image correspondence, but these have not been extensively evaluated in the context of domain knowledge (e.g., in CH). We conduct two experiments: (1) an image-to-text retrieval, analyzing how VLMs capture semantic overlap; and (2) a text-to-image visual grounding experiment evaluating how contextual clarity and language affect component-level detection accuracy. Our results show that current VLMs struggle to capture the fine-grained correspondences present in CH-related descriptive narratives due to semantic coarseness and limited relational understanding. Building on these insights, we adopt a grounding-based strategy that connects textual mentions and visual counterparts through prompt engineering and object-level anchoring that combines holistic and regional semantic alignment. Our study provides a structured evaluation of VLMs for multimodal integration in cultural heritage contexts, highlighting their limitations in fine-grained alignment. This contributes to a vision of interactive utilization of multimodal cultural heritage resources through improved connectivity between textual and visual content.
Authors
- Kourosh Khoshelham (ORCID: https://orcid.org/0000-0001-6639-1727)
- C. Chen (ORCID: https://orcid.org/0000-0002-6527-3902)
- Martin Tomko (ORCID: https://orcid.org/0000-0002-5736-4679)
Institutions
- The University of Melbourne (AU)
Publication Details
- Journal
- Journal on Computing and Cultural Heritage
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1145/3848029
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00