Linking Vision and Meaning: Text-Image Correspondence in Cultural Heritage Documentation

Cultural heritage (CH) reports contain rich descriptive content, but their rigid linear narrative structure necessitates the integration of disjointed parts of the text and auxiliary visual content (e.g., photographs) by the reader. To address this challenge, we explore the impact of the textual cues on cross-modal alignment performance. Recently, visual-language models (VLM) have been proposed for the automated establishment of text-image correspondence, but these have not been extensively evaluated in the context of domain knowledge (e.g., in CH). We conduct two experiments: (1) an image-to-text retrieval, analyzing how VLMs capture semantic overlap; and (2) a text-to-image visual grounding experiment evaluating how contextual clarity and language affect component-level detection accuracy. Our results show that current VLMs struggle to capture the fine-grained correspondences present in CH-related descriptive narratives due to semantic coarseness and limited relational understanding. Building on these insights, we adopt a grounding-based strategy that connects textual mentions and visual counterparts through prompt engineering and object-level anchoring that combines holistic and regional semantic alignment. Our study provides a structured evaluation of VLMs for multimodal integration in cultural heritage contexts, highlighting their limitations in fine-grained alignment. This contributes to a vision of interactive utilization of multimodal cultural heritage resources through improved connectivity between textual and visual content.

Authors

Institutions

Publication Details

Journal
Journal on Computing and Cultural Heritage
Published
2026-09-15
DOI
https://doi.org/10.1145/3848029
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Linking Vision and Meaning: Text-Image Correspondence in Cultural Heritage Documentation

Kourosh Khoshelham, C. Chen, Martin Tomko
Journal on Computing and Cultural Heritage
Multimodal Machine Learning Applications
article

Linking Vision and Meaning: Text-Image Correspondence in Cultural Heritage Documentation

Kourosh Khoshelham, C. Chen, Martin Tomko
article en

Abstract

Cultural heritage (CH) reports contain rich descriptive content, but their rigid linear narrative structure necessitates the integration of disjointed parts of the text and auxiliary visual content (e.g., photographs) by the reader. To address this challenge, we explore the impact of the textual cues on cross-modal alignment performance. Recently, visual-language models (VLM) have been proposed for the automated establishment of text-image correspondence, but these have not been extensively evaluated in the context of domain knowledge (e.g., in CH). We conduct two experiments: (1) an image-to-text retrieval, analyzing how VLMs capture semantic overlap; and (2) a text-to-image visual grounding experiment evaluating how contextual clarity and language affect component-level detection accuracy. Our results show that current VLMs struggle to capture the fine-grained correspondences present in CH-related descriptive narratives due to semantic coarseness and limited relational understanding. Building on these insights, we adopt a grounding-based strategy that connects textual mentions and visual counterparts through prompt engineering and object-level anchoring that combines holistic and regional semantic alignment. Our study provides a structured evaluation of VLMs for multimodal integration in cultural heritage contexts, highlighting their limitations in fine-grained alignment. This contributes to a vision of interactive utilization of multimodal cultural heritage resources through improved connectivity between textual and visual content.

Journal on Computing and Cultural Heritage
The University of Melbourne (AU)
Sustainable cities and communities
Openalex Percentile: Top 13%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Linking Vision and Meaning: Text-Image Correspondence in Cultural Heritage Documentation — Kourosh Khoshelham, C. Chen, et al. · Journal on Computing and Cultural Heritage (2026) | TGRS Research Map | TGRS