MOSAIC: A Multimodal Semantic-Oriented Alignment with Integrated Contrastive Learning for Multimodal Knowledge Graph Construction from Scientific Documents

Multimodal Knowledge Graphs (MMKGs) offer a promising paradigm for integrating heterogeneous sources into a unified, queryable, semantically structured representation. However, existing MMKG construction pipelines remain predominantly text-centric, extracting information from textual passages while leaving much of the visual and structural knowledge in scientific papers unrepresented. This produces fragmented graphs with disconnected components, isolated singleton nodes, and weak cross-modal connectivity, reducing the effectiveness of retrieval-augmented generation (RAG) systems that rely on interconnected graph traversal for multi-hop reasoning and evidence aggregation. We propose Multimodal Semantic-Oriented Alignment with Integrated Contrastive Learning (MOSAIC), an end-to-end framework for constructing semantically coherent MMKGs from heterogeneous scientific documents. MOSAIC combines four complementary contributions: (i) adaptive clustering via HDBSCAN, which infers data-driven entity boundaries without the brittle ε hyperparameters of conventional DBSCAN; (ii) a confidence-aware cross-modal alignment mechanism applying cosine-similarity gating to selectively invoke Large Language Models (LLMs), reducing spurious alignments and computational overhead; (iii) post-fusion semantic bridging, which links semantically related but structurally disconnected components through cosine-similarity-weighted bridge edges; and (iv) self-supervised contrastive embedding fine-tuning using an InfoNCE-style MultipleNegativesRankingLoss objective to specialise the embedding space for cross-modal entity representations. Empirical evaluation shows MOSAIC substantially improves graph topology and structural coherence, achieving a 110% increase in average clustering coefficient, reducing fragmentation, and strengthening intra-cluster semantic consistency. Evaluation across two challenging benchmark datasets, MMLongBench-Doc (134 documents, 1082 QA pairs) and DocBench (166 documents), shows that the MOSAIC-RAG engine significantly outperforms competing graph-based and dense retrieval systems, and that its advantage over lexical retrieval is concentrated in visually grounded questions, where it is statistically significant on both benchmarks. These results establish the effectiveness of our proposed MOSAIC framework for multimodal document understanding and its generalisability across diverse document categories.

Authors

Institutions

Publication Details

Journal
Big Data and Cognitive Computing
Published
2026-09-25
DOI
https://doi.org/10.3390/bdcc10100324
Primary Topic
Advanced Graph Neural Networks
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

MOSAIC: A Multimodal Semantic-Oriented Alignment with Integrated Contrastive Learning for Multimodal Knowledge Graph Construction from Scientific Documents

Jean Vincent Fonou-Dombeu, Busisani Mac Dube
Big Data and Cognitive Computing
Advanced Graph Neural Networks
article

MOSAIC: A Multimodal Semantic-Oriented Alignment with Integrated Contrastive Learning for Multimodal Knowledge Graph Construction from Scientific Documents

Jean Vincent Fonou-Dombeu, Busisani Mac Dube
article en

Abstract

Multimodal Knowledge Graphs (MMKGs) offer a promising paradigm for integrating heterogeneous sources into a unified, queryable, semantically structured representation. However, existing MMKG construction pipelines remain predominantly text-centric, extracting information from textual passages while leaving much of the visual and structural knowledge in scientific papers unrepresented. This produces fragmented graphs with disconnected components, isolated singleton nodes, and weak cross-modal connectivity, reducing the effectiveness of retrieval-augmented generation (RAG) systems that rely on interconnected graph traversal for multi-hop reasoning and evidence aggregation. We propose Multimodal Semantic-Oriented Alignment with Integrated Contrastive Learning (MOSAIC), an end-to-end framework for constructing semantically coherent MMKGs from heterogeneous scientific documents. MOSAIC combines four complementary contributions: (i) adaptive clustering via HDBSCAN, which infers data-driven entity boundaries without the brittle ε hyperparameters of conventional DBSCAN; (ii) a confidence-aware cross-modal alignment mechanism applying cosine-similarity gating to selectively invoke Large Language Models (LLMs), reducing spurious alignments and computational overhead; (iii) post-fusion semantic bridging, which links semantically related but structurally disconnected components through cosine-similarity-weighted bridge edges; and (iv) self-supervised contrastive embedding fine-tuning using an InfoNCE-style MultipleNegativesRankingLoss objective to specialise the embedding space for cross-modal entity representations. Empirical evaluation shows MOSAIC substantially improves graph topology and structural coherence, achieving a 110% increase in average clustering coefficient, reducing fragmentation, and strengthening intra-cluster semantic consistency. Evaluation across two challenging benchmark datasets, MMLongBench-Doc (134 documents, 1082 QA pairs) and DocBench (166 documents), shows that the MOSAIC-RAG engine significantly outperforms competing graph-based and dense retrieval systems, and that its advantage over lexical retrieval is concentrated in visually grounded questions, where it is statistically significant on both benchmarks. These results establish the effectiveness of our proposed MOSAIC framework for multimodal document understanding and its generalisability across diverse document categories.

Big Data and Cognitive ComputingVol. 10(10)
Umkhuseli Innovation and Research Management (ZA), University of KwaZulu-Natal (ZA)
Quality Education
Openalex Percentile: Top 9%
Advanced Graph Neural Networks
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.