SIHG-Rec: Unleashing the Power of Semantic and Interactive Homogeneous Graphs via Dual-Stage Fusion for Multimodal Recommendation

Recent studies in multimodal recommendation, which leverage diverse modal information to address data sparsity and enhance recommendation accuracy, have garnered significant interest. Two critical processes in this domain are modality fusion and representation learning. In representation learning, existing studies often extend beyond traditional heterogeneous graphs by constructing homogeneous graphs that incorporate multimodal information, aiming to capture finer-grained user preferences and item properties. However, these studies predominantly rely on semantic homogeneous graphs, overlooking interactive signals, which limits their capacity to capture high-quality user-item relationships. Moreover, prevailing modality fusion strategies are typically single-stage, employing either early or late fusion with simplistic attention or predefined strategies. Early fusion often impedes the learning of modality-specific features, while late fusion fails to sufficiently capture inter-modal relationships. In previous studies, modality fusion and representation learning have been treated as independent processes. In this paper, we argue that these two processes are complementary and can mutually enhance each other: powerful representation learning strengthens modality fusion, while effective fusion improves representation quality. To this end, we propose SIHG-Rec, a novel framework that jointly incorporates semantic information and historical interactions to construct S emantic and I nteractive H omogeneous G raphs, leading to more accurate and robust representations of user preferences and item properties. To fully exploit these graphs, we introduce a dual-stage fusion strategy that first refines all modalities using behavior-guided calibration at an early stage, then fuses their representations at a later stage. This design preserves modality-specific features while reducing task-irrelevant information through early-stage calibration. Furthermore, we adopt an adaptive optimization to ensure balanced and stable representations across modalities. Extensive experiments on three widely used datasets show that SIHG-Rec consistently outperforms a variety of strong baselines, especially in cold-start settings. Additional analyses confirm its superiority under varying levels of data sparsity.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Recommender Systems
Published
2026-09-14
DOI
https://doi.org/10.1145/3847662
Primary Topic
Recommender Systems and Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SIHG-Rec: Unleashing the Power of Semantic and Interactive Homogeneous Graphs via Dual-Stage Fusion for Multimodal Recommendation

Xiping Hu, Jinfeng Xu, Edith C.‐H. Ngai, Wei Wang et al.
ACM Transactions on Recommender Systems
Recommender Systems and Techniques
article

SIHG-Rec: Unleashing the Power of Semantic and Interactive Homogeneous Graphs via Dual-Stage Fusion for Multimodal Recommendation

Xiping Hu, Jinfeng Xu, Edith C.‐H. Ngai, Wei Wang, Sang‐Wook Kim, Zheyu Chen
article en

Abstract

Recent studies in multimodal recommendation, which leverage diverse modal information to address data sparsity and enhance recommendation accuracy, have garnered significant interest. Two critical processes in this domain are modality fusion and representation learning. In representation learning, existing studies often extend beyond traditional heterogeneous graphs by constructing homogeneous graphs that incorporate multimodal information, aiming to capture finer-grained user preferences and item properties. However, these studies predominantly rely on semantic homogeneous graphs, overlooking interactive signals, which limits their capacity to capture high-quality user-item relationships. Moreover, prevailing modality fusion strategies are typically single-stage, employing either early or late fusion with simplistic attention or predefined strategies. Early fusion often impedes the learning of modality-specific features, while late fusion fails to sufficiently capture inter-modal relationships. In previous studies, modality fusion and representation learning have been treated as independent processes. In this paper, we argue that these two processes are complementary and can mutually enhance each other: powerful representation learning strengthens modality fusion, while effective fusion improves representation quality. To this end, we propose SIHG-Rec, a novel framework that jointly incorporates semantic information and historical interactions to construct S emantic and I nteractive H omogeneous G raphs, leading to more accurate and robust representations of user preferences and item properties. To fully exploit these graphs, we introduce a dual-stage fusion strategy that first refines all modalities using behavior-guided calibration at an early stage, then fuses their representations at a later stage. This design preserves modality-specific features while reducing task-irrelevant information through early-stage calibration. Furthermore, we adopt an adaptive optimization to ensure balanced and stable representations across modalities. Extensive experiments on three widely used datasets show that SIHG-Rec consistently outperforms a variety of strong baselines, especially in cold-start settings. Additional analyses confirm its superiority under varying levels of data sparsity.

ACM Transactions on Recommender Systems
Beijing Institute of Technology (CN), Hanyang University (KR), Macao Polytechnic University (MO), University of Hong Kong (HK)
Openalex Percentile: Top 4%
Recommender Systems and Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.