Multi-modal knowledge graph reasoning and applications: A comprehensive survey
With the rapid growth of multi-modal data, multi-modal knowledge graphs (MMKGs) have emerged as an important research focus in intelligent reasoning by integrating structured knowledge with information from diverse modalities, such as text, images, and audio. This paper provides a comprehensive review of recent advances in multi-modal knowledge graph reasoning (MMKGR), systematically summarizing existing methods and applications while identifying key challenges and research opportunities. We first introduce the definitions, major tasks, and application scenarios, and then develop a unified taxonomy of existing methods. Specifically, according to how graph structures and multi-modal evidence are organized and processed during reasoning, MMKGR methods are broadly categorized into graph-centric approaches, including embedding-based and path-based methods, and sequence-centric approaches, including Transformer-based and LLM-based methods. We further review the major applications of MMKGR, including multi-modal question answering, multi-modal recommendation, and cross-modal information retrieval, and organize them into static graph enhancement and dynamic graph construction approaches across general and domain-specific scenarios. Finally, we identify the major challenges in MMKGR and present a staged roadmap for its short-, medium-, and long-term development.
Authors
- Zongmin Ma (ORCID: https://orcid.org/0000-0001-7780-6473)
- Ruizhe Ma (ORCID: https://orcid.org/0000-0003-2749-3063)
- Li Yan
- Bao Wang
- Xinyu Liang
Institutions
- University of Massachusetts Lowell (US)
- Nanjing University of Aeronautics and Astronautics (CN)
Publication Details
- Journal
- Engineering Applications of Artificial Intelligence
- Published
- 2026-09-10
- DOI
- https://doi.org/10.1016/j.engappai.2026.116142
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00