MKE: A Lightweight Framework for Multimodal Knowledge Extraction via Knowledge Distillation

Multimodal knowledge extraction has become increasingly important for applications that require compact and informative representations of text-image data, yet existing models often remain computationally expensive for practical use. To address this issue, we propose MKE, a lightweight framework for multimodal knowledge extraction based on knowledge distillation and efficient multimodal fusion. Specifically, MKE transfers linguistic knowledge from a large BART teacher model to a compact StuBART backbone and combines Flash-Attention-based cross-modal interaction with a lightweight nonlinear transformation module. Experiments on the MSMO English news dataset show that MKE achieves competitive multimodal summarization performance with 270M parameters and a computational cost of 155.63 GMACs (311.27 GFLOPs) per sample. We further evaluate inference efficiency under a controlled RTX 4090 setup, and we clarify that deployment on smartphones, IoT systems, and wearable devices remains a promising direction for future work rather than a scenario directly validated in this study.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-29
DOI
https://doi.org/10.3390/electronics15194474
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

MKE: A Lightweight Framework for Multimodal Knowledge Extraction via Knowledge Distillation

Xingyu Wang, Sijie Liu, Weiyu Dong, Lianshuai Wang et al.
Electronics
Multimodal Machine Learning Applications
article

MKE: A Lightweight Framework for Multimodal Knowledge Extraction via Knowledge Distillation

Xingyu Wang, Sijie Liu, Weiyu Dong, Lianshuai Wang, Meng An, Xuming Ye
article en

Abstract

Multimodal knowledge extraction has become increasingly important for applications that require compact and informative representations of text-image data, yet existing models often remain computationally expensive for practical use. To address this issue, we propose MKE, a lightweight framework for multimodal knowledge extraction based on knowledge distillation and efficient multimodal fusion. Specifically, MKE transfers linguistic knowledge from a large BART teacher model to a compact StuBART backbone and combines Flash-Attention-based cross-modal interaction with a lightweight nonlinear transformation module. Experiments on the MSMO English news dataset show that MKE achieves competitive multimodal summarization performance with 270M parameters and a computational cost of 155.63 GMACs (311.27 GFLOPs) per sample. We further evaluate inference efficiency under a controlled RTX 4090 setup, and we clarify that deployment on smartphones, IoT systems, and wearable devices remains a promising direction for future work rather than a scenario directly validated in this study.

ElectronicsVol. 15(19)
Minzu University of China (CN), Chinese Academy of Sciences (CN), Aerospace Information Research Institute (CN), Shanghai Artificial Intelligence Laboratory (CN)
Quality Education
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

MKE: A Lightweight Framework for Multimodal Knowledge Extraction via Knowledge Distillation — Xingyu Wang, Sijie Liu, et al. · Electronics (2026) | TGRS Research Map | TGRS