MKE: A Lightweight Framework for Multimodal Knowledge Extraction via Knowledge Distillation
Multimodal knowledge extraction has become increasingly important for applications that require compact and informative representations of text-image data, yet existing models often remain computationally expensive for practical use. To address this issue, we propose MKE, a lightweight framework for multimodal knowledge extraction based on knowledge distillation and efficient multimodal fusion. Specifically, MKE transfers linguistic knowledge from a large BART teacher model to a compact StuBART backbone and combines Flash-Attention-based cross-modal interaction with a lightweight nonlinear transformation module. Experiments on the MSMO English news dataset show that MKE achieves competitive multimodal summarization performance with 270M parameters and a computational cost of 155.63 GMACs (311.27 GFLOPs) per sample. We further evaluate inference efficiency under a controlled RTX 4090 setup, and we clarify that deployment on smartphones, IoT systems, and wearable devices remains a promising direction for future work rather than a scenario directly validated in this study.
Authors
- Xingyu Wang (ORCID: https://orcid.org/0009-0001-3822-3256)
- Sijie Liu (ORCID: https://orcid.org/0009-0000-7405-0121)
- Weiyu Dong
- Lianshuai Wang
- Meng An
- Xuming Ye
Institutions
- Minzu University of China (CN)
- Chinese Academy of Sciences (CN)
- Aerospace Information Research Institute (CN)
- Shanghai Artificial Intelligence Laboratory (CN)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-29
- DOI
- https://doi.org/10.3390/electronics15194474
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00