Sijie Liu, Xuming Ye, Xingyu Wang, Lianshuai Wang, Weiyu Dong, Meng An · Electronics 2026 · 2026
DOI: 10.3390/electronics15194474
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multimodal knowledge extraction has become increasingly important for applications that require compact and informative representations of text-image data, yet existing models often remain computationally expensive for practical use. To address this issue, we propose MKE, a lightweight framework for multimodal knowledge extraction based on knowledge distillation and efficient multimodal fusion. Specifically, MKE transfers linguistic knowledge from a large BART teacher model to a compact StuBART backbone and combines Flash-Attention-based cross-modal interaction with a lightweight nonlinear transformation module. Experiments on the MSMO English news dataset show that MKE achieves competitive multimodal summarization performance with 270M parameters and a computational cost of 155.63 GMACs (311.27 GFLOPs) per sample. We further evaluate inference efficiency under a controlled RTX 4090 setup, and we clarify that deployment on smartphones, IoT systems, and wearable devices remains a promising direction for future work rather than a scenario directly validated in this study.
No comments yet — start the discussion below.