GAO Yusi, WANG Siriguleng, SI Qintu · DOAJ (DOAJ: Directory of Open Access Journals) 2026 · 2026
DOI: 10.3778/j.issn.1673-9418.2509005
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
With the advancement of deep learning, neural machine translation models have achieved remarkable progress in translation tasks for high-resource language pairs and reached high levels on multiple automatic evaluation metrics. Low-resource machine translation, however, is often constrained by the scarcity of parallel corpora and insufficient linguistic research support, leading to a tendency for models to overfit and consequently making it difficult to improve translation quality. As an efficient model compression and knowledge transfer technique, knowledge distillation can transfer the knowledge of a teacher model to a student model, effectively enhancing the generalization and robustness of the student model in low-resource scenarios. This study presents a systematic review of knowledge distillation methods in the field of low-resource machine translation. Firstly, it elaborates on the concept and development history of knowledge distillation, sorts out its research progress thread from high-resource to low-resource scenarios, and introduces relevant evaluation metrics and datasets. Then, from three dimensions, distillation content, training methods, and model architectures, it analyzes in detail the principles, processes, advantages, disadvantages, and applicable scenarios of various knowledge distillation methods, and verifies their practical effectiveness through typical cases. Furthermore, it points the core challenges faced by low-resource machine translation in terms of?data scarcity, corpus noise, insufficient utilization?of monolingual resources and catastrophic forgetting, and systematically elaborates on the corresponding solutions of knowledge distillation technology to address the above challenges. Finally, it provides an outlook on future research directions, including robust distillation mechanisms, synergy between multi-modal and cross-domain distillation, self-distillation and teacher-free distillation, and joint optimization of data, models and distillation. This study provides a theoretical framework and development path for low-resource translation tasks, and offers a reference for the in-depth development of multilingual and multimodal intelligent translation research.
No comments yet — start the discussion below.