Yudhvir Singh Jogender Singh · Journal of Intelligent Decision Making and Information Science 2026 · 2026
DOI: 10.59543/jidmis.v3.2078
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Automatic Language Identification (LID) has become an essential research area in multilingual communication systems, speech analysis and intelligent human-computer interaction. The existing language identification techniques often suffer from limitations such as reduced recognition accuracy in noisy environments, ineffective extraction of contextual speech features, high computational complexity and poor generalization across multiple speech corpora. Conventional Machine Learning (ML) and Deep Learning (DL) methods frequently rely on handcrafted feature which increase processing time and memory requirements while failing to preserve long-range dependencies in speech sequences. This paper suggests an improved paradigm for language identification in order to overcome these limitations. Initially, the speech signals are pre-processed using a Butterworth Bandpass Filter (BBPF) to reduce environmental noise and convert the input into a clean digital representation. Linear Predictive Coding Cepstral Coefficients (LPCCs), Mel Frequency Cepstral Coefficients (MFCCs) and Perceptual Linear Prediction (PLP) coefficients are used for feature extraction. These techniques are used to capture complementary spectral, cepstral and perceptual speech features. It enables robust acoustic representation and improved speaker-independent discrimination. Then, the extracted features are embedded using the Recurrent Mask Bidirectional Attention-based Robustly Optimized BERT (RMB-RoBERTa) model to learn contextual and temporal relationships. The SMOTE is used to balance data, and the language prediction is performed through an Optimized Knowledge Distillation-based Axial Attention enclosed Depthwise Separable Long Short Network (OKD-ADSLSN) model with teacher and student models. It is optimized using the Random Tent Secant Optimization Algorithm (RT-SOA) to reduce the complexity of the model and inference cost. With a lower error rate, the suggested model obtained 99.68% accuracy, 99.68% precision, 99.68% recall, and 99.68% F1-score..
No comments yet — start the discussion below.