R. Thamarai Selvi · Dandao Xuebao/Journal of Ballistics 2026 · 2026
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multilingual speakers frequently mix English with regional languages during communication resulting in code mixed speech. In South India, People are using English and the Dravidian languages such as Tamil, Telugu, Kannada and Malayalam. The problem of identifying the languages based on the audio streams is not yet solved because of the mistakes of the automatic speech recognition (ASR), phonetic drift, and the variability of the Romanized transcription. The existing LID methods are mostly based on the trained supervised machine learning models or transformer-based networks that are trained on textual corpora that are clean, which generally deteriorate in the case of noisy ASR settings. This paper presents a ASR-conscious language identification system tailored towards Romanized Dravidian-English video/audio code-mixed speech. The system uses a hierarchical decision process comprising of phonetic normalization by rules, exact matching by a dictionary, conservative fuzzy matching, and disambiguation by context. It has a full automatic correction module that provides the means to correct ASR distortions and does not require manual transcript correction. This empirical analysis proves that the proposed model outperforms the models SVM, FastText, mBERT and XLM-R by attaining the accuracy 90.09% for language identification of code mixed dataset which contains English, Tamil, Malayalam, Kannada and Telugu languages.
No comments yet — start the discussion below.