M. Ramkumar, M. Marimuthu, R. Lakshminarayanan, V Sumanth · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-65518-0
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Cross-lingual speech recognition for real-time applications is constrained by the joint requirements of low latency, high recognition accuracy, and strict privacy over conversational audio. This research work addresses these constraints by designing an adaptive dual-stream neural encoder that explicitly models acoustic and linguistic information in parallel while operating in a streaming regime. The proposed architecture was developed with two coordinated streams: an acoustic encoder that processed log-Mel features in fixed-size chunks and a linguistic encoder that consumed partial token embeddings, with an adaptive gating module that fused both representations to minimize re-computation and buffering delay. A privacy-preserving data augmentation pipeline was further introduced, in which non-identifying signal transforms, feature-space perturbations, and cross-lingual masking were applied without storing raw sensitive speech or speaker metadata. The framework was trained and evaluated on multilingual corpora including CoVoST 2, Common Voice, and MuST-C, achieving relative word error rate reductions of approximately 19% on average and up to 22% in low-resource languages compared with a strong streaming Transformer baseline. Under streaming conditions, the dual-stream encoder attained an operating Real-Time Factor of 0.60 with end-to-end latency below 320 ms for typical utterances, while maintaining character error rates below 7% across the majority of evaluated languages. Ablation studies indicated that the adaptive gating contributed a relative WER improvement of about 20% and that privacy-preserving augmentation yielded a further gain of about 14% in noisy and domain-shifted test scenarios. This research work therefore establishes a practical pathway toward low-latency, privacy-aware, cross-lingual speech-to-text systems suitable for deployment in real-time multilingual services
No comments yet — start the discussion below.
