Agung Santosa, Asril Jarin, Lyla Ruslana Aini, Eko Mulyanto Yuniarno, Hammam Riza, Mauridhi Hery Purnomo · International journal of intelligent engineering and systems 2026 · 2026
DOI: 10.22266/ijies2026.0930.12
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Javanese and Sundanese, spoken by over 110 million people, remain poorly served by automatic speech recognition.Zero-shot Whisper large-v3 attains word error rates of 58.16% and 61.64% on read-speech recordings of these languages-high enough to limit practical utility-yet supervised fine-tuning is infeasible where transcribed speech is unavailable.This work fuses a frozen Whisper large-v3 with MambaByte, a 972-million-parameter bytelevel selective state-space language model adapted to each language by LoRA continual pre-training on freely available text alone (235.5 MB Javanese; 120.9 MB Sundanese).A prefix-trie cache reduces each byte-level hypothesis extension to a constant-time state update, making the system tractable on a single GPU.With decoding hyperparameters selected on a small labelled set and no labelled speech used to adapt the language model, GFD reduces WER to 51.87% for Javanese (a statistically significant 6.29 percentage-point reduction) and 54.72% for Sundanese (a statistically significant 6.92 percentage-point reduction).
No comments yet — start the discussion below.