Muharir, Edi Noersasongko, Abdul Syukur, R. Rizal Isnanto, Deshinta Arrova Dewi, Muljono Muljono · International journal of intelligent engineering and systems 2026 · 2026
DOI: 10.22266/ijies2026.0930.15
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
For low-resource language pairs, Phrase-Based SMT (PBSMT) remains a viable alternative where zeroshot off-the-shelf neural approaches underperform due to data scarcity.This paper presents the first reproducible benchmark for Banjar-Indonesian machine translation-an under-resourced Austronesian language pair with approximately six million speakers-evaluated on a parallel corpus of 8,770 sentence pairs drawn from Banjar folklore ("Cerita Si Palui").Nine direction-specific experimental runs were evaluated, comprising five system variants (baseline, factored morphological model, and MERT/MIRA optimization under 3-gram and 5-gram language models) across two translation directions, with the factored model evaluated only for Indonesian→Banjar.The baseline achieves 52.30BLEU (Banjar→Indonesian) and 48.70 BLEU (Indonesian→Banjar).For Indonesian→Banjar, complexity consistently degraded performance: factored models yielded +0.01 BLEU (p=0.532,ns) while MIRA degraded by 2.31 BLEU (p<0.001) and 5-gram models by 3.09-3.43BLEU (both p<0.001).For Banjar→Indonesian, a 5-gram model with MERT improved performance by +0.40 BLEU (p=0.010),attributed to Indonesian's orthographic standardization.Manual error analysis (100 sentences) identifies wrong lexical choices (84%) as the dominant challenge.Results reveal a complexity threshold dependent on target language regularity.
No comments yet — start the discussion below.