O. Ibrahimzade · arXiv (Cornell University) 2026 · 2026
DOI: 10.5281/zenodo.22742915
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Nearly every Turkic language sits beside a better-resourced neighbour that shares its script: Azerbaijani beside Turkish, Kazakh beside Russian, Turkmen and Gagauz beside Turkish. A multilingual speech recogniser that has seen far more of the neighbour will often transcribe the target language as the neighbour. The output is fluent and confident, and it reads as a working transcript to anyone who does not speak the target language. Neither standard metric exposes this failure: word error rate (WER) penalises it but says nothing about its cause, and character error rate (CER) actively hides it, because the phonemes are approximately right. In our evaluation, a system writing Turkish orthography for Azerbaijani scored 91.4% WER against 60.7% CER, a gap that reads as "difficult audio" rather than "wrong language." We propose the Language-Specific Grapheme Recall (LSGR), a training-free diagnostic derived mechanically from two alphabets: the set G_L of graphemes present in the target orthography but absent from the neighbour's, scored as the F1 of pooled precision and recall over reference–hypothesis pairs. Across nine systems evaluated on identical Azerbaijani telephone audio, LSGR is graded rather than binary (0.690–0.977 among systems that write the target orthography), is not recoverable from WER within that group, where two systems 1.7 WER points apart differ by 0.152 in LSGR, and drops to 0.000 on telephone speech for two 2026 systems that do not list Azerbaijani, whose CER alone would not have stopped a practitioner from deploying them. We derive G_L for eleven language pairs, show that for five Turkic pairs the neighbour's alphabet is a strict subset of the target's, which is exactly the condition under which one-sided recall suffices, and measure signal density on FLEURS to show that every Turkic pair tested is scorable on at least 88% of utterances while a Catalan/Spanish control is scorable on only 18%. We state the method's central limit plainly: LSGR measures orthographic conformance, not language understanding, and scores 0.000 equally for the wrong language and for the right language in the wrong script. The paper is the empirical companion to our theoretical framework for Turkic cross-lingual transfer, whose script-compatibility term, we argue, is double-edged: shared script raises transfer potential and simultaneously lowers the detectability of transfer failure.
No comments yet — start the discussion below.