Dr. Ali Siddiqui, Shauban Ali Solangi, Uzma Ehsan · Journal of Social Research Development 2026 · 2026
DOI: 10.59075/jsrd.v7i8.547
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models (LLMs) do not read words; they read tokens, and the tokenizer decides how many tokens a language needs to say the same thing. This paper introduces the notion of a Sindhi token tax and tests it empirically for Pakistan’s English–Sindhi context. Using 1,004 parallel sentences from the FLORES-200 benchmark in English, Sindhi and Urdu, the study measured five widely deployed tokenizers (GPT-2, Llama 2, Llama 3, Qwen2 and Command-R) with a reproducible Python pipeline. A convergent mixed-methods design combined quantitative indices with a rule-assisted qualitative analysis of how frequent Sindhi words are segmented. Quantitatively, Sindhi required 2.87 to 4.98 times as many tokens as English for the same content in total (Wilcoxon signed-rank, r = .87, p < .001), and a 4,096-token context window held about 770 to 1,335 Sindhi words against 2,890 to 3,330 English words. A new Orthographic Penalty Index showed that the 18 letters unique to the Sindhi alphabet, such as ڪ, ٿ, ڻ and ڳ, were split into byte fragments in 100% of their occurrences in four of the five tokenizers, while letters shared with Arabic and Urdu were almost never split. Qualitatively, five patterns emerged, including the fragmentation of core grammatical words, the “Urdu shortcut” by which Urdu code points are cheaper than Sindhi ones, and the invisibility of Sindhi morphology to subword boundaries. The study also found Unicode inconsistency in digital Sindhi, with 238 of 1,004 benchmark sentences mixing forms of the letter heh. The paper proposes the Script-Aware Tokenization Equity Framework (SATEF) and argues that script-level design, not only data scarcity, makes AI costlier and weaker for Sindhi speakers, with implications for language policy, education and technology in Pakistan.
No comments yet — start the discussion below.