Aman Mussa, Zhanseit Tuimebayev, Мадина Мансурова, Ahsan Habib Shihab · Information 2026 · 2026
DOI: 10.3390/info17090921
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Fine-tuning is often treated as a prerequisite for deploying large language models in low-resource languages, yet its value for recent open-weight models is unclear. We evaluate eleven checkpoints of 7–14 billion parameters before and after Low-Rank Adaptation on 5000 human-curated Kazakh instruction–response pairs. Evaluation is performed on the Kazakh splits of KazMMLU, ARC, GPQA, MMLU-Pro, and GSM8K, with Russian GSM8K as a higher-resource comparison. Our central finding is a language asymmetry: the same Kazakh-trained adapters improved Russian GSM8K more than the Kazakh GSM8K they targeted. Excluding one checkpoint that collapsed after fine-tuning by up to 58 percentage points, the remaining ten gained +2.6 percentage points on average on Kazakh (range −6.1 to +13.6) against +6.9 on Russian, significant under a paired signed-rank test. A manual audit of 400 responses identifies a contributing mechanism: Kazakh outputs were more often truncated before a final answer, and correct Kazakh reasoning more often scored wrong by answer extraction, so part of the gap reflects response form, not reasoning. Adaptation produced no meaningful average change on the four multiple-choice benchmarks. GSM8K is the suite’s only free-form task, so generality is untested. These results suggest that low-resource fine-tuning should be evaluated per checkpoint, benchmark, and language rather than being adopted by default.
No comments yet — start the discussion below.