Utkirbek Azimov · Interpretation and researches 2026 · 2026
DOI: 10.5281/zenodo.23085907
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
We benchmark Uzbek NER on a frozen Uzbek-NER-Gold test split (400 sentences; 5,712 tokens; 870 entities; 8 types; seed 7), with LLMs on a fixed 200-sentence subset. In-domain XLM-R-large fine-tuning (3,776 training sentences; five seeds) reaches 80.9% mean strict entity F1 on the subset (78.5–82.3) and 79.2% on all 400 (77.7–80.2), the best measured accuracy. An in-domain CRF reaches 77.1% on the subset [72.2, 81.6], trains in 11.4 s on CPU, and infers on 400 sentences in 0.04 s. Smaller Uzbek-pretrained encoders score 66.0% and 69.8%; transfer, zero-shot, and few-shot methods score 39.6–62.1%. Memorization scores 52.2%, spaCy 16.9%, and 350M few-shot models about 3%. We release the verified XLM-R checkpoint UAzimov/uzbek-ner-xlmr-large (0.7927 full / 0.8094 subset)on Hugginface.
No comments yet — start the discussion below.