Sharath Sathish, Harikishore Tadigotla · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202608.2220.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Under a fixed pretraining budget, which inductive biases actually improve a small languagemodel, and which only appear to until they are tested against a matched control? Prabhasa-BabyLManswers this for Paninian morphosyntax through a controlled study on the BabyLM 2026 benchmark.Three mechanisms are tested: morpheme-boundary N-hot embeddings, karaka role-stratified masking,and a supervised role-prediction objective. Because Paninian parsers exist for Sanskrit and not for English, the deterministic mapping from Universal Dependencies relations to karaka roles thatsupplies every English label in this work is specified in full, situated against semantic role labelling, and measured for what it loses: only 68.3% of the subword positions the model sees carry the roleits own parse assigned. A 114M-parameter bidirectional encoder is trained on the Strict (100M-word) and Strict-Small (10M-word) budgets and every claim is tested with three to five seeds, pairedbootstrap inference, and Holm–Bonferroni correction. The Strict entry scores 74.56 BLiMP againsta GPT-2 baseline of 74.53 and stands third on the Strict track by Overall Average (45.21) at the timeof writing. The gains are attributable to the training objective, the architecture, the training budget,and completeness of evaluation coverage, not to the Paninian mechanisms. Pure masked languagemodelling is indistinguishable from a hybrid MLM+CLM objective at 10M words but exceeds it by5.49 points at 100M. At matched mask budget, karaka-stratified masking is causally inert, and the karaka auxiliary objective yields only +0.76 points over five seeds. The result is a controlled separationof the design choices that help at this scale from a linguistically motivated prior that does not.
No comments yet — start the discussion below.