Alexandre Nikolaev, Yu‐Ying Chuang, R. Harald Baayen · Morphology 2026 · 2026
DOI: 10.1007/s11525-026-09470-9
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study shows that the Discriminative Lexicon Model (DLM) can learn Finnish nominal inflection from paired form and meaning embeddings, without being given access to stems, exponents, inflectional features, and information about inflectional class. It also shows that the model is productive. It generalizes to novel words, and does so with greater precision for more productive inflectional classes. This does not imply that productivity reduces simply to the absence of irregularity. Rather, generalization is made possible thanks to substantial isomorphies that exist between the space of form embeddings and the space of meaning embeddings. Usage-based DLM models that take token frequency into account have lower type accuracy: practice makes perfect, but with less practice, learning becomes increasingly problematic. Nevertheless, token-based DLMs perform well token-wise; they generalize to novel words, and reflect the differences in productivity of inflectional classes even better than type-based models. It is also shown that DLM comprehension models with form embeddings based on 4-grams outperform models with form embeddings based on 3-grams. This is due to 4-grams better covering the partial combinatorics of exponents and corresponding to more informative regions in the semantic embedding space.
No comments yet — start the discussion below.