Aina Garí Soler · INRIA a CCSD electronic archive server 2026 · 2026
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Masked language models output probability distributions over subword tokens, making it non-trivial to obtain reliable word-level probabilities -particularly when comparing words of different tokenization lengths. We present a systematic study of word-level probability estimation from MLMs, comparing multiple methods for probability calculation in two languages, multiple model families, and five tokenization lengths. Our most important finding is that the masking strategy used in pretraining strongly determines probability quality: models trained without whole-word masking show severely distorted within-length distributions, being outperformed even by a trigram language model for multi-token words. Among the methods we evaluate, geometric mean aggregation is the most robust and practical choice.
No comments yet — start the discussion below.