Syed Muntasir Mamun · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22989906
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Yavuz, Meister and Pimentel (2026) isolate two axes that have long been confounded in the practice of subword tokenisation: the optimisation objective (compression versus unigram log-likelihood) and the search procedure (bottom-up merging versus top-down pruning). By completing a 2 × 2 design with two new algorithms—BottomUpLL and TopDownComp—they show that bits-per-byte (BPB) is governed more by search than by objective, while grammatical minimal-pair accuracy (BLiMP and multilingual analogues) does not inherit that ordering. This review reconstructs the paper anatomically: definitions, lemmas, approximations, experimental skeleton, intrinsic and extrinsic findings, and declared limitations. It then situates the contribution against a genealogy of at least twelve competing models of vocabulary construction, from classical BPE and WordPiece through UnigramLM, PathPiece, S-BPE, SentencePiece, morphological analysers, convex-relaxation tokenisers, and vocabulary-refinement heuristics. Ontologically, the paper forces a choice between tokens as compressional atoms and tokens as likelihood atoms, and between vocabularies grown by fusion and vocabularies carved by deletion. Epistemologically, it replaces folk comparison of named algorithms with a factorial experiment, while still relying on a unigram proxy, a local-replacement deletion score, and a modest model scale. Also available on Academia.edu: https://www.academia.edu/attachments/134371150/download_file?s=portfolio
No comments yet — start the discussion below.