Arlinde C.E. Vrooman · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23010929
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This working paper, shared for early dissemination and feedback, investigates the use of contextual Transformer models for semantic disambiguation in large historical corpora, using the ambiguous Early Modern Dutch form weder in the GLOBALISE corpus as a case study. Whereas the construction of stable historical vocabularies can benefit from type-level representations that support interpretable category formation, semantic disambiguation of individual occurrences requires representations that are sensitive to textual context. A manually annotated dataset of 245 occurrences of weder, comprising 67 WEATHER and 178 NON-WEATHER examples from 35 inventories, was constructed through targeted candidate retrieval. To reduce the risk of inventory-specific information leaking across evaluation sets, the data were divided at the inventory level into fixed training, development, and test sets. A pretrained GloBERTise model was fine-tuned as a binary sequence classifier using target-centred textual contexts. Comparison of ±150, ±300, and ±750 character contexts shows that increasing the amount of surrounding documentary context does not necessarily improve classification. The ±300-character model achieved a test macro F1 of 0.6722 and identified 7 of 10 WEATHER occurrences in the held-out test set. Qualitative and distributional error analysis suggests that the principal difficulty lies in distinguishing the semantic interpretation of the target occurrence from misleading weather-related vocabulary in its surrounding context. The results illustrate the potential of contextual BERT-based disambiguation for historical Dutch while also highlighting the challenges posed by lexical ambiguity and documentary variation.
No comments yet — start the discussion below.