Omar Dawoud · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23106484
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Dependency parsers often receive morphological features as part of their input, but the source and completeness of those features are rarely treated as experimental variables. We study German Case in a controlled setting by comparing dependency parsers trained with no Case, independently predicted Case, and gold Case. Using pinned German-GSD and German-HDT Universal Dependencies resources, we audit Case coverage and provenance, construct matched feature conditions, and evaluate both in-domain parsing and cross-treebank transfer. In a three-seed, 1,000-step T4 pilot, predicted Case has higher mean labeled attachment than no Case and higher overall mean LAS than gold Case under HDT transfer for all three seeds. The difference is concentrated on tokens where gold Case is missing; on gold-Case-present tokens, predicted Case is 0.47 LAS points below gold Case on average. Sentence-level paired bootstrap intervals for the aggregate HDT comparison exclude zero, while the in-domain predicted-versus-gold comparison remains seed-sensitive. We use these results to study feature source and annotation completeness while treating the parser and training setup as a controlled feasibility study.
No comments yet — start the discussion below.