Shunxin Yao · · 2026
DOI: 10.31235/osf.io/tdg9h_v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Quantitative comparisons across languages have traditionally relied on manually curated typological features, while recent advances in multilingual language models provide new opportunities for large-scale cross-linguistic research. This paper presents a reproducible framework that integrates natural language processing and statistical modeling to examine lexical, syntactic, and semantic variation across English, Chinese, Spanish, French, and German. Drawing on a corpus of 1,860 Wikipedia articles, the study combines traditional linguistic indicators with LaBSE-based semantic representations. The results show that language differences dominate the observed variation, with Chinese forming a distinct outlier and English and Spanish exhibiting the greatest similarity. Moreover, lexical-syntactic distances are significantly correlated with semantic distances derived from Transformer embeddings. These findings suggest that traditional quantitative measures and modern representation-learning approaches capture complementary aspects of language structure and meaning. The framework offers a scalable methodology for multilingual corpus linguistics and quantitative typology.
No comments yet — start the discussion below.