Leila Nouri, Hassina Aliane · ACM Transactions on Asian and Low-Resource Language Information Processing 2026 · 2026
DOI: 10.1145/3856810
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Progress in natural language processing (NLP) has primarily benefited high-resource languages. Although Arabic now has access to substantial corpora and widely used pretrained language models such as AraBERT and MARBERT, it still exhibits a significant imbalance, particularly in the availability of structured lexical–semantic resources. Arabic WordNet (AWN) remains a cornerstone of Arabic NLP; however, its lexical coverage and semantic richness are substantially lower than those of Princeton WordNet (PWN). Researchers have explored various approaches to enrich AWN, from lexicon-based methods relying on bilingual dictionaries and pattern extraction to embedding-based approaches built on Word2Vec and transformer architectures. This paper systematically reviews Arabic WordNet (AWN) enrichment approaches, categorizes existing methods, and shows that hybrid strategies are largely underexplored. We examine how Arabic’s key linguistic features—non‑concatenative morphology, broken plurals, diacritic ambiguity, and root‑based semantics—affect enrichment outcomes and assess their effects across different enrichment techniques. Finally, we propose a linguistically grounded hybrid agenda for future AWN development.
No comments yet — start the discussion below.