Arabic WordNet Enrichment: A Comprehensive Survey
Progress in natural language processing (NLP) has primarily benefited high-resource languages. Although Arabic now has access to substantial corpora and widely used pretrained language models such as AraBERT and MARBERT, it still exhibits a significant imbalance, particularly in the availability of structured lexical–semantic resources. Arabic WordNet (AWN) remains a cornerstone of Arabic NLP; however, its lexical coverage and semantic richness are substantially lower than those of Princeton WordNet (PWN). Researchers have explored various approaches to enrich AWN, from lexicon-based methods relying on bilingual dictionaries and pattern extraction to embedding-based approaches built on Word2Vec and transformer architectures. This paper systematically reviews Arabic WordNet (AWN) enrichment approaches, categorizes existing methods, and shows that hybrid strategies are largely underexplored. We examine how Arabic’s key linguistic features—non‑concatenative morphology, broken plurals, diacritic ambiguity, and root‑based semantics—affect enrichment outcomes and assess their effects across different enrichment techniques. Finally, we propose a linguistically grounded hybrid agenda for future AWN development.
Authors
- Hassina Aliane (ORCID: https://orcid.org/0000-0003-3317-2741)
- Leila Nouri (ORCID: https://orcid.org/0009-0006-8007-1360)
Institutions
- Digital Science (United States) (US)
Publication Details
- Journal
- ACM Transactions on Asian and Low-Resource Language Information Processing
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1145/3856810
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00