Arabic WordNet Enrichment: A Comprehensive Survey

Progress in natural language processing (NLP) has primarily benefited high-resource languages. Although Arabic now has access to substantial corpora and widely used pretrained language models such as AraBERT and MARBERT, it still exhibits a significant imbalance, particularly in the availability of structured lexical–semantic resources. Arabic WordNet (AWN) remains a cornerstone of Arabic NLP; however, its lexical coverage and semantic richness are substantially lower than those of Princeton WordNet (PWN). Researchers have explored various approaches to enrich AWN, from lexicon-based methods relying on bilingual dictionaries and pattern extraction to embedding-based approaches built on Word2Vec and transformer architectures. This paper systematically reviews Arabic WordNet (AWN) enrichment approaches, categorizes existing methods, and shows that hybrid strategies are largely underexplored. We examine how Arabic’s key linguistic features—non‑concatenative morphology, broken plurals, diacritic ambiguity, and root‑based semantics—affect enrichment outcomes and assess their effects across different enrichment techniques. Finally, we propose a linguistically grounded hybrid agenda for future AWN development.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Asian and Low-Resource Language Information Processing
Published
2026-10-06
DOI
https://doi.org/10.1145/3856810
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Arabic WordNet Enrichment: A Comprehensive Survey

Hassina Aliane, Leila Nouri
ACM Transactions on Asian and Low-Resource Language Information Processing
Natural Language Processing Techniques
article

Arabic WordNet Enrichment: A Comprehensive Survey

Hassina Aliane, Leila Nouri
article en

Abstract

Progress in natural language processing (NLP) has primarily benefited high-resource languages. Although Arabic now has access to substantial corpora and widely used pretrained language models such as AraBERT and MARBERT, it still exhibits a significant imbalance, particularly in the availability of structured lexical–semantic resources. Arabic WordNet (AWN) remains a cornerstone of Arabic NLP; however, its lexical coverage and semantic richness are substantially lower than those of Princeton WordNet (PWN). Researchers have explored various approaches to enrich AWN, from lexicon-based methods relying on bilingual dictionaries and pattern extraction to embedding-based approaches built on Word2Vec and transformer architectures. This paper systematically reviews Arabic WordNet (AWN) enrichment approaches, categorizes existing methods, and shows that hybrid strategies are largely underexplored. We examine how Arabic’s key linguistic features—non‑concatenative morphology, broken plurals, diacritic ambiguity, and root‑based semantics—affect enrichment outcomes and assess their effects across different enrichment techniques. Finally, we propose a linguistically grounded hybrid agenda for future AWN development.

ACM Transactions on Asian and Low-Resource Language Information Processing
Digital Science (United States) (US)
Openalex Percentile: Top 11%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Arabic WordNet Enrichment: A Comprehensive Survey — Hassina Aliane, Leila Nouri · ACM Transactions on Asian and Low-Resource Language Information Processing (2026) | TGRS Research Map | TGRS