Enhancing Arabic Morphological Segmentation with Deep Recurrent Neural Architectures
Arabic morphological segmentation is a fundamental preprocessing step for many natural language processing tasks in Arabic. However, meaningful comparison across existing segmentation systems remains challenging, as prior work often relies on different preprocessing pipelines, segmentation conventions, output representations, and evaluation protocols. These inconsistencies make it difficult to interpret reported results and to assess progress in a consistent and reproducible manner. To address this limitation, we introduce a unified formulation of Arabic morphological segmentation as a character-level binary boundary prediction task. In this formulation, each character is labeled according to whether a morpheme boundary follows it, resulting in a simple and standardized binary (0/1) representation. This representation is model-agnostic, does not rely on external lexical resources, and provides a common evaluation space that enables direct and consistent comparison across heterogeneous segmentation systems. Within this unified framework, we conduct a controlled empirical study of bidirectional recurrent architectures, including RNN, LSTM, and GRU models with one to four stacked layers, evaluated under identical preprocessing, training, and scoring conditions on ATBv3. The results show that gated recurrent models consistently outperform standard RNNs. In particular, a four-layer bidirectional GRU achieves the best performance, reaching 98.33% F1, 97.69% precision, and 98.98% recall over 249,745 boundary decisions. Under the same evaluation setting, this model also outperforms Farasa and a CAMeLBERT-based baseline in F1 score while maintaining a substantially lower inference time. Overall, the proposed boundary-based formulation provides a reproducible and controlled evaluation framework for Arabic morphological segmentation. More broadly, it establishes a representation-invariant evaluation setting for consistent benchmarking, supporting more reliable comparison across systems and offering a practical foundation for future research in Arabic and other morphologically rich languages.
Authors
- Yasser Hifny (ORCID: https://orcid.org/0000-0002-5516-7225)
- Hossam Shamardan (ORCID: https://orcid.org/0000-0002-8161-1983)
- Naglaa Abdelrehem (ORCID: https://orcid.org/0009-0005-5840-1485)
Institutions
- October 6 University (EG)
- Helwan University (EG)
Publication Details
- Journal
- ACM Transactions on Asian and Low-Resource Language Information Processing
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1145/3846165
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00