Jawhar: Optimized Morphological Analysis and Contextual Reranking for Arabic Part-of-Speech Tagging
Part-of-speech (POS) tagging in Arabic is hard because its rich root-and-pattern morphology and the absence of short vowels make one unvoweled string compatible with many categories. This paper presents Jawhar, a hybrid framework that couples a high-performance morphological analyser with contextual reranking using a pretrained Arabic language model. Jawhar is an autonomous engine inspired by Al-Khalil MorphoSys and rebuilt in Python that replaces the original XML databases with optimised JSON structures for faster inference. It enumerates the morphologically valid candidates of each token, and a CAMeL-BERT stage then scores each candidate by its full morphological signature (type, POS, root, pattern, and voweled form). On the Prague Arabic Dependency Treebank, mapped to the universal 17-tag POS scheme, the fine-tuned scorer reached 96.4% token accuracy (macro-F1 0.921), on par with published neural taggers, while a candidate-constrained hybrid attached a full morphological analysis to 65.9% of tokens at the same accuracy and reached a 97.4% oracle ceiling. A rule-based configuration reached 54.4%, and the zero-shot reranker reached parity (54.3%), which showed that within a fixed candidate set, reordering could not cross the coverage ceiling. The main contribution was a token-level decomposition of the error budget that isolated candidate coverage and label mapping from contextual ranking, released with a public analyser and harness.
Authors
- Abdelkaher Ait Abdelouahad (ORCID: https://orcid.org/0000-0001-8887-7840)
- Mohamed Bouzahir (ORCID: https://orcid.org/0000-0002-9415-3600)
- Mohamed Nabil
Institutions
- Chouaib Doukkali University (MA)
Publication Details
- Journal
- Information
- Published
- 2026-09-09
- DOI
- https://doi.org/10.3390/info17090874
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00