Synthetic Data-Driven Transformer OCR for Kurdish Sorani via Dynamic Line Generation and Script-Aware Normalization
OCR for low-resource languages is still held back by the same small number of issues: too little labeled image-text data, too few benchmarks, and thin language-specific tooling. Kurdish Sorani is a particularly awkward case. It is written in a modified Arabic script, runs right to left, and has orthographic habits that standard Arabic OCR engines handle poorly. This paper describes a transformer OCR system for Sorani trained almost entirely on synthetic data, meaning line images rendered on the fly from a text corpus rather than manually transcribed scans. The pipeline has three parts: corpus-driven line synthesis, a deterministic script-aware normalization step based on character-level transliteration, and a TrOCR encoder–decoder recognizer. Text lines are rendered with randomly sampled fonts and sizes, then passed through stochastic augmentation to mimic realistic distortions. The system is evaluated twice. On an in-distribution synthetic set of 200 rendered lines, the best model reaches a character error rate of 0.0434, a word error rate of 0.1246, and 64.0% exact matches. More importantly, on a real-world test set of 19 scanned Kurdish documents (467 lines, 28,468 characters) processed end-to-end through detection and recognition, it reaches a character error rate of 0.0305 and a word error rate of 0.1770, beating both Arabic and Kurdish Tesseract baselines and an existing Kurdish TrOCR model while being considerably smaller than the latter. A controlled ablation, in which eight variants are trained under one shared budget and scored on identical images, then isolates what each design choice contributes. The label space is the largest design effect, and the reason is concrete: the decoder’s pre-trained tokenizer has no representation for seven common Sorani graphemes, which cover 14.7% of the corpus and place a floor under any model trained on native-script labels. Corpus size dominates overall and behaves as a threshold, font diversity helps with diminishing returns, and stochastic augmentation buys robustness at a small cost in in-distribution accuracy. Aligning detected lines against the transcribed ones further shows that line detection contributes under 1% of the reported character error on this material. The broader point, at least for Sorani, is that the synthetic training data and the label space in which the model predicts have to be designed together: a compact recognizer built that way outperforms a substantially larger released Kurdish model on genuine document images.
Authors
- Hawraz A. Ahmad
Institutions
- Salahaddin University-Erbil (IQ)
Publication Details
- Journal
- Algorithms
- Published
- 2026-09-16
- DOI
- https://doi.org/10.3390/a19090795
- Primary Topic
- Handwritten Text Recognition Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00