Does Script Matter? Evaluating Large Language Models for Urdu Sentiment Analysis across Nastaliq, Roman Urdu, and Code-Switched Text
Urdu is spoken by more than 230 million people, but very little of the Urdu that people actually write online is in the language's formal Nastaliq script. Most of it is either Roman Urdu, which is Urdu typed in Latin letters, or a mixture of Urdu and English. Almost all existing benchmarks for Urdu, however, are built only on Nastaliq text, so it is not clear how well large language models handle the way people really write. In this paper we test six instruction-tuned models, three from the Llama family, two Gemma models, and Gemini Flash-Lite, on binary sentiment classification in all three registers, using 250 balanced examples per register and 4,500 predictions in total. Every model did best on Roman Urdu (mean accuracy 83.9%) and worst on code-switched text (75.8%), with Nastaliq in between (77.5%). We then went through 63 of the misclassifications by hand and found that almost half of them were not model mistakes at all, but wrong labels in the original datasets. Once these are taken into account, the true accuracy of the models is closer to 89–91% in every register. Our results suggest that Nastaliq-only benchmarks understate how useful these models are for ordinary Urdu users, and that the quality of Urdu sentiment datasets is at least as much of a bottleneck as the models themselves.
Authors
- Yousaf Khawar Raja (ORCID: https://orcid.org/0009-0000-0931-2850)
Institutions
- National University of Computer and Emerging Sciences (PK)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-12
- DOI
- https://doi.org/10.5281/zenodo.22729907
- Primary Topic
- Sentiment Analysis and Opinion Mining
- Type
- preprint