Does Script Matter? Evaluating Large Language Models for Urdu Sentiment Analysis across Nastaliq, Roman Urdu, and Code-Switched Text

Urdu is spoken by more than 230 million people, but very little of the Urdu that people actually write online is in the language's formal Nastaliq script. Most of it is either Roman Urdu, which is Urdu typed in Latin letters, or a mixture of Urdu and English. Almost all existing benchmarks for Urdu, however, are built only on Nastaliq text, so it is not clear how well large language models handle the way people really write. In this paper we test six instruction-tuned models, three from the Llama family, two Gemma models, and Gemini Flash-Lite, on binary sentiment classification in all three registers, using 250 balanced examples per register and 4,500 predictions in total. Every model did best on Roman Urdu (mean accuracy 83.9%) and worst on code-switched text (75.8%), with Nastaliq in between (77.5%). We then went through 63 of the misclassifications by hand and found that almost half of them were not model mistakes at all, but wrong labels in the original datasets. Once these are taken into account, the true accuracy of the models is closer to 89–91% in every register. Our results suggest that Nastaliq-only benchmarks understate how useful these models are for ordinary Urdu users, and that the quality of Urdu sentiment datasets is at least as much of a bottleneck as the models themselves.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-12
DOI
https://doi.org/10.5281/zenodo.22729907
Primary Topic
Sentiment Analysis and Opinion Mining
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Does Script Matter? Evaluating Large Language Models for Urdu Sentiment Analysis across Nastaliq, Roman Urdu, and Code-Switched Text

Yousaf Khawar Raja
Zenodo (CERN European Organization for Nuclear Research)
Sentiment Analysis and Opinion Mining
preprint

Does Script Matter? Evaluating Large Language Models for Urdu Sentiment Analysis across Nastaliq, Roman Urdu, and Code-Switched Text

Yousaf Khawar Raja
preprint en

Abstract

Urdu is spoken by more than 230 million people, but very little of the Urdu that people actually write online is in the language's formal Nastaliq script. Most of it is either Roman Urdu, which is Urdu typed in Latin letters, or a mixture of Urdu and English. Almost all existing benchmarks for Urdu, however, are built only on Nastaliq text, so it is not clear how well large language models handle the way people really write. In this paper we test six instruction-tuned models, three from the Llama family, two Gemma models, and Gemini Flash-Lite, on binary sentiment classification in all three registers, using 250 balanced examples per register and 4,500 predictions in total. Every model did best on Roman Urdu (mean accuracy 83.9%) and worst on code-switched text (75.8%), with Nastaliq in between (77.5%). We then went through 63 of the misclassifications by hand and found that almost half of them were not model mistakes at all, but wrong labels in the original datasets. Once these are taken into account, the true accuracy of the models is closer to 89–91% in every register. Our results suggest that Nastaliq-only benchmarks understate how useful these models are for ordinary Urdu users, and that the quality of Urdu sentiment datasets is at least as much of a bottleneck as the models themselves.

Zenodo (CERN European Organization for Nuclear Research)
National University of Computer and Emerging Sciences (PK)
Quality Education
Sentiment Analysis and Opinion Mining
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.