A Translated Singlish SMS Dataset: Multi-Step LLM Language Detection and Translation for Code-Mixed Creole
Singlish is a code-mixed creole formed from English and multiple Asian languages and dialects, predominantly spoken in Singapore. Translation of Singlish to standard UK English broadens comprehension of the creole and supports downstream analyses such as sentiment analysis, yet few annotated resources exist for this purpose. We present a dataset of 2788 Singlish Short Messaging Service (SMS) texts sampled from the NUS SMS Corpus, each paired with (1) native-speaker annotations of the non-English languages present and their corresponding spans, (2) native-speaker translations into standard English, and (3) Large Language Model (LLM)-generated language-detection and translation outputs produced by five open-source models (Mistral-7B, LLaMA-3.1, Gemma-2, Qwen-2.5 and Phi-3.5) using a multi-step prompting pipeline. We describe the data collection, annotation, and LLM-generation protocols, and technically validate the LLM-generated layers against the native-speaker references using BLEU, ROUGE, and BERTScore, finding an average BERTScore F1 of 0.90. This dataset supports research on code-mixed and creole machine translation, automatic language identification, and the evaluation of LLMs on low-resource, code-mixed language varieties, and enables downstream applications such as sentiment analysis of Singlish text.
Authors
- Lynnette Hui Xian Ng (ORCID: https://orcid.org/0000-0002-2740-7818)
- Luoqi Chan
Institutions
- Carnegie Mellon University (US)
Publication Details
- Journal
- Data
- Published
- 2026-09-11
- DOI
- https://doi.org/10.3390/data11090235
- Primary Topic
- Authorship Attribution and Profiling
- Type
- article
- Field-Weighted Citation Impact
- 0.00