A Translated Singlish SMS Dataset: Multi-Step LLM Language Detection and Translation for Code-Mixed Creole

Singlish is a code-mixed creole formed from English and multiple Asian languages and dialects, predominantly spoken in Singapore. Translation of Singlish to standard UK English broadens comprehension of the creole and supports downstream analyses such as sentiment analysis, yet few annotated resources exist for this purpose. We present a dataset of 2788 Singlish Short Messaging Service (SMS) texts sampled from the NUS SMS Corpus, each paired with (1) native-speaker annotations of the non-English languages present and their corresponding spans, (2) native-speaker translations into standard English, and (3) Large Language Model (LLM)-generated language-detection and translation outputs produced by five open-source models (Mistral-7B, LLaMA-3.1, Gemma-2, Qwen-2.5 and Phi-3.5) using a multi-step prompting pipeline. We describe the data collection, annotation, and LLM-generation protocols, and technically validate the LLM-generated layers against the native-speaker references using BLEU, ROUGE, and BERTScore, finding an average BERTScore F1 of 0.90. This dataset supports research on code-mixed and creole machine translation, automatic language identification, and the evaluation of LLMs on low-resource, code-mixed language varieties, and enables downstream applications such as sentiment analysis of Singlish text.

Authors

Institutions

Publication Details

Journal
Data
Published
2026-09-11
DOI
https://doi.org/10.3390/data11090235
Primary Topic
Authorship Attribution and Profiling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Translated Singlish SMS Dataset: Multi-Step LLM Language Detection and Translation for Code-Mixed Creole

Lynnette Hui Xian Ng, Luoqi Chan
Data
Authorship Attribution and Profiling
article

A Translated Singlish SMS Dataset: Multi-Step LLM Language Detection and Translation for Code-Mixed Creole

Lynnette Hui Xian Ng, Luoqi Chan
article en

Abstract

Singlish is a code-mixed creole formed from English and multiple Asian languages and dialects, predominantly spoken in Singapore. Translation of Singlish to standard UK English broadens comprehension of the creole and supports downstream analyses such as sentiment analysis, yet few annotated resources exist for this purpose. We present a dataset of 2788 Singlish Short Messaging Service (SMS) texts sampled from the NUS SMS Corpus, each paired with (1) native-speaker annotations of the non-English languages present and their corresponding spans, (2) native-speaker translations into standard English, and (3) Large Language Model (LLM)-generated language-detection and translation outputs produced by five open-source models (Mistral-7B, LLaMA-3.1, Gemma-2, Qwen-2.5 and Phi-3.5) using a multi-step prompting pipeline. We describe the data collection, annotation, and LLM-generation protocols, and technically validate the LLM-generated layers against the native-speaker references using BLEU, ROUGE, and BERTScore, finding an average BERTScore F1 of 0.90. This dataset supports research on code-mixed and creole machine translation, automatic language identification, and the evaluation of LLMs on low-resource, code-mixed language varieties, and enables downstream applications such as sentiment analysis of Singlish text.

DataVol. 11(9)
Carnegie Mellon University (US)
Quality Education
Openalex Percentile: Top 8%
Authorship Attribution and Profiling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Translated Singlish SMS Dataset: Multi-Step LLM Language Detection and Translation for Code-Mixed Creole — Lynnette Hui Xian Ng, Luoqi Chan · Data (2026) | TGRS Research Map | TGRS