Robustness Analysis Of Lightweight Machine Learning Models For Sentiment Classification Under Noisy Text

Nowadays, sentimental analysis is an important application of natural language processing and widely used to understand whether a review expresses a positive or negative opinion.Many sentiment classification models are trained using clean data.Realworld reviews often contain spelling mistakes, missing or repeated characters, abbreviations and informal language.A model which executes well on clean data may lose accuracy on noisy text, making robustness essential for evaluation.Previous research has mainly focused on the robustness of neural and transformer-based NLP models, while lightweight machine-learning models have received less attention.These traditional models are simple, fast, resource-efficient and easier to interpret compared to complex deep-learning models.Therefore, this study focuses on evaluating the robustness of lightweight machine-learning classifiers when handling noisy text.The study assesses Multinomial Naive Bayes, Logistic Regression, Linear SVM using TF-IDF representation for fair comparison.The Amazon Cell-Phone Reviews dataset is used to classify reviews into positive and negative sentiments.To simulate real-world text, five types of noise-deletion, insertion, substitution, swapping and abbreviation are introduced at 10%, 20% and 30% severity levels.The noise is programmatically added while keeping the original sentiment labels unchanged.The models are evaluated on noisy and clean data to measure their sturdiness.Performance is examined using precision, recall, accuracy, performance degradation and F1 score.The study addresses which model remains most robust and which type, level of noise causes greatest performance loss.The study helps to determine the level and type of noise which leads to highest performance loss.This finding helps developers and researchers in choosing lightweight model for sentiment analysis.

Authors

Institutions

Publication Details

Journal
International Journal of Innovative Research in Technology
Published
2026-09-15
DOI
https://doi.org/10.64643/ijirt.208465-459
Primary Topic
Sentiment Analysis and Opinion Mining
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Robustness Analysis Of Lightweight Machine Learning Models For Sentiment Classification Under Noisy Text

Ms.Anjali Jawale, Mrs. Shital Pashankar, Mr. Shreyash Dodekar
International Journal of Innovative Research in Technology
Sentiment Analysis and Opinion Mining
article

Robustness Analysis Of Lightweight Machine Learning Models For Sentiment Classification Under Noisy Text

Ms.Anjali Jawale, Mrs. Shital Pashankar, Mr. Shreyash Dodekar
article en

Abstract

Nowadays, sentimental analysis is an important application of natural language processing and widely used to understand whether a review expresses a positive or negative opinion.Many sentiment classification models are trained using clean data.Realworld reviews often contain spelling mistakes, missing or repeated characters, abbreviations and informal language.A model which executes well on clean data may lose accuracy on noisy text, making robustness essential for evaluation.Previous research has mainly focused on the robustness of neural and transformer-based NLP models, while lightweight machine-learning models have received less attention.These traditional models are simple, fast, resource-efficient and easier to interpret compared to complex deep-learning models.Therefore, this study focuses on evaluating the robustness of lightweight machine-learning classifiers when handling noisy text.The study assesses Multinomial Naive Bayes, Logistic Regression, Linear SVM using TF-IDF representation for fair comparison.The Amazon Cell-Phone Reviews dataset is used to classify reviews into positive and negative sentiments.To simulate real-world text, five types of noise-deletion, insertion, substitution, swapping and abbreviation are introduced at 10%, 20% and 30% severity levels.The noise is programmatically added while keeping the original sentiment labels unchanged.The models are evaluated on noisy and clean data to measure their sturdiness.Performance is examined using precision, recall, accuracy, performance degradation and F1 score.The study addresses which model remains most robust and which type, level of noise causes greatest performance loss.The study helps to determine the level and type of noise which leads to highest performance loss.This finding helps developers and researchers in choosing lightweight model for sentiment analysis.

International Journal of Innovative Research in TechnologyVol. 13(5)
Indira Gandhi Institute of Technology (IN)
Openalex Percentile: Top 8%
Sentiment Analysis and Opinion Mining
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Robustness Analysis Of Lightweight Machine Learning Models For Sentiment Classification Under Noisy Text — Ms.Anjali Jawale, Mrs. Shital Pashankar, et al. · International Journal of Innovative Research in Technology (2026) | TGRS Research Map | TGRS