Augmentation of Sentiment Analysis Models
Sentiment analysis is widely used to quantify the emotional tone of text. However, many pretrained sentiment models are only designed for short inputs and can struggle when applied to longer compilations of text or even full documents. These limitations are especially important when a continuous sentiment score range is needed rather than a simple positive, neutral, or negative label. In this research, sixteen pretrained sentiment models were evaluated on sentence-level, paragraph-level, and document-level benchmark datasets to determine their compatibility for continuous sentiment scoring on a normalized output scale. To mitigate the effect of input truncation and preserve sentiment granularity on longer texts, an augmentation layer that combines sentence-level, paragraph-level, and document-level sentiment estimates through weighted aggregation and sliding window sentence aggregation was added to each compatible model. Model performance was calculated using RMSE against ground-truth benchmark datasets, and model behavior was further examined through score distributions and residual correlations. Results show that the augmentation layer generally improved performance for paragraph-level and document-level text, with more noticeable improvements at the document level. Among the evaluated pretrained models, DeBERTa provided the best overall balance of benchmark performance and continuous score resolution, making it the most defensible model for downstream sentiment analysis, although substantial rank agreement was observed among the strongest candidate models.
Authors
- Alan J. Michaels (ORCID: https://orcid.org/0000-0003-2437-3410)
- Christopher Henshaw (ORCID: https://orcid.org/0009-0009-5769-2590)
- Jared Byers (ORCID: https://orcid.org/0009-0008-7092-3173)
Institutions
- Virginia Tech (US)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-20
- DOI
- https://doi.org/10.3390/electronics15184310
- Primary Topic
- Sentiment Analysis and Opinion Mining
- Type
- article
- Field-Weighted Citation Impact
- 0.00