Character-level Perturbations in Software Requirements : Evaluating Transformer Models Robustness Across Learning Paradigms and Classification Tasks

Abstract Transformer models have shown promising results in software requirements classification. However, their robustness to real-world textual noise remains underexplored. In practice, requirements are frequently affected by typographical errors, transmission artifacts, and optical character recognition mistakes, which can degrade model reliability. This study investigates the robustness of transformer models for requirements classification under character-level adversarial perturbations from two complementary analytical perspectives: transfer learning paradigm and task type. We conduct an empirical evaluation of four models (BERT, RoBERTa, All-MiniLM, and BART) across three transfer learning paradigms (zero-shot, few-shot, and fine-tuning) and three classification tasks of increasing complexity (functional binary, security binary, and NFR multi-class). Experiments are performed on benchmark software requirement datasets perturbed at three character-level noise intensities, where 5%, 10%, and 15% of the characters in each requirement statement are modified. Our results indicate a performance–robustness trade-off across learning paradigms. Zero-shot learning shows notable robustness, with no statistically significant performance degradation, whereas fine-tuning, despite achieving the highest baseline performance, exhibits considerable vulnerability, with degradation of up to 41.5%. With respect to task complexity, security classification shows relatively high robustness (15.2% degradation), while functional and NFR classification are more affected, with degradation rates of 35.4% and 38.2%, respectively. Furthermore, no single transformer model performs consistently well across all settings. These findings offer practical guidance for selecting robust transformer-based configurations for processing software requirements from noisy and imperfect sources.

Authors

Institutions

Publication Details

Journal
International Journal of Computational Intelligence Systems
Published
2026-09-19
DOI
https://doi.org/10.1007/s44196-026-01589-1
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Character-level Perturbations in Software Requirements : Evaluating Transformer Models Robustness Across Learning Paradigms and Classification Tasks

Waad Alhoshan, Wejdan Alqahtani
International Journal of Computational Intelligence Systems
Software Engineering Research
article

Character-level Perturbations in Software Requirements : Evaluating Transformer Models Robustness Across Learning Paradigms and Classification Tasks

Waad Alhoshan, Wejdan Alqahtani
article en

Abstract

Abstract Transformer models have shown promising results in software requirements classification. However, their robustness to real-world textual noise remains underexplored. In practice, requirements are frequently affected by typographical errors, transmission artifacts, and optical character recognition mistakes, which can degrade model reliability. This study investigates the robustness of transformer models for requirements classification under character-level adversarial perturbations from two complementary analytical perspectives: transfer learning paradigm and task type. We conduct an empirical evaluation of four models (BERT, RoBERTa, All-MiniLM, and BART) across three transfer learning paradigms (zero-shot, few-shot, and fine-tuning) and three classification tasks of increasing complexity (functional binary, security binary, and NFR multi-class). Experiments are performed on benchmark software requirement datasets perturbed at three character-level noise intensities, where 5%, 10%, and 15% of the characters in each requirement statement are modified. Our results indicate a performance–robustness trade-off across learning paradigms. Zero-shot learning shows notable robustness, with no statistically significant performance degradation, whereas fine-tuning, despite achieving the highest baseline performance, exhibits considerable vulnerability, with degradation of up to 41.5%. With respect to task complexity, security classification shows relatively high robustness (15.2% degradation), while functional and NFR classification are more affected, with degradation rates of 35.4% and 38.2%, respectively. Furthermore, no single transformer model performs consistently well across all settings. These findings offer practical guidance for selecting robust transformer-based configurations for processing software requirements from noisy and imperfect sources.

International Journal of Computational Intelligence Systems
Imam Mohammad ibn Saud Islamic University (SA)
Openalex Percentile: Top 4%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Character-level Perturbations in Software Requirements : Evaluating Transformer Models Robustness Across Learning Paradigms and Classification Tasks — Waad Alhoshan, Wejdan Alqahtani · International Journal of Computational Intelligence Systems (2026) | TGRS Research Map | TGRS