Efficient Certified Robustness Assessment for Transformer Models via Verifier-Aware Perturbation-Position Scheduling

Certified robustness provides guarantees that a model’s prediction is stable within a specified perturbation region, but certified-radius assessment for Transformer text classifiers is computationally expensive because sentence-level certificates require repeated radius searches over multiple perturbation positions. Existing approaches primarily improve the precision or runtime of individual verifier calls, whereas repeated position-level search is an additional source of computational cost. We address this bottleneck with verifier-aware perturbation-position scheduling: selected positions undergo full radius search to obtain a candidate sentence-level radius, skipped positions are validated at that radius, and only failed validations receive fallback search. Thus, scheduling reduces repeated search effort but preserves verifier coverage of all eligible perturbation tasks. Experiments include an expanded 100-example SST-2 evaluation and a 30-example Yelp evaluation, together with one-token and two-token embedding-space perturbations, early-exit verifier cascades, and scheduling ablations, to assess certificate preservation and verification efficiency across different settings. On the expanded SST-2 all-position evaluation, the scheduler reproduces the exhaustive sentence-level radius on all examples, while substantially reduces the number of full radius searches. Additional sensitivity experiments show similar scheduling behavior under the evaluated depth and longer-sequence settings. The presented results suggest that perturbation-position scheduling provides a verifier-compatible way to reduce repeated work in certified robustness assessment. Next research steps should expand the evaluation by addressing large pretrained language models, alternative verification backends, and broader discrete text perturbation settings.

Authors

Institutions

Publication Details

Journal
Machine Learning and Knowledge Extraction
Published
2026-09-20
DOI
https://doi.org/10.3390/make8090290
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Efficient Certified Robustness Assessment for Transformer Models via Verifier-Aware Perturbation-Position Scheduling

Luca Lazzaroni, Alessandro Pighetti, Riccardo Berta, Francesco Bellotti et al.
Machine Learning and Knowledge Extraction
Natural Language Processing Techniques
article

Efficient Certified Robustness Assessment for Transformer Models via Verifier-Aware Perturbation-Position Scheduling

Luca Lazzaroni, Alessandro Pighetti, Riccardo Berta, Francesco Bellotti, Vafali Soltanmuradov, David Martin Gomez
article en

Abstract

Certified robustness provides guarantees that a model’s prediction is stable within a specified perturbation region, but certified-radius assessment for Transformer text classifiers is computationally expensive because sentence-level certificates require repeated radius searches over multiple perturbation positions. Existing approaches primarily improve the precision or runtime of individual verifier calls, whereas repeated position-level search is an additional source of computational cost. We address this bottleneck with verifier-aware perturbation-position scheduling: selected positions undergo full radius search to obtain a candidate sentence-level radius, skipped positions are validated at that radius, and only failed validations receive fallback search. Thus, scheduling reduces repeated search effort but preserves verifier coverage of all eligible perturbation tasks. Experiments include an expanded 100-example SST-2 evaluation and a 30-example Yelp evaluation, together with one-token and two-token embedding-space perturbations, early-exit verifier cascades, and scheduling ablations, to assess certificate preservation and verification efficiency across different settings. On the expanded SST-2 all-position evaluation, the scheduler reproduces the exhaustive sentence-level radius on all examples, while substantially reduces the number of full radius searches. Additional sensitivity experiments show similar scheduling behavior under the evaluated depth and longer-sequence settings. The presented results suggest that perturbation-position scheduling provides a verifier-compatible way to reduce repeated work in certified robustness assessment. Next research steps should expand the evaluation by addressing large pretrained language models, alternative verification backends, and broader discrete text perturbation settings.

Machine Learning and Knowledge ExtractionVol. 8(9)
Universidad Carlos III de Madrid (ES), University of Genoa (IT)
Openalex Percentile: Top 8%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.