Can Offline-Optimized Latent Adversarial Perturbations Be Amortized? A Study of Teacher–Student Distillation for Real-Time Speech Protection
Real-time adversarial speech protection requires feed-forward inference, whereas the highest-quality adversarial perturbations are typically obtained through expensive per-utterance optimization. We investigate whether these optimized perturbations can be amortized into a lightweight feed-forward generator without sacrificing intelligibility or attack effectiveness. Operating in the latent space of a neural audio codec (EnCodec), we present three experiments. (1) Offline per-utterance latent optimization, regularized for smoothness, produces perturbations that are both effective (WER gain 60 ± 13 pp) and of high measured intelligibility and signal quality (STOI ≈ 0.97, SNR ≈ 18 dB). (2) A directly trained feed-forward universal generator transfers across unseen speakers and an unseen (though same-family) architecture (wav2vec2-Conformer: 65.8 ± 2.1 pp WER gain over 3 seeds) but at a perceptual cost (STOI ≈ 0.77). (3) We attempt to obtain a clean feed-forward generator by distilling the offline teacher into a lightweight student. Across a dataset-size sweep and a per-clip norm-projection diagnostic that fixes the student's budget to the teacher's own, the student fails to transfer the attack (WER gain ≤ 2 pp vs +75 pp for the teacher on the same held-out set). A target-stability test shows the teacher is not a stable, uniquely selected regression target: a 0.1%-of-norm change to the optimizer initialization already yields a comparably effective but substantially different solution (cosine ≈ 0.30). A student-to-modes comparison shows the student recovers no sampled teacher solution and produces no effective perturbation, failing even in-sample. Our experiments do not separate whether this failure reflects a student capacity/optimization limit or mis-specification of the single-target objective under this non-unique target; mode-aware distillation is the direct experiment that would separate them. We frame this as an empirical limitation of the evaluated single-speaker setting, not a fundamental impossibility. All WER figures are measured on EnCodec-reconstructed audio.
Authors
- Pradeep Ramesh Chauhan (ORCID: https://orcid.org/0009-0008-0909-0184)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22847289
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- preprint