Adversarial vulnerabilities and mitigations in unsupervised machine learning: a systematic review
Abstract The use of unsupervised models is growing with the desire to use unlabelled data, and ensuring their robustness against adversaries has become increasingly essential. Although there are surveys that cover some attacks and defences on unsupervised models, a comprehensive review on this topic is missing. We systematically analysed 93 published papers from an initial list of 21,195 papers, identified via a systematic search and reduced through systematic filtration. Our study covered six unsupervised model categories: clustering, super-resolution, autoencoders, generative adversarial networks, diffusion models, and transformers. We systematically synthesised the eight types of attacks and nine types of defences identified across these models, providing top-level descriptions, comparisons of methods across unsupervised models, and the essential differences of these attacks and defences in supervised versus unsupervised settings. Our results provide a comprehensive review of current research in this area, identify critical gaps, and offer recommendations for future studies to mitigate the security risks associated with unsupervised learning.
Authors
- Mathias Lundteigen Mohus (ORCID: https://orcid.org/0000-0002-1105-6515)
- Jingyue Li
Institutions
- Norwegian University of Science and Technology (NO)
Publication Details
- Journal
- Artificial Intelligence Review
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1007/s10462-026-11695-3
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00