NetFormer: A Dual-Stream Interpretable Transformer Autoencoder for Unsupervised Network Intrusion Detection
Background The growing complexity and frequency of cyberattacks demand intrusion detection systems (IDS) that accurately identify malicious activity with very low false-positive rates and minimal latency. Traditional rule-based and classical machine learning methods are unable to capture the long-range temporal dependencies inherent in multi-stage attacks, and even recurrent neural networks struggle with vanishing gradients over long sequences. Transformers, with self-attention, can model such dependencies, but their application to unsupervised network anomaly detection with mixed data types remains limited. Methods In this work, we introduce NetFormer, a novel Transformer-based unsupervised anomaly detection framework for network traffic time-series. The model features (1) a dual-stream embedding system that separately handles categorical and numerical features, (2) a reconstruction-based autoencoder trained exclusively on normal traffic to compute anomaly scores, (3) an interpretability framework that visualizes attention maps to explain detection decisions. With respect to the results, it can be noted that. Results Evaluated on the CSE-CIC-IDS2018 benchmark, NetFormer achieves F1-score of 0.851, precision of 0.842, recall of 0.861, and 1.24%, as the false positive rate, making it more efficient than the existing classical, LSTM-based, Transformer technologies tested. It excels at detecting volumetric attacks (DDoS F1 = 0.913) and also shows strong performance on slow-rate and subtle anomalies. Cross-dataset validation on UNSW-NB15 confirms robust generalization (F1 = 0.839). Testing on UNSW-NB15 produced 0.839 for F1-score which means that NetFormer is capable of generalizing, however, the transition from one dataset to another is not seamless. Based on the analysis of the attention map it can be concluded that NetFormer detects relevant attack time and traffic features being focused on detecting attacks. Conclusions To sum up, the research shows that the Transformer-based autoencoder is able to learn long-lasting time series, thus offering an excellent outcome during detection without supervision while ensuring low false positive rate at same time.
Authors
- Hiba A. Abu-Alsaad (ORCID: https://orcid.org/0000-0002-5034-0113)
- Mohammed A.S Al-Hitawi (ORCID: https://orcid.org/0009-0009-7905-0978)
- Osama Mohammed (ORCID: https://orcid.org/0009-0006-7591-1660)
- Omar Altalebi
Institutions
- Mustansiriyah University (IQ)
- University Of Fallujah (IQ)
- Middle Technical University (IQ)
Publication Details
- Journal
- F1000Research
- Published
- 2026-09-28
- DOI
- https://doi.org/10.12688/f1000research.182153.2
- Primary Topic
- Network Security and Intrusion Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00