Contrastive Sequence Learning for DoH Tunnel Detection and Fine-Grained Traffic Classification
DNS over HTTPS (DoH) protects name-resolution traffic but can also conceal command-and-control and data-exfiltration channels. Detection is challenging because benign and malicious flows share similar statistics, packet relationships span long sequences, and labeled malicious examples are scarce. We present the Contrastive Learning–Transformer–Bidirectional Gated Recurrent Unit–Attention model (CL–TBiGRU–Attention). The model combines contrastive pretraining with global and bidirectional temporal modeling. We evaluate binary DoH/non-DoH detection, tunneling-tool classification on two datasets, and malware-family classification. Across these four tasks, accuracy ranges from 94.91% to 99.98%, and the F1-score ranges from 95.02% to 99.98%. The model outperforms the listed neural baselines in binary detection and all listed methods in the three fine-grained tasks. On malware-family classification, it reaches 97.00% accuracy and 97.01% F1-score, with most residual confusion occurring between Sisron and Zloader. These results show the value of contrastive sequence modeling for coarse- and fine-grained DoH traffic analysis on the evaluated benchmarks.
Authors
- Jun Yin (ORCID: https://orcid.org/0000-0002-9085-3925)
- Peng Zhang (ORCID: https://orcid.org/0000-0001-9518-5914)
- Yanlei Liu
- Yi Gu (ORCID: https://orcid.org/0000-0002-2320-5848)
- Tongjie Wei
- Peng Wang
Institutions
- Nanjing University of Science and Technology (CN)
- Inner Mongolia Electric Power (China) (CN)
Publication Details
- Journal
- Computers
- Published
- 2026-09-14
- DOI
- https://doi.org/10.3390/computers15090617
- Primary Topic
- Internet Traffic Analysis and Secure E-voting
- Type
- article
- Field-Weighted Citation Impact
- 0.00