Deep feature selection and neural architectures for robust darknet traffic classification

Abstract Accurately categorizing darknet traffic remains a significant cybersecurity challenge because of the encrypted and anonymous nature of communications made possible by technologies like Virtual Private Networks (VPNs) and The Onion Router (Tor). This work proposes a structured deep learning (DL)-based framework for darknet traffic classification using both individual and combinatorial feature selection (FS) techniques on the CIC-Darknet2020 dataset. We employed five FS techniques, Mutual Information (MI), Random Forest (RF), Recursive Feature Elimination (RFE), Autoencoder-based FS, and XGBoost (XGB), to identify the top 20 discriminative features from the dataset. 31 unique feature subsets with both individual and intersected feature combinations were generated using a combinatorial FS approach. The CNN model trained on RFE-selected features outperformed all other models evaluated, with a testing accuracy of 95.2% and a logarithmic loss of 0.1257. Kuncheva’s Consistency Index was used to measure the stability of each FS technique’s own feature rankings, and the results showed that tree-based methods yielded more repeatable rankings than the Autoencoder-based approach. Using multi-seed retraining and formal statistical significance testing, this ranking was further validated. Furthermore, it was shown that the top-performing configurations were practically suitable for operational deployment because they were computationally efficient and well-calibrated for real-time inference. Additionally, the results demonstrate that optimized FS techniques greatly enhance classification performance when compared to increasing architectural complexity. Although combinatorial FS methods outperformed MI and Autoencoder-based strategies in terms of robustness and generalization, the best-performing models were primarily derived from individual tree-based FS techniques. The analysis also discovered that Tor traffic was the most difficult to classify due to its encrypted and obfuscated communication patterns. The proposed framework offers a solid foundation for cybersecurity and scalable darknet traffic analysis, and it demonstrates a strong capacity for generalization.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-03
DOI
https://doi.org/10.1038/s41598-026-73927-4
Primary Topic
Internet Traffic Analysis and Secure E-voting
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Deep feature selection and neural architectures for robust darknet traffic classification

Waleed M. Ismael, G. Kirubavathi, Amal Ajayan, Jaeyoung Choi et al.
Scientific Reports
Internet Traffic Analysis and Secure E-voting
article

Deep feature selection and neural architectures for robust darknet traffic classification

Waleed M. Ismael, G. Kirubavathi, Amal Ajayan, Jaeyoung Choi, Akash Biju
article en

Abstract

Abstract Accurately categorizing darknet traffic remains a significant cybersecurity challenge because of the encrypted and anonymous nature of communications made possible by technologies like Virtual Private Networks (VPNs) and The Onion Router (Tor). This work proposes a structured deep learning (DL)-based framework for darknet traffic classification using both individual and combinatorial feature selection (FS) techniques on the CIC-Darknet2020 dataset. We employed five FS techniques, Mutual Information (MI), Random Forest (RF), Recursive Feature Elimination (RFE), Autoencoder-based FS, and XGBoost (XGB), to identify the top 20 discriminative features from the dataset. 31 unique feature subsets with both individual and intersected feature combinations were generated using a combinatorial FS approach. The CNN model trained on RFE-selected features outperformed all other models evaluated, with a testing accuracy of 95.2% and a logarithmic loss of 0.1257. Kuncheva’s Consistency Index was used to measure the stability of each FS technique’s own feature rankings, and the results showed that tree-based methods yielded more repeatable rankings than the Autoencoder-based approach. Using multi-seed retraining and formal statistical significance testing, this ranking was further validated. Furthermore, it was shown that the top-performing configurations were practically suitable for operational deployment because they were computationally efficient and well-calibrated for real-time inference. Additionally, the results demonstrate that optimized FS techniques greatly enhance classification performance when compared to increasing architectural complexity. Although combinatorial FS methods outperformed MI and Autoencoder-based strategies in terms of robustness and generalization, the best-performing models were primarily derived from individual tree-based FS techniques. The analysis also discovered that Tor traffic was the most difficult to classify due to its encrypted and obfuscated communication patterns. The proposed framework offers a solid foundation for cybersecurity and scalable darknet traffic analysis, and it demonstrates a strong capacity for generalization.

Scientific Reports
Gachon University (KR), Azal University for Human Development (YE), Amrita Vishwa Vidyapeetham (IN)
Openalex Percentile: Top 9%
Internet Traffic Analysis and Secure E-voting
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.