A benchmark dataset for powershell fileless malware detection: methodology, validation, and n-Gram feature extraction

Abstract PowerShell is now an integral part of the Windows operating system and includes robust system management, automation, and configuration capabilities. Yet the PowerShell-based malware detection is a challenge as its malicious functionalities can be covered with obfuscation and get executed without any conventional executable artifacts. Many existing detection approaches are limits with proper justification on validation and labeling. Also many of them focus on the detection of PowerShell based malicious script without give emphasis on the preprocessing and feature extraction process. This study proposed a benchmark-oriented detection framework considering the advanced multi-phase deobfuscation, script normalization, VirusTotal based label validation, n-gram feature extraction and Functional Link Artificial Neural Network based classification. The proposed framework is measured using multiple n-gram configurations and standard classification metrics with a regulated train-test protocol. The Trigonometric Functional Link Artificial Neural Network (TFLANN) model achieve an accuracy of 98.61% and 0.9930 ROC-AUC using 5-gram feature sets under the reported 70:30 evaluation setting. Repeated seed trails and paired statistical tests validated the model’s rank among the four functional expansions and conventional linear and ensemble classifiers were used to put the reported accuracy into perspective. The result represent that the efficient preprocessing and lexical representation support accurate detection of malicious PowerShell script to a grate extent.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-08
DOI
https://doi.org/10.1038/s41598-026-74828-2
Primary Topic
Advanced Malware Detection Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A benchmark dataset for powershell fileless malware detection: methodology, validation, and n-Gram feature extraction

Aswini Kumar Samantaray, Ganapati Panda, Adyasha Rath, Manish Kumar Meher et al.
Scientific Reports
Advanced Malware Detection Techniques
article

A benchmark dataset for powershell fileless malware detection: methodology, validation, and n-Gram feature extraction

Aswini Kumar Samantaray, Ganapati Panda, Adyasha Rath, Manish Kumar Meher, Prabodh Kumar Sahoo
article en

Abstract

Abstract PowerShell is now an integral part of the Windows operating system and includes robust system management, automation, and configuration capabilities. Yet the PowerShell-based malware detection is a challenge as its malicious functionalities can be covered with obfuscation and get executed without any conventional executable artifacts. Many existing detection approaches are limits with proper justification on validation and labeling. Also many of them focus on the detection of PowerShell based malicious script without give emphasis on the preprocessing and feature extraction process. This study proposed a benchmark-oriented detection framework considering the advanced multi-phase deobfuscation, script normalization, VirusTotal based label validation, n-gram feature extraction and Functional Link Artificial Neural Network based classification. The proposed framework is measured using multiple n-gram configurations and standard classification metrics with a regulated train-test protocol. The Trigonometric Functional Link Artificial Neural Network (TFLANN) model achieve an accuracy of 98.61% and 0.9930 ROC-AUC using 5-gram feature sets under the reported 70:30 evaluation setting. Repeated seed trails and paired statistical tests validated the model’s rank among the four functional expansions and conventional linear and ensemble classifiers were used to put the reported accuracy into perspective. The result represent that the efficient preprocessing and lexical representation support accurate detection of malicious PowerShell script to a grate extent.

Scientific Reports
Openalex Percentile: Top 12%
Advanced Malware Detection Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.