Multimodal generative cyber forensics network for unified malware intelligence synthesis through temporal evidence alignment, neuro-symbolic hypergraph reasoning, and cognitive threat attributions
Abstract Cyber-forensic investigations increasingly involve heterogeneous evidence, including phishing audio, malicious text and scripts, executable binaries, memory artifacts, and network traffic. Existing approaches generally analyze these sources independently or perform sample-level fusion without explicitly addressing temporal asynchrony, higher-order forensic relationships, missing attack stages, counterfactual threat evolution, and attribution uncertainty. This paper presents a Multimodal Generative Cyber Forensics Network comprising CTEEN for temporal-semantic alignment, NSFRH for higher-order forensic reasoning, GASRE for evidence-constrained attack-stage reconstruction, CMES for counterfactual malware-evolution analysis, and CTAECO for confidence-aware candidate-category attribution. The evaluation used VirusShare, CIC-MalMem-2022, the Android Malware Corpus, the Malware Traffic Repository, and an Integrated Controlled Multimodal Benchmark constructed from source pools containing approximately 52,000 malware records, 1.8 million network-flow records, 24,000 textual or script artifacts, and 12,500 phishing or social-engineering audio recordings. Comparisons were conducted against Cyber Code Intelligence, AI-Based Malware Analysis, Multimodal Pre-trained Fusion, and LLM-MalDetect using identical eligible evidence groups, preprocessing outputs, partitions, and metric definitions. Across 25 matched partition–seed observations, the proposed framework achieved 99.26 ± 0.18% classification accuracy (95% CI: 99.19–99.33%), 99.20 ± 0.21% F1-score, and 99.66 ± 0.12% AUC. It further obtained 98.46% mean controlled alignment accuracy, 98.68% evidence-constrained reconstruction accuracy, 99.04% benchmark-defined candidate-category attribution accuracy, and 99.0% overall forensic-intelligence reliability. Accuracy decreased to 95.18% under temporally disjoint testing and 92.34% under family-disjoint testing. Leakage audits, negative controls, matched statistical tests, and alternative-design ablations further supported the findings. The complete framework contained 96.3 million trainable parameters and required 12.8 GB of peak memory, with a mean inference latency of 73.2 ms per complete multimodal episode (13.7 episodes/s) under the reported computational environment. These requirements indicate that deployment on resource-constrained platforms would require model compression or inference optimization. Reviewer 3, Comment 3) The results demonstrate effective evidence-constrained multimodal reasoning under the controlled benchmark but do not establish natural incident-level correspondence, definitive real-world attribution, or universal forensic performance.
Authors
- Rijvan Beg (ORCID: https://orcid.org/0009-0008-6895-3576)
- Ghanshyam Raghuwanshi (ORCID: https://orcid.org/0000-0003-4880-3171)
- Rajesh Kumar Pateriya (ORCID: https://orcid.org/0000-0001-7163-0024)
- Surendra Solanki (ORCID: https://orcid.org/0000-0002-5067-7621)
- Deepak Singh Tomar (ORCID: https://orcid.org/0000-0001-9025-1679)
Institutions
- Manipal University Jaipur
- Maulana Azad National Institute of Technology (IN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1038/s41598-026-69869-6
- Primary Topic
- Advanced Malware Detection Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00