Auditing VirusTotal consensus thresholds in malware benchmarks: A reproducible security-measurement study on EMBER2024

VirusTotal (VT) consensus is widely used to construct malware benchmarks, yet the consensus threshold is usually fixed without sensitivity analysis. We treat this threshold as a benchmark-design variable, not a new detector. On EMBER2024, released binary labels remain fixed; raising the threshold T only removes low-consensus malware from the training pool. A fixed-label audit sweeps T from the released default 0.065 to 0.50 across four classifier families, an 11-point LightGBM grid, calibration, fixed-form weighting, file-type analyses, and three control arms, totaling 600 model fits over ten seeds. LightGBM identifies an empirical positive effect region T โˆˆ [0.08, 0.18], peaking at ๐‘‡ = 0 . 1 2 with a 4.73 percentage point gain in challenge TPR@1%FPR (Cohenโ€™s ๐‘‘ = 2 . 0 3 , Holm-adjusted ๐‘ = . 0 2 0 ) and 5โ€“8 ร— lower seed dispersion; XGBoost shows the same coarse-grid pattern. A matched-size control finds no discrimination gain under random removal; random removal accounts for about one fifth of the ECE reduction. ECE decreases through this region but continues improving after discrimination degrades, so calibration alone does not select an operating point. All eight tested fixed-form weighting schemes reduce low-consensus-malware detection. The numerical range is specific to EMBER2024, a nominal denominator of approximately 77 engines, and gradient-boosted classifiers. The reusable contribution is the audit procedure and panel-relative reporting rule ๐‘˜ = โŒˆ ๐‘‡ โข ๐‘ โŒ‰ ; each new corpus requires re-audit against labels fixed independently of the candidate training thresholds.

Authors

Institutions

Publication Details

Journal
Journal of Information Security and Applications
Published
2026-09-29
DOI
https://doi.org/10.1016/j.jisa.2026.104659
Primary Topic
Advanced Malware Detection Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Auditing VirusTotal consensus thresholds in malware benchmarks: A reproducible security-measurement study on EMBER2024

Deโ€Thu Huynh, Trong Thua Huynh, Van-Quynh Trinh, Ngoc-Hieu Le
Journal of Information Security and Applications
Advanced Malware Detection Techniques
article

Auditing VirusTotal consensus thresholds in malware benchmarks: A reproducible security-measurement study on EMBER2024

Deโ€Thu Huynh, Trong Thua Huynh, Van-Quynh Trinh, Ngoc-Hieu Le
article en

Abstract

VirusTotal (VT) consensus is widely used to construct malware benchmarks, yet the consensus threshold is usually fixed without sensitivity analysis. We treat this threshold as a benchmark-design variable, not a new detector. On EMBER2024, released binary labels remain fixed; raising the threshold T only removes low-consensus malware from the training pool. A fixed-label audit sweeps T from the released default 0.065 to 0.50 across four classifier families, an 11-point LightGBM grid, calibration, fixed-form weighting, file-type analyses, and three control arms, totaling 600 model fits over ten seeds. LightGBM identifies an empirical positive effect region T โˆˆ [0.08, 0.18], peaking at ๐‘‡ = 0 . 1 2 with a 4.73 percentage point gain in challenge TPR@1%FPR (Cohenโ€™s ๐‘‘ = 2 . 0 3 , Holm-adjusted ๐‘ = . 0 2 0 ) and 5โ€“8 ร— lower seed dispersion; XGBoost shows the same coarse-grid pattern. A matched-size control finds no discrimination gain under random removal; random removal accounts for about one fifth of the ECE reduction. ECE decreases through this region but continues improving after discrimination degrades, so calibration alone does not select an operating point. All eight tested fixed-form weighting schemes reduce low-consensus-malware detection. The numerical range is specific to EMBER2024, a nominal denominator of approximately 77 engines, and gradient-boosted classifiers. The reusable contribution is the audit procedure and panel-relative reporting rule ๐‘˜ = โŒˆ ๐‘‡ โข ๐‘ โŒ‰ ; each new corpus requires re-audit against labels fixed independently of the candidate training thresholds.

Journal of Information Security and ApplicationsVol. 103
Posts and Telecommunications Institute of Technology - Ho Chi Minh City (VN), Saigon International University (VN)
Gender equality
Openalex Percentile: Top 10%
Advanced Malware Detection Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Auditing VirusTotal consensus thresholds in malware benchmarks: A reproducible security-measurement study on EMBER2024 โ€” Deโ€Thu Huynh, Trong Thua Huynh, et al. ยท Journal of Information Security and Applications (2026) | TGRS Research Map | TGRS