WinAPIReplay: Safely Re-Executing Win32 and NT-Native Malware API-Call Logs to Measure Behavioral Reproducibility and Generate Labeled Endpoint Telemetry

Behavioral malware analysis relies on dynamic-analysis logs—sequences of Windows API calls recorded by sandboxes such as CAPEv2—assumed to represent the malware’s effects faithfully. To the best of our knowledge, this assumption has never been tested by re-executing the recorded calls. We propose WinAPIReplay, which re-executes each recorded Win32 and NT-native call as a real operating-system call, without the malware binary, using a unified handle map for cross-layer handle chains and a three-tier sandbox confining every side effect to disposable places. Across 500 WinMET samples from five families, 74.04% of modeled Layer-1 behavior-domain calls are reproduced (95% bootstrap CI [71.02, 76.75]); the Layer-1 domain covers 57.98% of all recorded calls, and a conservative rate excluding substituted calls is 71.15%. An ablation attributes this causally to the per-category executors—handle-validity falls from 94.9% to 10.2% without them—and no side effect escapes the sandbox on the channels the tool models and monitors. A four-way taxonomy assigns most of the residual to intrinsic, environment-dependent behavior; re-execution is near-deterministic (98.90% stable). Under Sysmon, it safely generates family-labeled file/registry telemetry for the reproduced subset (46,150 events) from static logs alone—while its result record recovers the malware’s process arguments and network destinations—and a classifier over it reaches 92.0% leave-one-out accuracy on 100 samples, statistically indistinguishable from an API-category baseline. WinAPIReplay thus provides, to the best of our knowledge, the first quantitative measurement of dynamic-log reproducibility within a demonstrated safety envelope, and a safe route to labeled, environment-consistent endpoint telemetry.

Authors

Institutions

Publication Details

Journal
Information
Published
2026-09-16
DOI
https://doi.org/10.3390/info17090905
Primary Topic
Advanced Malware Detection Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

WinAPIReplay: Safely Re-Executing Win32 and NT-Native Malware API-Call Logs to Measure Behavioral Reproducibility and Generate Labeled Endpoint Telemetry

Masami Mohri, Youji Fukuta, Masanori Hirotomo, Yoshiaki Shiraishi
Information
Advanced Malware Detection Techniques
article

WinAPIReplay: Safely Re-Executing Win32 and NT-Native Malware API-Call Logs to Measure Behavioral Reproducibility and Generate Labeled Endpoint Telemetry

Masami Mohri, Youji Fukuta, Masanori Hirotomo, Yoshiaki Shiraishi
article en

Abstract

Behavioral malware analysis relies on dynamic-analysis logs—sequences of Windows API calls recorded by sandboxes such as CAPEv2—assumed to represent the malware’s effects faithfully. To the best of our knowledge, this assumption has never been tested by re-executing the recorded calls. We propose WinAPIReplay, which re-executes each recorded Win32 and NT-native call as a real operating-system call, without the malware binary, using a unified handle map for cross-layer handle chains and a three-tier sandbox confining every side effect to disposable places. Across 500 WinMET samples from five families, 74.04% of modeled Layer-1 behavior-domain calls are reproduced (95% bootstrap CI [71.02, 76.75]); the Layer-1 domain covers 57.98% of all recorded calls, and a conservative rate excluding substituted calls is 71.15%. An ablation attributes this causally to the per-category executors—handle-validity falls from 94.9% to 10.2% without them—and no side effect escapes the sandbox on the channels the tool models and monitors. A four-way taxonomy assigns most of the residual to intrinsic, environment-dependent behavior; re-execution is near-deterministic (98.90% stable). Under Sysmon, it safely generates family-labeled file/registry telemetry for the reproduced subset (46,150 events) from static logs alone—while its result record recovers the malware’s process arguments and network destinations—and a classifier over it reaches 92.0% leave-one-out accuracy on 100 samples, statistically indistinguishable from an API-category baseline. WinAPIReplay thus provides, to the best of our knowledge, the first quantitative measurement of dynamic-log reproducibility within a demonstrated safety envelope, and a safe route to labeled, environment-consistent endpoint telemetry.

InformationVol. 17(9)
Saga University (JP), Kobe University (JP), Kindai University (JP)
Openalex Percentile: Top 9%
Advanced Malware Detection Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.