Pwnagent: a knowledge-guided multi-agent system for automatic exploit generation

Abstract Automatic Exploit Generation (AEG) plays an important role in proactive assessment of software threats by identifying vulnerabilities and constructing functional payloads. Existing Large Language Model (LLM)-based methods, however, often struggle to reason about complex exploit logic and to perform runtime introspection, leaving a gap between static vulnerability analysis and dynamic memory behavior. We present PwnAgent, an LLM-driven multi-agent framework for end-to-end exploit generation that combines offensive domain knowledge with active runtime introspection. PwnAgent uses a hierarchical knowledge base for multi-stage exploit reasoning and a feedback-driven self-correction engine to calibrate dynamic memory parameters during execution. Because broad Capture The Flag (CTF) benchmarks offer limited binary-exploitation depth and pwn-specific evaluation must balance reproducibility, difficulty progression, and exploit diversity, we construct a 66-task pwn benchmark from public CTF-style challenges. The benchmark is primarily composed of Linux x86/x86-64 ELF binaries and stack-oriented tasks, with smaller format-string, heap, integer-overflow, ARM, and MIPS subsets used as limited probes beyond the dominant setting. Under the same recent Kimi-K2.6 backend, PwnAgent achieves a 62.12% end-to-end success rate, compared with 31.82% for the evaluated PwnGPT baseline, a 30.30 percentage-point gain. These paired results indicate that structured knowledge guidance, execution-grounded measurement, and feedback repair improve LLM-based exploit generation in the evaluated setting, while the absolute success rate shows that fully autonomous exploitation remains challenging.

Authors

Institutions

Publication Details

Journal
Cybersecurity
Published
2026-09-28
DOI
https://doi.org/10.1186/s42400-026-00649-5
Primary Topic
Software Testing and Debugging Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Pwnagent: a knowledge-guided multi-agent system for automatic exploit generation

Qianqiong Wu, Yangyang Geng, Qilong Wu, Qiang Wei et al.
Cybersecurity
Software Testing and Debugging Techniques
article

Pwnagent: a knowledge-guided multi-agent system for automatic exploit generation

Qianqiong Wu, Yangyang Geng, Qilong Wu, Qiang Wei, Chaojie Wei, Jing Huang, Yunfeng Wang
article en

Abstract

Abstract Automatic Exploit Generation (AEG) plays an important role in proactive assessment of software threats by identifying vulnerabilities and constructing functional payloads. Existing Large Language Model (LLM)-based methods, however, often struggle to reason about complex exploit logic and to perform runtime introspection, leaving a gap between static vulnerability analysis and dynamic memory behavior. We present PwnAgent, an LLM-driven multi-agent framework for end-to-end exploit generation that combines offensive domain knowledge with active runtime introspection. PwnAgent uses a hierarchical knowledge base for multi-stage exploit reasoning and a feedback-driven self-correction engine to calibrate dynamic memory parameters during execution. Because broad Capture The Flag (CTF) benchmarks offer limited binary-exploitation depth and pwn-specific evaluation must balance reproducibility, difficulty progression, and exploit diversity, we construct a 66-task pwn benchmark from public CTF-style challenges. The benchmark is primarily composed of Linux x86/x86-64 ELF binaries and stack-oriented tasks, with smaller format-string, heap, integer-overflow, ARM, and MIPS subsets used as limited probes beyond the dominant setting. Under the same recent Kimi-K2.6 backend, PwnAgent achieves a 62.12% end-to-end success rate, compared with 31.82% for the evaluated PwnGPT baseline, a 30.30 percentage-point gain. These paired results indicate that structured knowledge guidance, execution-grounded measurement, and feedback repair improve LLM-based exploit generation in the evaluated setting, while the absolute success rate shows that fully autonomous exploitation remains challenging.

CybersecurityVol. 9(1)
PLA Information Engineering University (CN)
Openalex Percentile: Top 6%
Software Testing and Debugging Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Pwnagent: a knowledge-guided multi-agent system for automatic exploit generation — Qianqiong Wu, Yangyang Geng, et al. · Cybersecurity (2026) | TGRS Research Map | TGRS