Raw CVE/CWE Retrieval Does Not Improve LLM-Based Vulnerability Detection in Python: A Pre-Specified Null Result and a Retrieval-Dose Audit
Retrieval-Augmented Generation (RAG) over vulnerability databases is widely expected to improve LLM-based vulnerability detection. We report a pre-specified evaluation in which it does not, and derive from it an evaluation protocol. On a near-balanced benchmark of 100 Python 3.12 snippets, raw CVE/CWE retrieval never measurably beat four open-weight models (8B–480B) without retrieval (pooled ΔF1 = +0.006, 95% CI [−0.042; +0.051]; per-model deltas all ≤0). This article makes three contributions. First, the pre-specified null itself: bounded by its confidence interval, homogeneous across models, and invariant to tie-break, model-subset, and pair-dependence conventions. Second, a knowledge-base-overlap audit protocol—exposure identification, retrieval-trace (dose) verification, class-conditional splits, and stratified re-analysis—applied first to our own results, where it localises the null-sized advantage onto the nine snippets exposed to their own CVE entry, eight of which have retrieval traces confirming the entry actually reached the model. Third, exploratory evidence from 164 functions of real advisory-linked fix commits: the null appears there too, and the absolute performance of every tool collapses to near-chance, with a bridge cell attributing that collapse principally to the benchmark rather than to the models. In an exploratory live-corpus arm the pooled retrieval effect remains null, while a benefit appears on the 31% of functions for which the retriever surfaced the matching advisory (dosed ΔF1 = +0.336; dose-stratified difference in recovery, Fisher p < 0.0001), at a specificity cost. Because retrieval dose is confounded with retrievability, we read that benefit as known-vulnerability re-identification rather than as improved detection. We accordingly recommend that RAG security evaluations report overlap strata and retrieval dose as routinely as they report F1. We release both benchmarks, all 19,540 per-run results, and every script.
Authors
- Klaus Gebeshuber
- Patrick Deininger (ORCID: https://orcid.org/0009-0007-9625-3094)
- Stefan Rappl
- Helmut Lindner
Institutions
- FH JOANNEUM University of Applied Sciences (AT)
- Graz University of Technology (AT)
Publication Details
- Journal
- Journal of Cybersecurity and Privacy
- Published
- 2026-09-16
- DOI
- https://doi.org/10.3390/jcp6050163
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00