Raw CVE/CWE Retrieval Does Not Improve LLM-Based Vulnerability Detection in Python: A Pre-Specified Null Result and a Retrieval-Dose Audit

Retrieval-Augmented Generation (RAG) over vulnerability databases is widely expected to improve LLM-based vulnerability detection. We report a pre-specified evaluation in which it does not, and derive from it an evaluation protocol. On a near-balanced benchmark of 100 Python 3.12 snippets, raw CVE/CWE retrieval never measurably beat four open-weight models (8B–480B) without retrieval (pooled ΔF1 = +0.006, 95% CI [−0.042; +0.051]; per-model deltas all ≤0). This article makes three contributions. First, the pre-specified null itself: bounded by its confidence interval, homogeneous across models, and invariant to tie-break, model-subset, and pair-dependence conventions. Second, a knowledge-base-overlap audit protocol—exposure identification, retrieval-trace (dose) verification, class-conditional splits, and stratified re-analysis—applied first to our own results, where it localises the null-sized advantage onto the nine snippets exposed to their own CVE entry, eight of which have retrieval traces confirming the entry actually reached the model. Third, exploratory evidence from 164 functions of real advisory-linked fix commits: the null appears there too, and the absolute performance of every tool collapses to near-chance, with a bridge cell attributing that collapse principally to the benchmark rather than to the models. In an exploratory live-corpus arm the pooled retrieval effect remains null, while a benefit appears on the 31% of functions for which the retriever surfaced the matching advisory (dosed ΔF1 = +0.336; dose-stratified difference in recovery, Fisher p < 0.0001), at a specificity cost. Because retrieval dose is confounded with retrievability, we read that benefit as known-vulnerability re-identification rather than as improved detection. We accordingly recommend that RAG security evaluations report overlap strata and retrieval dose as routinely as they report F1. We release both benchmarks, all 19,540 per-run results, and every script.

Authors

Institutions

Publication Details

Journal
Journal of Cybersecurity and Privacy
Published
2026-09-16
DOI
https://doi.org/10.3390/jcp6050163
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Raw CVE/CWE Retrieval Does Not Improve LLM-Based Vulnerability Detection in Python: A Pre-Specified Null Result and a Retrieval-Dose Audit

Klaus Gebeshuber, Patrick Deininger, Stefan Rappl, Helmut Lindner
Journal of Cybersecurity and Privacy
Adversarial Robustness in Machine Learning
article

Raw CVE/CWE Retrieval Does Not Improve LLM-Based Vulnerability Detection in Python: A Pre-Specified Null Result and a Retrieval-Dose Audit

Klaus Gebeshuber, Patrick Deininger, Stefan Rappl, Helmut Lindner
article en

Abstract

Retrieval-Augmented Generation (RAG) over vulnerability databases is widely expected to improve LLM-based vulnerability detection. We report a pre-specified evaluation in which it does not, and derive from it an evaluation protocol. On a near-balanced benchmark of 100 Python 3.12 snippets, raw CVE/CWE retrieval never measurably beat four open-weight models (8B–480B) without retrieval (pooled ΔF1 = +0.006, 95% CI [−0.042; +0.051]; per-model deltas all ≤0). This article makes three contributions. First, the pre-specified null itself: bounded by its confidence interval, homogeneous across models, and invariant to tie-break, model-subset, and pair-dependence conventions. Second, a knowledge-base-overlap audit protocol—exposure identification, retrieval-trace (dose) verification, class-conditional splits, and stratified re-analysis—applied first to our own results, where it localises the null-sized advantage onto the nine snippets exposed to their own CVE entry, eight of which have retrieval traces confirming the entry actually reached the model. Third, exploratory evidence from 164 functions of real advisory-linked fix commits: the null appears there too, and the absolute performance of every tool collapses to near-chance, with a bridge cell attributing that collapse principally to the benchmark rather than to the models. In an exploratory live-corpus arm the pooled retrieval effect remains null, while a benefit appears on the 31% of functions for which the retriever surfaced the matching advisory (dosed ΔF1 = +0.336; dose-stratified difference in recovery, Fisher p < 0.0001), at a specificity cost. Because retrieval dose is confounded with retrievability, we read that benefit as known-vulnerability re-identification rather than as improved detection. We accordingly recommend that RAG security evaluations report overlap strata and retrieval dose as routinely as they report F1. We release both benchmarks, all 19,540 per-run results, and every script.

Journal of Cybersecurity and PrivacyVol. 6(5)
FH JOANNEUM University of Applied Sciences (AT), Graz University of Technology (AT)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.