When Retrieved Knowledge Overrides Correct Reasoning: A Knowledge-Conflict Failure Mode in RAG-Augmented Security Code Audit
Retrieval-augmented generation (RAG) is often proposed as a low-cost way to give large language models (LLMs) up-to-date, curated domain knowledge without retraining. We build a pipeline that continuously ingests CVE records from the National Vulnerability Database (NVD) and OSV, organizes them by CWE (Common Weakness Enumeration) category, and injects the most relevant records into an LLM's context when it audits source code for vulnerabilities. Contrary to the assumption that such retrieval can only help or be neutral, we find a concrete, replicated failure mode: retrieved CVE knowledge can cause the model to override its own correct data-flow reasoning in favor of surface pattern-matching against the retrieved examples. On a stratified sample from the OWASP Benchmark, we observe the model correctly identify that a given ternary-guarded value never reaches a dangerous sink -- and then, once given retrieved CVE context, explicitly state that the mitigation exists yet flag the code as vulnerable anyway because it "matches the vulnerable pattern demonstrated in the provided CVEs." Across our (currently small, n=28) evaluation, this manifests as a precision regression with no corresponding recall gain: RAG did not catch anything the base model missed, but did introduce false positives the base model did not make. We describe the system architecture, situate this failure mode within the broader knowledge-conflict literature in RAG research, and report preliminary precision/recall/F1 results, explicit about the small sample size and the further experiments needed before drawing firm conclusions.
Authors
- Seenuvasan T
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23005922
- Primary Topic
- Information and Cyber Security
- Type
- preprint