FramePoison: attacks on medical RAG that target framing, not just facts
Retrieval-augmented generation (RAG) is increasingly used in high-stakes domains like medicine, where answers are grounded in retrieved literature and can shape clinical understanding. This opens a new attack surface: an adversary can shape model outputs by injecting a few documents into the retrieval corpus, without touching its weights. Most prior work on corpus poisoning has focused on factual substitution, forcing the model to output attacker-specified false content. But medical answers are explanations, not single facts, and readers are influenced by both what is claimed and how it is presented. We introduce framing-level corpus poisoning, a class of attacks targeting how retrieved evidence is framed, regardless of whether the underlying facts change. We instantiate this class through FramePoison in two realistic settings. In political misinformation, fabricated claims are presented as regulatory directives. In commercial promotion, the core medical content is preserved while the attack steers the answer toward a specific product. Across three biomedical corpora and five open-weight language models, FramePoison is consistently effective under both original and rephrased queries. We inject five poisoned documents per question, and the retrieval prefix typically places at least three in the top-5 results; in successful attacks, one or two poisoned documents in the retrieved context are often sufficient to change the answer.
Authors
- Jinhao Duan (ORCID: https://orcid.org/0000-0001-7045-9376)
- Shan Xie
- Tianshuo Wei (ORCID: https://orcid.org/0009-0008-2470-6688)
- Kaidi Xu
- Ye Wei
- Zijun Zhang
- Hualou Liang
- Xiangyu Zhao
Institutions
- University of North Carolina at Chapel Hill (US)
- Hong Kong Polytechnic University (HK)
- City University of Hong Kong (HK)
Publication Details
- Journal
- npj Artificial Intelligence
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1038/s44387-026-00163-6
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00