Query Perturbations in RAG: Retrieval Instability, Evidence Availability, and Downstream Accuracy

Query perturbations can change retrieved evidence without changing the intended question. We report a frozen Natural Questions (NQ) study combining manually screened query pairs, BM25s top-๐‘˜ set/rank diagnostics, a lexical evidence-availability proxy, downstream EM/F1, and paired comparisons of four cached query/context inputs: ๐‘„๐‘ + ๐ท๐‘ , ๐‘„๐‘ + ๐ท๐‘ , ๐‘„๐‘ + ๐ท๐‘ , and ๐‘„๐‘ +๐ท๐‘ . The balanced primary cohort contains 100 pairs (25 per perturbation family), selected by frozen rules from 185 pairs labeled ACCEPT in manual semantic-preservation review. Its top-ranked passage changed in 50%. The answer-alias-in-passage@5 lexical proxy was 0.47 for clean versus 0.41 for perturbed queries (perturbed-minus-clean difference โˆ’0.06, 95% CI [โˆ’0.13, 0.01]). The same 100 pairs underwent 400 DeepSeek generation requests. Exact match was 0.38 under ๐‘„๐‘ + ๐ท๐‘ and 0.36 under ๐‘„๐‘ + ๐ท๐‘ ; the prespecified equal-family-weighted point difference was 0.02, while the frozen pooled base-query 95% bootstrap resampling interval was [โˆ’0.06, 0.10]. All overall query-surface, evidence, and interaction contrast intervals included zero. Separately, a 500-base/2,000-pair automatic-candidate retrieval diagnostic, whose inclusion did not require manual semantic-review acceptance, showed a top-1 passage change in 52.7% of pairs, mean Jaccard@5 of 0.428, and a lexical proxy of 0.474 versus 0.3685. The two analyses have different membership and validation status. Retrieval-list instability is observed; the final-100 lexical proxy and downstream answer-score differences are imprecisely estimated and do not resolve an overall decrease.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23118265
Primary Topic
Information Retrieval and Search Behavior
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Query Perturbations in RAG: Retrieval Instability, Evidence Availability, and Downstream Accuracy

Xu Zhou
Zenodo (CERN European Organization for Nuclear Research)
Information Retrieval and Search Behavior
preprint

Query Perturbations in RAG: Retrieval Instability, Evidence Availability, and Downstream Accuracy

Xu Zhou
preprint en

Abstract

Query perturbations can change retrieved evidence without changing the intended question. We report a frozen Natural Questions (NQ) study combining manually screened query pairs, BM25s top-๐‘˜ set/rank diagnostics, a lexical evidence-availability proxy, downstream EM/F1, and paired comparisons of four cached query/context inputs: ๐‘„๐‘ + ๐ท๐‘ , ๐‘„๐‘ + ๐ท๐‘ , ๐‘„๐‘ + ๐ท๐‘ , and ๐‘„๐‘ +๐ท๐‘ . The balanced primary cohort contains 100 pairs (25 per perturbation family), selected by frozen rules from 185 pairs labeled ACCEPT in manual semantic-preservation review. Its top-ranked passage changed in 50%. The answer-alias-in-passage@5 lexical proxy was 0.47 for clean versus 0.41 for perturbed queries (perturbed-minus-clean difference โˆ’0.06, 95% CI [โˆ’0.13, 0.01]). The same 100 pairs underwent 400 DeepSeek generation requests. Exact match was 0.38 under ๐‘„๐‘ + ๐ท๐‘ and 0.36 under ๐‘„๐‘ + ๐ท๐‘ ; the prespecified equal-family-weighted point difference was 0.02, while the frozen pooled base-query 95% bootstrap resampling interval was [โˆ’0.06, 0.10]. All overall query-surface, evidence, and interaction contrast intervals included zero. Separately, a 500-base/2,000-pair automatic-candidate retrieval diagnostic, whose inclusion did not require manual semantic-review acceptance, showed a top-1 passage change in 52.7% of pairs, mean Jaccard@5 of 0.428, and a lexical proxy of 0.474 versus 0.3685. The two analyses have different membership and validation status. Retrieval-list instability is observed; the final-100 lexical proxy and downstream answer-score differences are imprecisely estimated and do not resolve an overall decrease.

Zenodo (CERN European Organization for Nuclear Research)
Information Retrieval and Search Behavior
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Query Perturbations in RAG: Retrieval Instability, Evidence Availability, and Downstream Accuracy โ€” Xu Zhou ยท Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS