When Should RAG Abstain? Predicting Document Sufficiency Before Generation

Retrieval-augmented generation (RAG) systems may retrieve documents that are relevant to a technical question without containing the evidence needed to answer it. This study evaluates a lightweight probabilistic gate that predicts, before generation, whether the retrieved context contains the annotated evidence required to answer a technical-support question. Using the public TechQA dataset, this study combines 22 retrieval-time features derived from lexical BM25, semantic retrieval using BAAI/bge-small-en-v1.5, and reciprocal rank fusion (RRF) in a logistic regression classifier. A predefined protocol uses 360 training, 120 calibration, and 120 internal-validation cases, with the 310-question official development partition reserved as a local holdout. On the TechQA holdout, the proposed gate improved the identification of evidence-bearing contexts compared with the RRF baseline (AP 0.604 versus 0.504). The benefit was particularly relevant under conservative authorization: at 10% coverage, selective risk decreased from 45.3% with RRF to 32.3% with logistic regression. However, the conservative threshold selected during internal validation produced a higher-than-expected risk on the holdout, highlighting the importance of evaluating decision thresholds separately from ranking performance. These results indicate that retrieval-time signals can improve the prioritization of contexts for automatic answering, while discrimination and the transferability of decision thresholds must be assessed separately. The operational label is based on literal inclusion of the annotated answer span; the study does not evaluate generated answers or establish their factual correctness.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-09
DOI
https://doi.org/10.5281/zenodo.23267441
Primary Topic
Information Retrieval and Search Behavior
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

When Should RAG Abstain? Predicting Document Sufficiency Before Generation

Luigi De Facci
Zenodo (CERN European Organization for Nuclear Research)
Information Retrieval and Search Behavior
preprint

When Should RAG Abstain? Predicting Document Sufficiency Before Generation

Luigi De Facci
preprint en

Abstract

Retrieval-augmented generation (RAG) systems may retrieve documents that are relevant to a technical question without containing the evidence needed to answer it. This study evaluates a lightweight probabilistic gate that predicts, before generation, whether the retrieved context contains the annotated evidence required to answer a technical-support question. Using the public TechQA dataset, this study combines 22 retrieval-time features derived from lexical BM25, semantic retrieval using BAAI/bge-small-en-v1.5, and reciprocal rank fusion (RRF) in a logistic regression classifier. A predefined protocol uses 360 training, 120 calibration, and 120 internal-validation cases, with the 310-question official development partition reserved as a local holdout. On the TechQA holdout, the proposed gate improved the identification of evidence-bearing contexts compared with the RRF baseline (AP 0.604 versus 0.504). The benefit was particularly relevant under conservative authorization: at 10% coverage, selective risk decreased from 45.3% with RRF to 32.3% with logistic regression. However, the conservative threshold selected during internal validation produced a higher-than-expected risk on the holdout, highlighting the importance of evaluating decision thresholds separately from ranking performance. These results indicate that retrieval-time signals can improve the prioritization of contexts for automatic answering, while discrimination and the transferability of decision thresholds must be assessed separately. The operational label is based on literal inclusion of the annotated answer span; the study does not evaluate generated answers or establish their factual correctness.

Zenodo (CERN European Organization for Nuclear Research)
Information Retrieval and Search Behavior
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

When Should RAG Abstain? Predicting Document Sufficiency Before Generation — Luigi De Facci · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS