RAGBench-X: A Reproducible Framework for Controlled Evaluation of Retrieval-Augmented Question Answering

RAGBench-X is a reproducible framework for controlled evaluation of retrieval-augmented question answering (RAG) systems. It is designed to make retrieval and augmentation choices explicit, configurable, traceable, and repeatable, enabling systematic investigation of how different retrieval components affect question-answering performance. The framework supports a retrieval-free language-model baseline, dense retrieval using Sentence Transformers, sparse BM25 retrieval, reciprocal-rank fusion (RRF), cross-encoder reranking, query reformulation, multi-query retrieval, context compression, source attribution, evidence verification, and structured experiment logging. It provides a FastAPI service and Streamlit research interface, together with persistent experiment storage and raw per-question artifacts. The current evaluation corpus is a pinned subset of FastAPI documentation containing 60 source-linked candidate questions. These questions are currently marked for human review and are not yet a human-verified evaluation benchmark. Engineering validation included 36 automated tests, seven Streamlit dashboard views, and an eight-question retrieval-only evaluation using real Sentence Transformers embeddings, Qdrant, BM25, reciprocal-rank fusion, and cross-encoder reranking. The retrieval check recorded Recall@5 = 1.00, MRR@5 = 0.8125, and mean retrieval latency of approximately 0.442 seconds. These measurements are reported as preliminary engineering validation and should not be interpreted as generalizable research findings. RAGBench-X provides the experimental infrastructure for a subsequent controlled empirical study comparing retrieval and augmentation strategies under a fixed corpus, question set, generation model, and evaluation protocol. The framework is intended to support reproducible research in retrieval-augmented generation, information retrieval, and domain-specific question answering.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.22846184
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

RAGBench-X: A Reproducible Framework for Controlled Evaluation of Retrieval-Augmented Question Answering

Syed Saquib Ali
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
article

RAGBench-X: A Reproducible Framework for Controlled Evaluation of Retrieval-Augmented Question Answering

Syed Saquib Ali
article en

Abstract

RAGBench-X is a reproducible framework for controlled evaluation of retrieval-augmented question answering (RAG) systems. It is designed to make retrieval and augmentation choices explicit, configurable, traceable, and repeatable, enabling systematic investigation of how different retrieval components affect question-answering performance. The framework supports a retrieval-free language-model baseline, dense retrieval using Sentence Transformers, sparse BM25 retrieval, reciprocal-rank fusion (RRF), cross-encoder reranking, query reformulation, multi-query retrieval, context compression, source attribution, evidence verification, and structured experiment logging. It provides a FastAPI service and Streamlit research interface, together with persistent experiment storage and raw per-question artifacts. The current evaluation corpus is a pinned subset of FastAPI documentation containing 60 source-linked candidate questions. These questions are currently marked for human review and are not yet a human-verified evaluation benchmark. Engineering validation included 36 automated tests, seven Streamlit dashboard views, and an eight-question retrieval-only evaluation using real Sentence Transformers embeddings, Qdrant, BM25, reciprocal-rank fusion, and cross-encoder reranking. The retrieval check recorded Recall@5 = 1.00, MRR@5 = 0.8125, and mean retrieval latency of approximately 0.442 seconds. These measurements are reported as preliminary engineering validation and should not be interpreted as generalizable research findings. RAGBench-X provides the experimental infrastructure for a subsequent controlled empirical study comparing retrieval and augmentation strategies under a fixed corpus, question set, generation model, and evaluation protocol. The framework is intended to support reproducible research in retrieval-augmented generation, information retrieval, and domain-specific question answering.

Zenodo (CERN European Organization for Nuclear Research)
Industry, innovation and infrastructure
Openalex Percentile: Top 8%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

RAGBench-X: A Reproducible Framework for Controlled Evaluation of Retrieval-Augmented Question Answering — Syed Saquib Ali · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS