How Much Can Query Routing Help Hybrid Retrieval? A Multi-Benchmark Study with Latency Analysis

Hybrid retrieval pipelines for retrieval-augmented generation (RAG) combine lexical and dense retrieval, often followed by a cross-encoder reranker, yet no single configuration is best for every query. Routing each query to the cheapest adequate strategy is therefore attractive. We ask how much such routing can actually help. On a small synthetic Nginx corpus (180 questions), a cross-validated learned router appeared to improve Recall@1 over fixed hybrid retrieval by 2.2 points, but the gain was not significant, vanished with a second embedding model, and did not replicate at scale. On four BEIR benchmarks (SciFact, NFCorpus, ArguAna, FiQA-2018) with three embedding models and measured latency, fixed-weight hybrid retrieval was the best fixed strategy (average nDCG@10 of 0.497, 0.528 and 0.515, versus 0.459, 0.487 and 0.510 for dense retrieval), beat Reciprocal Rank Fusion in all nine cells tested, and a fixed weight of 0.75 was on average within 0.6 points of the best per-cell weight. A cross-encoder reranker was 1.1 to 3.2 points worse than hybrid retrieval at 4.5 to 7 times the latency. A learned cost-aware router improved on the best fixed strategy in none of the 12 dataset--model cells and was significantly worse in two leave-one-dataset-out cases. Yet choosing the best method for each query would add 8.0 to 8.6 points, and a query-perturbation analysis shows that 74 to 93% of this headroom is stable rather than selection noise. Which method wins on a query is thus a stable property that cheap, label-free features fail to predict, which we present as an open problem. Code and results are provided in a supplementary package.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23114658
Primary Topic
Information Retrieval and Search Behavior
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

How Much Can Query Routing Help Hybrid Retrieval? A Multi-Benchmark Study with Latency Analysis

Shashi Kant
Zenodo (CERN European Organization for Nuclear Research)
Information Retrieval and Search Behavior
preprint

How Much Can Query Routing Help Hybrid Retrieval? A Multi-Benchmark Study with Latency Analysis

Shashi Kant
preprint en

Abstract

Hybrid retrieval pipelines for retrieval-augmented generation (RAG) combine lexical and dense retrieval, often followed by a cross-encoder reranker, yet no single configuration is best for every query. Routing each query to the cheapest adequate strategy is therefore attractive. We ask how much such routing can actually help. On a small synthetic Nginx corpus (180 questions), a cross-validated learned router appeared to improve Recall@1 over fixed hybrid retrieval by 2.2 points, but the gain was not significant, vanished with a second embedding model, and did not replicate at scale. On four BEIR benchmarks (SciFact, NFCorpus, ArguAna, FiQA-2018) with three embedding models and measured latency, fixed-weight hybrid retrieval was the best fixed strategy (average nDCG@10 of 0.497, 0.528 and 0.515, versus 0.459, 0.487 and 0.510 for dense retrieval), beat Reciprocal Rank Fusion in all nine cells tested, and a fixed weight of 0.75 was on average within 0.6 points of the best per-cell weight. A cross-encoder reranker was 1.1 to 3.2 points worse than hybrid retrieval at 4.5 to 7 times the latency. A learned cost-aware router improved on the best fixed strategy in none of the 12 dataset--model cells and was significantly worse in two leave-one-dataset-out cases. Yet choosing the best method for each query would add 8.0 to 8.6 points, and a query-perturbation analysis shows that 74 to 93% of this headroom is stable rather than selection noise. Which method wins on a query is thus a stable property that cheap, label-free features fail to predict, which we present as an open problem. Code and results are provided in a supplementary package.

Zenodo (CERN European Organization for Nuclear Research)
Indian Institute of Technology Patna (IN)
Information Retrieval and Search Behavior
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

How Much Can Query Routing Help Hybrid Retrieval? A Multi-Benchmark Study with Latency Analysis — Shashi Kant · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS