Multi-objective evidence construction in retrieval-enhanced question answering: from ranking to subset optimization

Retrieval-augmented question answering commonly constructs the final context by directly truncating an item-level candidate ranking, but this strategy does not explicitly control evidence redundancy, complementarity, or evidence-chain completeness. We propose MORSE ( M ulti- O bjective R etrieval S ubset E volution), a training-free post-retrieval method for selecting an evidence subset under a fixed context budget. MORSE constructs a question-adaptive search domain, organizes selected passages into functional core and support groups, and evaluates candidate subsets through three objectives: core-evidence quality, support complementarity, and package efficiency. NSGA-II searches for non-dominated evidence packages, while a conservative gate replaces the upstream ranking prefix only when protected quantities remain within predefined margins and at least one improvement is obtained. We evaluate MORSE on SQuAD v1.1, HotpotQA, and TriviaQA against direct Top-5 selection, cross-encoder reranking, MMR, submodular selection, and DPP. Compared with Top-5, MORSE yields higher mean F1 across all three datasets under both downstream QA models. In the SQuAD candidate source analysis, Hybrid+CE+MORSE obtains the strongest overall configuration, reaching 95.72 Recall and 88.54 MRR, with F1 scores of 69.96 for Qwen2.5:7B and 77.40 for Llama3.1:8B. The Hybrid Top-100 ablation identifies the core-quality pathway and conservative gate as the most influential components. Under the primary Top-100 setting, MORSE uses an average active search pool of 40.33 candidates and requires 120.76 ms per question. These results indicate that post-retrieval evidence-package optimization can complement upstream retrieval and reranking, particularly when useful complementary evidence remains distributed beyond the leading candidate prefix.

Authors

Institutions

Publication Details

Journal
Complex & Intelligent Systems
Published
2026-09-21
DOI
https://doi.org/10.1007/s40747-026-02520-z
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multi-objective evidence construction in retrieval-enhanced question answering: from ranking to subset optimization

Q.L. Ye, Chi Chen, Yan Wang
Complex & Intelligent Systems
Topic Modeling
article

Multi-objective evidence construction in retrieval-enhanced question answering: from ranking to subset optimization

Q.L. Ye, Chi Chen, Yan Wang
article en

Abstract

Retrieval-augmented question answering commonly constructs the final context by directly truncating an item-level candidate ranking, but this strategy does not explicitly control evidence redundancy, complementarity, or evidence-chain completeness. We propose MORSE ( M ulti- O bjective R etrieval S ubset E volution), a training-free post-retrieval method for selecting an evidence subset under a fixed context budget. MORSE constructs a question-adaptive search domain, organizes selected passages into functional core and support groups, and evaluates candidate subsets through three objectives: core-evidence quality, support complementarity, and package efficiency. NSGA-II searches for non-dominated evidence packages, while a conservative gate replaces the upstream ranking prefix only when protected quantities remain within predefined margins and at least one improvement is obtained. We evaluate MORSE on SQuAD v1.1, HotpotQA, and TriviaQA against direct Top-5 selection, cross-encoder reranking, MMR, submodular selection, and DPP. Compared with Top-5, MORSE yields higher mean F1 across all three datasets under both downstream QA models. In the SQuAD candidate source analysis, Hybrid+CE+MORSE obtains the strongest overall configuration, reaching 95.72 Recall and 88.54 MRR, with F1 scores of 69.96 for Qwen2.5:7B and 77.40 for Llama3.1:8B. The Hybrid Top-100 ablation identifies the core-quality pathway and conservative gate as the most influential components. Under the primary Top-100 setting, MORSE uses an average active search pool of 40.33 candidates and requires 120.76 ms per question. These results indicate that post-retrieval evidence-package optimization can complement upstream retrieval and reranking, particularly when useful complementary evidence remains distributed beyond the leading candidate prefix.

Complex & Intelligent Systems
Wenzhou University (CN), Xidian University (CN), Wenzhou Municipal Sci-Tech Bureau (CN)
Openalex Percentile: Top 8%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Multi-objective evidence construction in retrieval-enhanced question answering: from ranking to subset optimization — Q.L. Ye, Chi Chen, et al. · Complex & Intelligent Systems (2026) | TGRS Research Map | TGRS