Knowing How Many to Show: Training-Free, Evidence-Gated Result Sets for Semantic Search over Any Content, from Text Collections to On-Device Media

Semantic search returns a ranked list and leaves the cutoff to the application, which almost always shows a fixed top-k. A fixed top-k can never answer "nothing found", shows the same number of items whether a query has one answer or twenty, and, when its output is handed to a language model in retrieval-augmented generation, supplies unrelated passages as evidence. Methods that learn where to cut a ranked list need relevance judgments or click logs, which do not exist in the contexts where private search matters most: a person's device, an organization's on-premises document store, a newly launched service, or the context-selection step of a local assistant. We present a retrieval core for this label-free regime. Every item is turned into language (its own text, or a description written by a local model) and stored as several vectors, one for the whole item and one per phrase, alongside a keyword index. A staged evidence-gated rule then decides, per query, whether to return anything and how many items: every threshold is anchored to the embedding model's noise floor and derived from the semantic score alone, while keyword evidence may reorder only items the embedding already judges related, and otherwise corroborates, tightens or vouches for a result but never raises a threshold. The rule consumes only scores and lexical statistics, so it is independent of content type and of who consumes the result. On three unseen BEIR collections (science, consumer medicine, finance), with every system's settings chosen on three other collections and frozen, it raises a task score that credits correct lists and honest empty answers by +8.7 to +28.0 points over fixed top-k cutoffs and by +1.7 to +13.4 over the strongest adaptive baseline, a similarity threshold within the top five (clearly on two collections, marginally on the third). Both return nothing for nearly every unanswerable question, but the threshold also returns nothing for 8.5–72% of answerable questions, against 0.6–22% for the evidence-gated rule. Used unchanged as the search engine of an Android application over 600 photos, photographed documents and screenshots, it scores 63.1 against 55.7 for the strongest CLIP-family model at any deployable cutoff, entirely on a commodity tablet. The implementation is released as an open-source library with adapters for any content source, description model and embedding model.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-09
DOI
https://doi.org/10.5281/zenodo.23270662
Primary Topic
Information Retrieval and Search Behavior
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Knowing How Many to Show: Training-Free, Evidence-Gated Result Sets for Semantic Search over Any Content, from Text Collections to On-Device Media

Mehran Iranpour
Zenodo (CERN European Organization for Nuclear Research)
Information Retrieval and Search Behavior
preprint

Knowing How Many to Show: Training-Free, Evidence-Gated Result Sets for Semantic Search over Any Content, from Text Collections to On-Device Media

Mehran Iranpour
preprint en

Abstract

Semantic search returns a ranked list and leaves the cutoff to the application, which almost always shows a fixed top-k. A fixed top-k can never answer "nothing found", shows the same number of items whether a query has one answer or twenty, and, when its output is handed to a language model in retrieval-augmented generation, supplies unrelated passages as evidence. Methods that learn where to cut a ranked list need relevance judgments or click logs, which do not exist in the contexts where private search matters most: a person's device, an organization's on-premises document store, a newly launched service, or the context-selection step of a local assistant. We present a retrieval core for this label-free regime. Every item is turned into language (its own text, or a description written by a local model) and stored as several vectors, one for the whole item and one per phrase, alongside a keyword index. A staged evidence-gated rule then decides, per query, whether to return anything and how many items: every threshold is anchored to the embedding model's noise floor and derived from the semantic score alone, while keyword evidence may reorder only items the embedding already judges related, and otherwise corroborates, tightens or vouches for a result but never raises a threshold. The rule consumes only scores and lexical statistics, so it is independent of content type and of who consumes the result. On three unseen BEIR collections (science, consumer medicine, finance), with every system's settings chosen on three other collections and frozen, it raises a task score that credits correct lists and honest empty answers by +8.7 to +28.0 points over fixed top-k cutoffs and by +1.7 to +13.4 over the strongest adaptive baseline, a similarity threshold within the top five (clearly on two collections, marginally on the third). Both return nothing for nearly every unanswerable question, but the threshold also returns nothing for 8.5–72% of answerable questions, against 0.6–22% for the evidence-gated rule. Used unchanged as the search engine of an Android application over 600 photos, photographed documents and screenshots, it scores 63.1 against 55.7 for the strongest CLIP-family model at any deployable cutoff, entirely on a commodity tablet. The implementation is released as an open-source library with adapters for any content source, description model and embedding model.

Zenodo (CERN European Organization for Nuclear Research)
Information Retrieval and Search Behavior
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.