PRECISE: Reducing the bias of LLM evaluations using prediction‐powered ranking estimation

Abstract Evaluating the quality of search systems traditionally requires a significant number of human relevance annotations. In recent times, several systems have explored the usage of Large Language Models (LLMs) as automated judges for this task while their inherent biases prevent direct use for metric estimation. We present a statistical framework extending Prediction‐Powered Inference (PPI) (Angelopoulos, Duchi, and Zrnic 2024) that combines minimal human annotations with LLM judgments to produce reliable estimates of metrics which require sub‐instance annotations. Our method requires as few as 100 human‐annotated queries and 10,000 unlabeled examples, reducing annotation requirements by significantly compared to traditional approaches. We formulate our proposed framework ( PRECISE ) for inference of relevance uplift for an LLM‐based query reformulation application, extending PPI to sub‐instance annotations at the query‐document level. By reformulating the metric‐integration space, we reduced the computational complexity from to , where represents corpus size (in order of millions). Detailed experiments across prominent retrieval datasets demonstrate that our method reduces the variance of estimates for the business‐critical Precision@ K metric, while effectively correcting for LLM bias in low‐resource settings.

Authors

Institutions

Publication Details

Journal
AI Magazine
Published
2026-10-08
DOI
https://doi.org/10.1002/aaai.70066
Primary Topic
Information Retrieval and Search Behavior
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

PRECISE: Reducing the bias of LLM evaluations using prediction‐powered ranking estimation

Anirban Majumder, Abhishek Divekar
AI Magazine
Information Retrieval and Search Behavior
article

PRECISE: Reducing the bias of LLM evaluations using prediction‐powered ranking estimation

Anirban Majumder, Abhishek Divekar
article en

Abstract

Abstract Evaluating the quality of search systems traditionally requires a significant number of human relevance annotations. In recent times, several systems have explored the usage of Large Language Models (LLMs) as automated judges for this task while their inherent biases prevent direct use for metric estimation. We present a statistical framework extending Prediction‐Powered Inference (PPI) (Angelopoulos, Duchi, and Zrnic 2024) that combines minimal human annotations with LLM judgments to produce reliable estimates of metrics which require sub‐instance annotations. Our method requires as few as 100 human‐annotated queries and 10,000 unlabeled examples, reducing annotation requirements by significantly compared to traditional approaches. We formulate our proposed framework ( PRECISE ) for inference of relevance uplift for an LLM‐based query reformulation application, extending PPI to sub‐instance annotations at the query‐document level. By reformulating the metric‐integration space, we reduced the computational complexity from to , where represents corpus size (in order of millions). Detailed experiments across prominent retrieval datasets demonstrate that our method reduces the variance of estimates for the business‐critical Precision@ K metric, while effectively correcting for LLM bias in low‐resource settings.

AI MagazineVol. 47(4)
Amazon (United States) (US)
Openalex Percentile: Top 96%
Information Retrieval and Search Behavior
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

PRECISE: Reducing the bias of LLM evaluations using prediction‐powered ranking estimation — Anirban Majumder, Abhishek Divekar · AI Magazine (2026) | TGRS Research Map | TGRS