Rater State Bias in RLHF Preference Data: An Audit Framework

We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs. They may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, a rater's preferences may shift over time, so that preference data encodes rater state alongside judgments about response quality. We argue that, if present, such shifts would differ from random label noise. They could be correlated across annotators under shared conditions, and would not be guaranteed to cancel under aggregation. We propose rater state shift as a plausible, testable source of bias, and outline an audit framework for studying it. We do not infer the training history of any specific deployed model.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23023305
Primary Topic
Mobile Crowdsensing and Crowdsourcing
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Rater State Bias in RLHF Preference Data: An Audit Framework

Elena Kopteva, Vitaliy Hlynianyi-Zhuk
Zenodo (CERN European Organization for Nuclear Research)
Mobile Crowdsensing and Crowdsourcing
preprint

Rater State Bias in RLHF Preference Data: An Audit Framework

Elena Kopteva, Vitaliy Hlynianyi-Zhuk
preprint en

Abstract

We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs. They may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, a rater's preferences may shift over time, so that preference data encodes rater state alongside judgments about response quality. We argue that, if present, such shifts would differ from random label noise. They could be correlated across annotators under shared conditions, and would not be guaranteed to cancel under aggregation. We propose rater state shift as a plausible, testable source of bias, and outline an audit framework for studying it. We do not infer the training history of any specific deployed model.

Zenodo (CERN European Organization for Nuclear Research)
University of Illinois Urbana-Champaign (US)
Mobile Crowdsensing and Crowdsourcing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Rater State Bias in RLHF Preference Data: An Audit Framework — Elena Kopteva, Vitaliy Hlynianyi-Zhuk · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS