Rater State Bias in RLHF Preference Data: An Audit Framework
We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs. They may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, a rater's preferences may shift over time, so that preference data encodes rater state alongside judgments about response quality. We argue that, if present, such shifts would differ from random label noise. They could be correlated across annotators under shared conditions, and would not be guaranteed to cancel under aggregation. We propose rater state shift as a plausible, testable source of bias, and outline an audit framework for studying it. We do not infer the training history of any specific deployed model.
Authors
- Elena Kopteva (ORCID: https://orcid.org/0000-0001-8364-0481)
- Vitaliy Hlynianyi-Zhuk (ORCID: https://orcid.org/0009-0003-2635-0121)
Institutions
- University of Illinois Urbana-Champaign (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23023305
- Primary Topic
- Mobile Crowdsensing and Crowdsourcing
- Type
- preprint