Hybrid RLHF-AIF: Uncertainty-Guided Routing of Human and AI Feedback for Cost-Efficient Safety Alignment of Large Language Models
Safety alignment of large language models faces a feedback dilemma: human judgment is reliable but costly to scale; automated judgment is cheap but less trustworthy. Existing methods commit to one source per run, so misjudged responses are never corrected. This paper proposes Hybrid RLHF-AIF, which picks the feedback source for each sample online. First, a multi-label DistilBERT discriminator scores every response across 19 harm categories, reporting confidence through Monte Carlo dropout. Second, an uncertainty-aware scheduler sends only ambiguous or low-confidence responses to the reliable reviewer channel under a fixed budget and the rest to an artificial intelligence (AI) evaluator, recomputing the decision inside every Proximal Policy Optimization (PPO) iteration so that routing follows the policy. Collected annotations are replayed so the discriminator tracks the shifting policy. The reviewer channel is a stronger language model, not a real annotator, so the savings concern simulated human feedback. On PKU-SafeRLHF and HH-RLHF, over five seeds, the framework matches the safety of full simulated-human supervision on 10% of that budget while staying markedly more helpful and natural. Because that discriminator would otherwise judge its own objective, all final policies are re-scored with Llama Guard 3, and utility is re-judged by Claude Opus 5, a model from a different vendor; the ordering of methods survives both.
Authors
- Tianyu Sun (ORCID: https://orcid.org/0009-0001-3918-2329)
- Yanzhou Li (ORCID: https://orcid.org/0009-0008-2263-7383)
Institutions
- Dongguan University of Technology (CN)
- New York University (US)
Publication Details
- Journal
- Mathematics
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/math14193523
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00