Hybrid RLHF-AIF: Uncertainty-Guided Routing of Human and AI Feedback for Cost-Efficient Safety Alignment of Large Language Models

Safety alignment of large language models faces a feedback dilemma: human judgment is reliable but costly to scale; automated judgment is cheap but less trustworthy. Existing methods commit to one source per run, so misjudged responses are never corrected. This paper proposes Hybrid RLHF-AIF, which picks the feedback source for each sample online. First, a multi-label DistilBERT discriminator scores every response across 19 harm categories, reporting confidence through Monte Carlo dropout. Second, an uncertainty-aware scheduler sends only ambiguous or low-confidence responses to the reliable reviewer channel under a fixed budget and the rest to an artificial intelligence (AI) evaluator, recomputing the decision inside every Proximal Policy Optimization (PPO) iteration so that routing follows the policy. Collected annotations are replayed so the discriminator tracks the shifting policy. The reviewer channel is a stronger language model, not a real annotator, so the savings concern simulated human feedback. On PKU-SafeRLHF and HH-RLHF, over five seeds, the framework matches the safety of full simulated-human supervision on 10% of that budget while staying markedly more helpful and natural. Because that discriminator would otherwise judge its own objective, all final policies are re-scored with Llama Guard 3, and utility is re-judged by Claude Opus 5, a model from a different vendor; the ordering of methods survives both.

Authors

Institutions

Publication Details

Journal
Mathematics
Published
2026-09-28
DOI
https://doi.org/10.3390/math14193523
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Hybrid RLHF-AIF: Uncertainty-Guided Routing of Human and AI Feedback for Cost-Efficient Safety Alignment of Large Language Models

Tianyu Sun, Yanzhou Li
Mathematics
Adversarial Robustness in Machine Learning
article

Hybrid RLHF-AIF: Uncertainty-Guided Routing of Human and AI Feedback for Cost-Efficient Safety Alignment of Large Language Models

Tianyu Sun, Yanzhou Li
article en

Abstract

Safety alignment of large language models faces a feedback dilemma: human judgment is reliable but costly to scale; automated judgment is cheap but less trustworthy. Existing methods commit to one source per run, so misjudged responses are never corrected. This paper proposes Hybrid RLHF-AIF, which picks the feedback source for each sample online. First, a multi-label DistilBERT discriminator scores every response across 19 harm categories, reporting confidence through Monte Carlo dropout. Second, an uncertainty-aware scheduler sends only ambiguous or low-confidence responses to the reliable reviewer channel under a fixed budget and the rest to an artificial intelligence (AI) evaluator, recomputing the decision inside every Proximal Policy Optimization (PPO) iteration so that routing follows the policy. Collected annotations are replayed so the discriminator tracks the shifting policy. The reviewer channel is a stronger language model, not a real annotator, so the savings concern simulated human feedback. On PKU-SafeRLHF and HH-RLHF, over five seeds, the framework matches the safety of full simulated-human supervision on 10% of that budget while staying markedly more helpful and natural. Because that discriminator would otherwise judge its own objective, all final policies are re-scored with Llama Guard 3, and utility is re-judged by Claude Opus 5, a model from a different vendor; the ordering of methods survives both.

MathematicsVol. 14(19)
Dongguan University of Technology (CN), New York University (US)
Reduced inequalities
Openalex Percentile: Top 9%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.