Tianyu Sun, Yanzhou Li · Mathematics 2026 · 2026
DOI: 10.3390/math14193523
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Safety alignment of large language models faces a feedback dilemma: human judgment is reliable but costly to scale; automated judgment is cheap but less trustworthy. Existing methods commit to one source per run, so misjudged responses are never corrected. This paper proposes Hybrid RLHF-AIF, which picks the feedback source for each sample online. First, a multi-label DistilBERT discriminator scores every response across 19 harm categories, reporting confidence through Monte Carlo dropout. Second, an uncertainty-aware scheduler sends only ambiguous or low-confidence responses to the reliable reviewer channel under a fixed budget and the rest to an artificial intelligence (AI) evaluator, recomputing the decision inside every Proximal Policy Optimization (PPO) iteration so that routing follows the policy. Collected annotations are replayed so the discriminator tracks the shifting policy. The reviewer channel is a stronger language model, not a real annotator, so the savings concern simulated human feedback. On PKU-SafeRLHF and HH-RLHF, over five seeds, the framework matches the safety of full simulated-human supervision on 10% of that budget while staying markedly more helpful and natural. Because that discriminator would otherwise judge its own objective, all final policies are re-scored with Llama Guard 3, and utility is re-judged by Claude Opus 5, a model from a different vendor; the ordering of methods survives both.
No comments yet — start the discussion below.