PPO-based Reinforcement Learning with Human Feedback with Hybrid Oversight and Predictive Reward Evaluation for AGI

Authors

DOI:

https://doi.org/10.62411/faith.3048-3719-276

Keywords:

Artificial General Intelligence, AGI, Human-in-the-Loop Learning, Predictive Reward Evaluation, Reinforcement Learning with Human Feedback, Reward Modeling, RLHF, Safe AI, Value Alignment

Abstract

The pursuit of Artificial General Intelligence (AGI) requires learning frameworks that not only optimize task performance but also align with complex human values. Reinforcement Learning with Human Feedback (RLHF) has emerged as a promising approach to address this challenge; however, conventional RLHF pipelines face scalability issues, reward-model brittleness, and safety concerns. In this study, we propose a PPO-based RLHF framework enhanced with hybrid human–AI oversight and predictive reward evaluation metrics. The framework integrates human annotations with AI-generated critiques, improving data efficiency and robustness against reward hacking. Experimental evaluations on language alignment and control benchmarks demonstrate that the proposed approach achieves a preference win-rate of 78% (vs. 65% in standard RLHF and 54% in supervised fine-tuning), improves task accuracy to 83% (a 12% increase over Sparrow), and reduces safety violations by 31% compared to baseline RLHF. Furthermore, the hybrid oversight strategy enhanced sample efficiency by 1.5×, reducing overall training episodes and annotation costs. These results confirm that the proposed method significantly improves alignment, efficiency, and safety, positioning RLHF with hybrid oversight and predictive evaluation as a practical substrate for advancing safe and scalable AGI systems.

Downloads

Download data is not yet available.

Author Biography

Atul Sharma, Independent Researcher

Independent Researcher, Kudus 421312, Maharashtra, India

References

R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction. MIT Press, 2018.

V. Mnih et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, Feb. 2015, doi: 10.1038/nature14236.

D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané, “Concrete Problems in AI Safety,” ArXiv. Jul. 25, 2016. [Online]. Available: http://arxiv.org/abs/1606.06565

J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel, “Trust Region Policy Optimization,” ArXiv. Apr. 20, 2017. [Online]. Available: http://arxiv.org/abs/1502.05477

L. Ouyang et al., “Training language models to follow instructions with human feedback,” ArXiv. Mar. 04, 2022. [Online]. Available: http://arxiv.org/abs/2203.02155

A. Glaese et al., “Improving alignment of dialogue agents via targeted human judgements,” ArXiv. Sep. 28, 2022. [Online]. Available: http://arxiv.org/abs/2209.14375

J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal Policy Optimization Algorithms,” ArXiv. Aug. 28, 2017. [Online]. Available: http://arxiv.org/abs/1707.06347

Y. Bai et al., “Constitutional AI: Harmlessness from AI Feedback,” ArXiv. Dec. 15, 2022. [Online]. Available: http://arxiv.org/abs/2212.08073

H. Lee et al., “RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback,” ArXiv. Sep. 03, 2024. [Online]. Available: http://arxiv.org/abs/2309.00267

E. Frick et al., “How to Evaluate Reward Models for RLHF,” ArXiv. Oct. 22, 2024. [Online]. Available: http://arxiv.org/abs/2410.14872

L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement Learning: A Survey,” Arxiv. 1996. doi: 10.48550/arXiv.cs/9605103.

R. S. S. Sutton and A. G. G. Barto, “Reinforcement Learning: An Introduction,” IEEE Trans. Neural Networks, vol. 9, no. 5, pp. 1054–1054, Sep. 1998, doi: 10.1109/TNN.1998.712192.

D. Silver et al., “Mastering the game of Go with deep neural networks and tree search,” Nature, vol. 529, no. 7587, pp. 484–489, Jan. 2016, doi: 10.1038/nature16961.

S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-End Training of Deep Visuomotor Policies,” J. Mach. Learn. Res., vol. 17, pp. 1–40, Apr. 2016, [Online]. Available: http://arxiv.org/abs/1504.00702

S. J. Russell, Human Compatible: Artificial Intelligence and the Problem of Control. Viking, 2019.

P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Twenty-first international conference on Machine learning - ICML ’04, 2004, p. 1. doi: 10.1145/1015330.1015430.

J. Ho and S. Ermon, “Generative Adversarial Imitation Learning,” ArXiv. Jun. 10, 2016. [Online]. Available: http://arxiv.org/abs/1606.03476

P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” ArXiv. Feb. 17, 2023. [Online]. Available: http://arxiv.org/abs/1706.03741

G. Irving, P. Christiano, and D. Amodei, “AI safety via debate,” arXiv. Oct. 22, 2018. [Online]. Available: http://arxiv.org/abs/1805.00899

T. Korbak et al., “Pretraining Language Models with Human Preferences,” in ICML’23: Proceedings of the 40th International Conference on Machine Learning, Jun. 2023, pp. 17506–17533. [Online]. Available: http://arxiv.org/abs/2302.08582

Y. Hu et al., “Towards Comprehensive Preference Data Collection for Reward Modeling,” ArXiv. Jun. 24, 2024. [Online]. Available: http://arxiv.org/abs/2406.16486

Downloads

Published

2025-10-24

How to Cite

[1]
A. Sharma, “PPO-based Reinforcement Learning with Human Feedback with Hybrid Oversight and Predictive Reward Evaluation for AGI”, J. Fut. Artif. Intell. Tech., vol. 2, no. 3, pp. 493–503, Oct. 2025.

Similar Articles

<< < 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.