Back close

Balancing Reward and Safety: An Analysis of Reinforcement Learning Agents

Publication Type : Conference Proceedings

Publisher : IEEE

Source : 2024 IEEE Conference on Engineering Informatics (ICEI)

Url : https://doi.org/10.1109/icei64305.2024.10912281

Campus : Coimbatore

School : School of Artificial Intelligence - Coimbatore

Year : 2024

Abstract : This work presents a comparative analysis of Trust Region Policy Optimization (TRPO), Proximal Policy Optimization (PPO), and Constrained Policy Optimization (CPO) algorithms for robotic path planning, using a panda robotic arm in the PyBullet simulation environment. The work aims to identify the most suitable reinforcement learning (RL) agent by evaluating each algorithm based on learning efficiency, safety adherence, and effectiveness in policy development. Safety metrics such as constraint violations, recovery time, and risk of unsafe actions were measured to assess safety adherence in various scenarios. Through rigorous simulation-based testing, this work explores key factors such as convergence rates, sample efficiency, safety performance, and generalization capabilities. Interestingly, while some agents demonstrate strong performance during training, they may struggle in application scenarios, revealing the potential disconnect between training success and practical application. This work also sheds light on the unique manifestation of overfitting in RL, where agents tend to rely heavily on learned actions, becoming vulnerable to minor environmental changes. This underscores the importance of balancing exploration and exploitation to avoid local optima and foster robust generalization. By offering a detailed comparison of these algorithms, the findings provide insights into the strengths and limitations of each, offering recommendations for selecting appropriate RL techniques for applications on robotic path planning, with a focus on safety and robustness in dynamic environments.

Cite this Research Publication : Udayagiri Varun, P S S Sai Keerthana, Komal Sai Anurag Pasumarthy, Sejal Singh, Mithun Kumar Kar, Balancing Reward and Safety: An Analysis of Reinforcement Learning Agents, 2024 IEEE Conference on Engineering Informatics (ICEI), IEEE, 2024, https://doi.org/10.1109/icei64305.2024.10912281

Admissions Apply Now