In recent years, the rapid development of deep reinforcement learning has provided a new way to solve the robot control problems. However, the low sample efficiency and slow convergence speed of deep reinforcement learning have become one of the obstacles when transitioning from simulation to the real world. In this paper, we propose a quadrotor control policy using Gaussian ensemble model-based reinforcement learning. Unlike traditional control methods, this method uses an actor-critic deep neural network which is updated with a reward function to achieve end-to-end control of the quadrotor by establishing a mapping between the quadrotor's states and motor control signals. Additionally, we improve sample efficiency by constructing an ensemble model following a Gaussian normal distribution, which differs from conventional model-free RL methods. The environment model is trained using data from the agent's interaction with the real environment and reduces the number of interactions with the real environment by generating simulated data. The approach is evaluated in the AirSim which is a high-fidelity visual and physical simulator. The results show that the proposed approach improves the sample efficiency, eliminates oscillations and steady error, and demonstrates robustness to external disturbances.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

End-To-End Control of a Quadrotor Using Gaussian Ensemble Model-Based Reinforcement Learning

  • Qiwen Zheng,
  • Qingyuan Xia,
  • Haonan Luo,
  • Bohai Deng,
  • Shengwei Li

摘要

In recent years, the rapid development of deep reinforcement learning has provided a new way to solve the robot control problems. However, the low sample efficiency and slow convergence speed of deep reinforcement learning have become one of the obstacles when transitioning from simulation to the real world. In this paper, we propose a quadrotor control policy using Gaussian ensemble model-based reinforcement learning. Unlike traditional control methods, this method uses an actor-critic deep neural network which is updated with a reward function to achieve end-to-end control of the quadrotor by establishing a mapping between the quadrotor's states and motor control signals. Additionally, we improve sample efficiency by constructing an ensemble model following a Gaussian normal distribution, which differs from conventional model-free RL methods. The environment model is trained using data from the agent's interaction with the real environment and reduces the number of interactions with the real environment by generating simulated data. The approach is evaluated in the AirSim which is a high-fidelity visual and physical simulator. The results show that the proposed approach improves the sample efficiency, eliminates oscillations and steady error, and demonstrates robustness to external disturbances.