The Twin Delayed Deep Deterministic policy gradient (TD3) algorithm, a quintessential instance of reinforcement learning (RL) that employs Actor-Critic networks, has demonstrated exceptional performance in both training velocity and efficacy, making it extensively applicable to a spectrum of continuous control problems. In this paper, we have built a RL training environment tailored to the attitude control of spacecraft. By devising an apt reward function, the paper expedites the convergence of the algorithm, enhancing its overall performance. Furthermore, the discourse extends to the proficiency of the trained controllers in navigating a variety of test environments replete with disturbances, including attitude reorientation, impulsive torque disturbances, and loss of effectiveness faults. The trained controller exhibits a commendable degree of generalizability, underscoring their potential for real-world applications. The study advances the understanding of RL in spacecraft attitude control, and paves the way for the development of more sophisticated and adaptive control systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attitude Control of Spacecraft Based on Deep Reinforcement Learning TD3 Algorithm

  • Zhuoyue Peng,
  • Qiang Shen,
  • Chi Song

摘要

The Twin Delayed Deep Deterministic policy gradient (TD3) algorithm, a quintessential instance of reinforcement learning (RL) that employs Actor-Critic networks, has demonstrated exceptional performance in both training velocity and efficacy, making it extensively applicable to a spectrum of continuous control problems. In this paper, we have built a RL training environment tailored to the attitude control of spacecraft. By devising an apt reward function, the paper expedites the convergence of the algorithm, enhancing its overall performance. Furthermore, the discourse extends to the proficiency of the trained controllers in navigating a variety of test environments replete with disturbances, including attitude reorientation, impulsive torque disturbances, and loss of effectiveness faults. The trained controller exhibits a commendable degree of generalizability, underscoring their potential for real-world applications. The study advances the understanding of RL in spacecraft attitude control, and paves the way for the development of more sophisticated and adaptive control systems.