Reinforcement-Learning-Based Online PID Gain Adaptation for Fault-Tolerant Quadrotor Control Under Rotor Loss-of-Effectiveness
摘要
Actuator loss-of-effectiveness (LoE) in a single rotor reduces control authority and can destabilize fixed-gain quadrotor controllers. This paper investigates a bounded online gain-adaptation layer, implemented via reinforcement learning, for a standard cascaded proportional–integral–derivative (PID) quadrotor controller. The emphasis is not RL-based PID tuning per se, but a practical integration in which four decentralized agents update roll, pitch, yaw, and altitude PID gains in real time within preset bounds while the underlying mixer, outer-loop structure, and rigid-body model remain unchanged. A six-degree-of-freedom Newton–Euler quadrotor model with first-order motor dynamics is implemented in MATLAB/Simulink, and a Soft Actor-Critic (SAC) instantiation of the adaptation layer is evaluated under nominal flight and multiple transient and sustained single-rotor LoE profiles. For context, Twin Delayed Deep Deterministic Policy Gradient (TD3) and Proximal Policy Optimization (PPO) are evaluated under the same observation/action definitions, reward structure, gain limits, and 500-episode training budget; an additional 1000-episode PPO check is included to assess training-budget sensitivity. Offline-optimized fixed-gain PID baselines obtained using particle swarm optimization (PSO) and grey wolf optimization (GWO) are included to separate the benefit of online adaptation from static tuning and to address sensitivity to the selected metaheuristic baseline. Step responses and three-dimensional trajectory-tracking simulations show comparable nominal accuracy across methods but clear differences under degradation: online gain adaptation yields smaller post-fault deviations than fixed gains, while SAC provides the most consistent fault accommodation among the evaluated configurations. In the tested setup, these results support bounded online gain adaptation as a low-integration-cost way to improve LoE robustness without replacing a familiar PID-based quadrotor control architecture.