Dynamic obstacle avoidance and grasping planning for mobile robotic arm in complex environment based on improved TD3
摘要
To address the challenges of insufficient dynamic obstacle avoidance capability and limited grasping planning capabilities in complex environments for mobile robotic arms, this paper introduces a hybrid algorithm, COQNLS-TD3(GRU)-PER. This algorithm integrates the modified Constrained Optimization Quasi-Newton Least Squares (COQNLS) method with the enhanced Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, which is further augmented by incorporating the Gated Recurrent Unit (GRU) module and a Prioritized Experience Replay mechanism (PER). Initially, COQNLS planning is designed for the robotic arm to obtain the pre-grasp posture. Subsequently, the problem is defined as a Markov Decision Process involving the design of state space, action space, and reward-penalty functions. Subsequently, the GRU module is integrated into the neural network to process temporal state space features, thus enhancing dynamic obstacle avoidance capabilities. The planning process unfolds in phases: the dynamic obstacle avoidance phase is driven by training with TD3(GRU), and the target planning phase is steered by the pre-grasp posture, complemented by a cooperative guidance mechanism for transitional control outputs. To enhance training sample efficiency and accelerate algorithm convergence, a Prioritized Experience Replay mechanism is incorporated, facilitating the efficient development of an effective policy model for planning. Finally, to validate the efficacy of the proposed algorithm, a three-dimensional experimental scenario is designed for comparative experiments with other algorithms. Experimental results reveal that compared to the traditional TD3 algorithm, the COQNLS-TD3(GRU)-PER algorithm offers substantial improvements in both training efficiency and the control performance of the policy model.