Two Time Scale Partial Unknown Dynamics System Tracking Control Based on Off-Policy Inverse Reinforcement Learning
摘要
This article, integrating the singular perturbation technique with inverse reinforcement learning, proposes a novel linear two time scale system tracking control method grounded in off-policy inverse reinforcement learning. This method addresses the challenge of unknown cost functions prevalent in industrial processes. First, the singular perturbation method is leveraged to decompose the original problem into fast and slow subsystem issues. Without manually designing a cost function, the method learns from known optimal behavioral data by reconstructing cost functions tailored to each subsystem, enabling the system to mimic optimal behaviors. Then, for the fast time scale system, a model-based inverse reinforcement learning method is adopted, while for the slow time scale system, a model-free off-policy inverse reinforcement learning strategy is employed, which reconstructs the system’s cost function solely using measured expert behavioral data inputs. Finally, using a mixed separation thickening industrial process to illustrate the effectiveness of this method in two time scale tracking.