Model-Agnostic Hessian-Free Meta-policy Optimization via Zeroth-Order Estimation: A Linear Quadratic Regulator Perspective
摘要
Meta-learning offers a promising paradigm for control problems with uncertain and diverse dynamics, but its application is often hindered by the high computational cost and instability of estimating policy Hessians. To address this, we introduce a Hessian-free meta-learning algorithm for ergodic Linear Quadratic Regulator (LQR) tasks. By directly approximating the meta-policy gradient with zeroth-order information via Gaussian smoothing and Stein’s identity, the proposed method remains tractable and applicable to more complex adaptation schemes. Our core contribution is the development and rigorous analysis of this approach. We establish formal guarantees on the stability of the learned policies and prove the algorithm’s convergence to the global optimum, which are complemented by the sample complexity analysis that quantifies the trade-offs between estimation bias control and computational resources. Empirical validation on a modified Boeing system benchmark confirms our theoretical findings, showing that the proposed method achieves performance comparable to state-of-the-art Hessian-based approaches.