<p>Meta-learning offers a promising paradigm for control problems with uncertain and diverse dynamics, but its application is often hindered by the high computational cost and instability of estimating policy Hessians. To address this, we introduce a Hessian-free meta-learning algorithm for ergodic Linear Quadratic Regulator (LQR) tasks. By directly approximating the meta-policy gradient with zeroth-order information via Gaussian smoothing and Stein’s identity, the proposed method remains tractable and applicable to more complex adaptation schemes. Our core contribution is the development and rigorous analysis of this approach. We establish formal guarantees on the stability of the learned policies and prove the algorithm’s convergence to the global optimum, which are complemented by the sample complexity analysis that quantifies the trade-offs between estimation bias control and computational resources. Empirical validation on a modified Boeing system benchmark confirms our theoretical findings, showing that the proposed method achieves performance comparable to state-of-the-art Hessian-based approaches.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Model-Agnostic Hessian-Free Meta-policy Optimization via Zeroth-Order Estimation: A Linear Quadratic Regulator Perspective

  • Yunian Pan,
  • Tao Li,
  • Quanyan Zhu

摘要

Meta-learning offers a promising paradigm for control problems with uncertain and diverse dynamics, but its application is often hindered by the high computational cost and instability of estimating policy Hessians. To address this, we introduce a Hessian-free meta-learning algorithm for ergodic Linear Quadratic Regulator (LQR) tasks. By directly approximating the meta-policy gradient with zeroth-order information via Gaussian smoothing and Stein’s identity, the proposed method remains tractable and applicable to more complex adaptation schemes. Our core contribution is the development and rigorous analysis of this approach. We establish formal guarantees on the stability of the learned policies and prove the algorithm’s convergence to the global optimum, which are complemented by the sample complexity analysis that quantifies the trade-offs between estimation bias control and computational resources. Empirical validation on a modified Boeing system benchmark confirms our theoretical findings, showing that the proposed method achieves performance comparable to state-of-the-art Hessian-based approaches.