<p>This paper proposes a new multi-objective reinforcement learning (MORL) algorithm for robotics by extending policy improvement with path integral (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10015_2025_1027_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="24" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {PI}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>PI</mtext> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation>) algorithm. For a robot motion acquisition problem, most existing MORL algorithms are hard to apply, because of the high-dimensional and continuous state and action spaces. However, policy-based algorithms such as <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10015_2025_1027_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="24" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {PI}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>PI</mtext> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> can be applied to solve this problem in single-objective cases. Based on the similarity of <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10015_2025_1027_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="24" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {PI}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>PI</mtext> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> and evolution strategies (ESs) and the fact that ESs are well-suited for multi-objective optimization, we propose an extension of <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10015_2025_1027_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="24" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {PI}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>PI</mtext> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> and some techniques to speed up the learning. The effectiveness is shown via numerical simulations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-objective path integral policy improvement for learning robotic motion

  • Hayato Sago,
  • Ryo Ariizumi,
  • Toru Asai,
  • Shun-ichi Azuma

摘要

This paper proposes a new multi-objective reinforcement learning (MORL) algorithm for robotics by extending policy improvement with path integral ( \(\text {PI}^2\) PI 2 ) algorithm. For a robot motion acquisition problem, most existing MORL algorithms are hard to apply, because of the high-dimensional and continuous state and action spaces. However, policy-based algorithms such as \(\text {PI}^2\) PI 2 can be applied to solve this problem in single-objective cases. Based on the similarity of \(\text {PI}^2\) PI 2 and evolution strategies (ESs) and the fact that ESs are well-suited for multi-objective optimization, we propose an extension of \(\text {PI}^2\) PI 2 and some techniques to speed up the learning. The effectiveness is shown via numerical simulations.