<p>This work considers the offline evaluation problem for indefinite-horizon Markov Decision Processes. A minimax Q-function learning algorithm is proposed, which, instead of i.i.d. tuples <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10463_2025_924_Article_IEq1.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="73" /> </InlineMediaObject> <EquationSource Format="TEX">\((s,a,s',r)\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo stretchy="false">(</mo> <mi>s</mi> <mo>,</mo> <mi>a</mi> <mo>,</mo> <msup> <mi>s</mi> <mo>′</mo> </msup> <mo>,</mo> <mi>r</mi> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation>, evaluates undiscounted expected return based by i.i.d. trajectories truncated at a given time step. The confidence error bounds are developed. Experiments using Open AI’s Cart Pole environment are employed to demonstrate the algorithm.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Offline minimax Q-function learning for undiscounted indefinite-horizon MDPs

  • Fengying Li,
  • Yuqiang Li,
  • Xianyi Wu,
  • Wei Bai

摘要

This work considers the offline evaluation problem for indefinite-horizon Markov Decision Processes. A minimax Q-function learning algorithm is proposed, which, instead of i.i.d. tuples \((s,a,s',r)\) ( s , a , s , r ) , evaluates undiscounted expected return based by i.i.d. trajectories truncated at a given time step. The confidence error bounds are developed. Experiments using Open AI’s Cart Pole environment are employed to demonstrate the algorithm.