<p>Sharpe ratio (SR) is a critical parameter in characterizing financial time series as it jointly considers the reward and the volatility of any stock/portfolio through its mean and standard deviation. Deriving online algorithms for optimizing the SR is particularly challenging since even offline policies experience constant regret with respect to the best expert (Even-Dar et al., <CitationRef CitationID="CR11">2006</CitationRef>). This paper focuses on optimizing the regularized square SR (RSSR) by considering two settings: regret minimization (RM) and best arm identification (BAI). In this regard, we propose a novel multiarmed bandit (MAB) algorithm for RM called <Emphasis FontCategory="NonProportional">UCB-RSSR</Emphasis> for RSSR maximization. We derive a path-dependent concentration bound for the estimate of the RSSR. Based on that, we derive the regret guarantees of <Emphasis FontCategory="NonProportional">UCB-RSSR</Emphasis> and show that it evolves as <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10994_2024_6680_Article_IEq1.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="64" /> </InlineMediaObject> <EquationSource Format="TEX">\({\mathcal {O}}\left( \log {n}\right)\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi mathvariant="script">O</mi> <mfenced close=")" open="("> <mo>log</mo> <mi>n</mi> </mfenced> </mrow> </math></EquationSource> </InlineEquation> for the two-armed bandit case played for a horizon <i>n</i>. We also consider algorithms for the fixed budget setting of the BAI problems, i.e., sequential halving and successive rejects, and propose <Emphasis FontCategory="NonProportional">SHSR</Emphasis> and <Emphasis FontCategory="NonProportional">SuRSR</Emphasis> algorithms. We derive the upper bound for the error probability of BAI algorithms. We demonstrate that <Emphasis FontCategory="NonProportional">UCB-RSSR</Emphasis> outperforms the only other known SR optimizing bandit algorithm, <Emphasis FontCategory="NonProportional">U-UCB</Emphasis> (Cassel et al., <CitationRef CitationID="CR9">2023</CitationRef>). We also study the efficacy of proposed BAI algorithms for 6 different setups and discuss the cases where our proposed algorithms are suitable. Our research highlights that our proposed algorithms will find extensive applications in risk-aware portfolio management problems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing sharpe ratio: risk-adjusted decision-making in multi-armed bandits

  • Sabrina Khurshid,
  • Mohammed Shahid Abdulla,
  • Gourab Ghatak

摘要

Sharpe ratio (SR) is a critical parameter in characterizing financial time series as it jointly considers the reward and the volatility of any stock/portfolio through its mean and standard deviation. Deriving online algorithms for optimizing the SR is particularly challenging since even offline policies experience constant regret with respect to the best expert (Even-Dar et al., 2006). This paper focuses on optimizing the regularized square SR (RSSR) by considering two settings: regret minimization (RM) and best arm identification (BAI). In this regard, we propose a novel multiarmed bandit (MAB) algorithm for RM called UCB-RSSR for RSSR maximization. We derive a path-dependent concentration bound for the estimate of the RSSR. Based on that, we derive the regret guarantees of UCB-RSSR and show that it evolves as \({\mathcal {O}}\left( \log {n}\right)\) O log n for the two-armed bandit case played for a horizon n. We also consider algorithms for the fixed budget setting of the BAI problems, i.e., sequential halving and successive rejects, and propose SHSR and SuRSR algorithms. We derive the upper bound for the error probability of BAI algorithms. We demonstrate that UCB-RSSR outperforms the only other known SR optimizing bandit algorithm, U-UCB (Cassel et al., 2023). We also study the efficacy of proposed BAI algorithms for 6 different setups and discuss the cases where our proposed algorithms are suitable. Our research highlights that our proposed algorithms will find extensive applications in risk-aware portfolio management problems.