As learning-enabled Cyber-Physical Systems (CPSs) are increasingly used in safety-critical settings, there is a growing need to ensure their safety. For example, to tackle the problem of rate-adaptive pacemakers which correct Sinus Node Dysfunction, a Reinforcement Learning (RL) approach may be used to mimic the natural pacing rhythm of the heart. However, this is currently not done and there are no known approaches to ensure the safety of combining RL with conventional pacing algorithms. While there is growing interest on ensuring the safety of AI-enabled CPS, the issue of safe RL for CPS, using light-weight formal methods, has drawn scant attention. Therefore, we present an approach which combines Runtime Enforcement with RL. To guarantee safety throughout the RL agent’s learning and execution stages, an enforcer is constructed from a set of safety policies expressed using a variant of timed automata. The RL agent’s outputs are observed by the enforcer, which ensures that only safe actions are delivered to the environment by correcting the outputs which would violate the safety policies. In order to evaluate the proposed approach, we have implemented a rate-adaptive pacemaker which learns the natural pacing rhythm through an RL agent, allowing it to pace appropriately during disease. This model is executed in closed-loop with a real-time heart model where various diseases can be exhibited in order to test its efficacy. Furthermore, we illustrate the benefits of this system by contrasting it with traditional pacing techniques and RL without the use of an enforcer.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Formal Approach for Safe Reinforcement Learning: A Rate-Adaptive Pacemaker Case Study

  • Sai Rohan Harshavardhan Vuppala,
  • Nathan Allen,
  • Srinivas Pinisetty,
  • Partha Roop

摘要

As learning-enabled Cyber-Physical Systems (CPSs) are increasingly used in safety-critical settings, there is a growing need to ensure their safety. For example, to tackle the problem of rate-adaptive pacemakers which correct Sinus Node Dysfunction, a Reinforcement Learning (RL) approach may be used to mimic the natural pacing rhythm of the heart. However, this is currently not done and there are no known approaches to ensure the safety of combining RL with conventional pacing algorithms. While there is growing interest on ensuring the safety of AI-enabled CPS, the issue of safe RL for CPS, using light-weight formal methods, has drawn scant attention. Therefore, we present an approach which combines Runtime Enforcement with RL. To guarantee safety throughout the RL agent’s learning and execution stages, an enforcer is constructed from a set of safety policies expressed using a variant of timed automata. The RL agent’s outputs are observed by the enforcer, which ensures that only safe actions are delivered to the environment by correcting the outputs which would violate the safety policies. In order to evaluate the proposed approach, we have implemented a rate-adaptive pacemaker which learns the natural pacing rhythm through an RL agent, allowing it to pace appropriately during disease. This model is executed in closed-loop with a real-time heart model where various diseases can be exhibited in order to test its efficacy. Furthermore, we illustrate the benefits of this system by contrasting it with traditional pacing techniques and RL without the use of an enforcer.