Data-Driven Shielding of Online Reinforcement Learning: A Stormwater Pond Case Study
摘要
To synthesize a safe and optimal controller for switched hybrid systems, one can first synthesize a shield that ensures safety, and then apply reinforcement learning within the constraints of the shield to obtain the desired controller. However, developing such a shield for switched hybrid systems typically requires a full model of the environment, which is not always available. Instead, historical data of the environment might be available. In this paper, we introduce a method for the construction of safety shields based on different scenarios captured in historical data. We show how individual shields for different scenarios can be combined to obtain a single shield that is provably safe within the bounds of the observed scenarios. We demonstrate the method using an industrial case study of a stormwater detention pond, which includes ten years of historical data of different rain events/scenarios. Our experimental results show that the shielded optimal controller ensures safety across all individual historical rain scenarios compared to the unshielded optimal controller. Additionally, we empirically show that the shield may also generalize for scenarios not covered by the historical data.