Adaptive counter-pulsation control using a reinforcement learning framework in pulsatile ECMO: Implementation and evaluation
摘要
Pulsatile extracorporeal membrane oxygenation (p-ECMO) reduces left ventricular afterload through counter-pulsation (CP), but stable CP requires accurate cardiac-phase detection and rapid adaptation to heart-rate changes. Conventional phase-locked-loop (PLL) controllers show prolonged transition times and reduced responsiveness under rapid cardiac fluctuations. To overcome these limitations, a supervised policy-learning controller inspired by reinforcement-learning (RL) frameworks was developed. State features included HE/EH/HH/EE delays and RS-based confidence values extracted from blood-pressure waveforms in a mock circulation system. A rule-derived policy provided optimal training labels for the DNN controller. Real-time experiments with ±5 and ±10 bpm perturbations demonstrated rapid stabilization, significantly reducing transient times compared to the PLL method (6.61–10.54 s, p < 0.05). Furthermore, while the conventional PLL-based method showed fluctuating performance, the RL-CP framework maintained a significantly higher and more stable CP rate (89.28–94.70 %, p < 0.01) across all experimental conditions. These results indicate that RL-framework-based supervised policy learning enables robust and adaptive real-time CP control for p-ECMO, significantly outperforming PLL-based methods and suggesting strong potential for clinical translation.