A distributional reinforcement learning model for temporal coagulation balance evaluation and state‑guided supportive suggestions
摘要
Bleeding and thrombosis represent a dynamic continuum in critically ill patients, where overlapping etiologies and heterogeneous pathophysiological mechanisms give rise to mixed-pattern coagulopathy. Conventional risk models usually focus on static, single-timepoint predictions and fail to adapt to rapidly evolving clinical states. Here, we retrospectively analyzed 2537 unique patients with 10,851 longitudinal records, constructing cycles based on clinical practice and engineering pharmacokinetic features. We then developed Adaptive Coagulopathy RL with BiLSTM-Attention (ACRLA), a distributional offline reinforcement learning framework designed to guide dynamic, individualized management of coagulopathy under laboratory-guided monitoring. Model performance was validated against clinician-adjudicated phenotypes, and sensitivity analyses to ensure robustness. The model showed stable convergence, achieving a positive per-patient test reward (0.077 ± 0.1651, 95% CI: 0.046–0.107) with a mean cumulative reward of 1.317. Model-derived tendencies aligned with clinical phenotypes, and feature importance identified D-dimer and platelet count as dominant drivers. The model’s therapeutic preference hierarchy mirrored clinical guidelines. These findings indicate that ACRLA effectively captures dynamic coagulopathy trajectories in a clinically interpretable state space and provides a data-driven tool for personalized, adaptive management of coagulopathy in critically ill patients.