A Reinforcement Learning Framework for Optimizing Kidney Allocation for Transplant Based on Survival and Ethical Criteria
摘要
We present a novel reinforcement learning (RL)–based approach for determining organ allocation for transplant to the most optimal candidate by adaptively balancing a number of utility and ethical criteria. Traditional deterministic policies—such as First-Come, First-Served (FCFS), Utility-First (UF), and Benefit-First (BF)—have well-known limitations in responding to day-to-day changes in candidate status, while other advanced methods lack a mechanism for dynamically prioritizing high-risk candidates. Our RL framework formulates the allocation process as a Markov decision problem in which each day’s incoming organ and the waitlist composition define a state, and the allocation choice defines the action. A reward function encompasses net benefit minus penalties for extended waiting, with a machine learning model predicting post-transplant survival for each donor-recipient pair. The RL agent learns to make optimal daily allocation decisions that reduce waitlist deaths and improve average post-transplant outcomes. Simulations on 5,000 donor-recipient pairs derived from SRTR data demonstrate that the RL-based policy consistently outperforms deterministic baselines in metrics such as death rate while waiting for an organ, average wait time, average post-transplant survival, and average net benefit. Our findings suggest that an RL-driven paradigm has the potential to increase both the survival gains and justice during kidney allocation in settings with limited organ supply.