Reduction of a Markov decision process with non-linear discounting to a stochastic game with standard total undiscounted criterion
摘要
We consider a Markov decision process (MDP), whose total discounted utility is aggregated recursively with a concave discount function that is not necessarily linear. The state and action spaces are Borel spaces, and the utility function is nonnegative. We show that it can be reduced to a turn-based stochastic game model with the total undiscounted utility. This reduction result is then applied to the MDP problem with recursively aggregated utility to be maximized or cost to be minimized.