First Passage Optimality for Continuous-Time Markov Decision Processes With Varying Discount Factors and History-Dependent Policies

Guo XP; Song XY; Zhang Y

IEEE Transactions on Automatic Control, Vol.59, No.1, 163-174, 2014

DOI10.1109/TAC.2013.2281475 Export Citation

First Passage Optimality for Continuous-Time Markov Decision Processes With Varying Discount Factors and History-Dependent Policies

Guo XP, Song XY, Zhang Y

This paper is an attempt to study the first passage optimality criterion for continuous-time Markov decision processes with state-dependent discount factors and history-dependent policies. The state space is denumerable, the action space is a Borel space, and the transition and reward rates are unbounded. Under suitable conditions, we show the existence of a deterministic stationary optimal policy, establish the Bellman (optimality) equation, to which the value function is the unique solution, and give the value and policy iteration algorithms for solving (at least approximating) the value function and an optimal policy. Furthermore, we give examples about reliability and controlled birth processes with killing to illustrate the potential applications of the results obtained here, and also to show the difference between the main results in this paper and those in the previous literature.

Keywords:Continuous-time Markov decision process;first passage criterion;varying discount factor