Sufficiency of Markov Policies for Continuous-Time Jump Markov Decision Processes

被引:4
|
作者
Feinberg, Eugene A. [1 ]
Mandava, Manasa [2 ]
Shiryaev, Albert N. [3 ]
机构
[1] SUNY Stony Brook, Dept Appl Math & Stat, Stony Brook, NY 11794 USA
[2] Indian Sch Business, Hyderabad 500032, India
[3] Steklov Math Inst, Dept Probabil Theory & Math Stat, Moscow 119991, Russia
关键词
continuous-time jump Markov process; Borel; state; action; Markov policy; COUNTABLE STATE; MODELS;
D O I
10.1287/moor.2021.1169
中图分类号
C93 [管理学]; O22 [运筹学];
学科分类号
070105 ; 12 ; 1201 ; 1202 ; 120202 ;
摘要
One of the basic facts known for discrete-time Markov decision processes is that, if the probability distribution of an initial state is fixed, then for every policy it is easy to construct a (randomized) Markov policy with the same marginal distributions of state-action pairs as for the original policy. This equality of marginal distributions implies that the values of major objective criteria, including expected discounted total costs and average rewards per unit time, are equal for these two policies. This paper investigates the validity of the similar fact for continuous-time jump Markov decision processes (CTJMDPs). It is shown in this paper that the equality of marginal distributions takes place for a CTJMDP if the corresponding Markov policy defines a nonexplosive jump Markov process. If this Markov process is explosive, then at each time instance, the marginal probability, that a state-action pair belongs to a measurable set of state-action pairs, is not greater for the described Markov policy than the same probability for the original policy. These results are applied in this paper to CTJMDPs with expected discounted total costs and average costs per unit time. It is shown for these criteria that, if the initial state distribution is fixed, then for every policy, there exists a Markov policy with the same or better value of the objective function.
引用
收藏
页码:1266 / 1286
页数:21
相关论文
共 50 条
  • [1] Sufficiency of Markov Policies for Continuous-Time Markov Decision Processes and Solutions to Kolmogorov's Forward Equation for Jump Markov Processes
    Feinberg, Eugene A.
    Mandava, Manasa
    Shiryaev, Albert N.
    2013 IEEE 52ND ANNUAL CONFERENCE ON DECISION AND CONTROL (CDC), 2013, : 5728 - 5732
  • [2] Optimality of Mixed Policies for Average Continuous-Time Markov Decision Processes with Constraints
    Guo, Xianping
    Zhang, Yi
    MATHEMATICS OF OPERATIONS RESEARCH, 2016, 41 (04) : 1276 - 1296
  • [3] Bias and overtaking optimality for continuous-time jump Markov decision processes in polish spaces
    Zhu, Quanxin
    Prieto-Rumeau, Tomas
    JOURNAL OF APPLIED PROBABILITY, 2008, 45 (02) : 417 - 429
  • [4] Heat Release by Controlled Continuous-Time Markov Jump Processes
    Muratore-Ginanneschi, Paolo
    Mejia-Monasterio, Carlos
    Peliti, Luca
    JOURNAL OF STATISTICAL PHYSICS, 2013, 150 (01) : 181 - 203
  • [5] Heat Release by Controlled Continuous-Time Markov Jump Processes
    Paolo Muratore-Ginanneschi
    Carlos Mejía-Monasterio
    Luca Peliti
    Journal of Statistical Physics, 2013, 150 : 181 - 203
  • [6] Impulsive control for continuous-time Markov decision processes
    Université Bordeaux, IMB, INRIA Bordeaux Sud-Ouest, 200 Avenue de la Vieille Tour, Talence Cedex
    33405, France
    不详
    L69 7ZL, United Kingdom
    Adv Appl Probab, 1 (106-127):
  • [7] The Transformation Method for Continuous-Time Markov Decision Processes
    Piunovskiy, Alexey
    Zhang, Yi
    JOURNAL OF OPTIMIZATION THEORY AND APPLICATIONS, 2012, 154 (02) : 691 - 712
  • [8] IMPULSIVE CONTROL FOR CONTINUOUS-TIME MARKOV DECISION PROCESSES
    Dufour, Francois
    Piunovskiy, Alexei B.
    ADVANCES IN APPLIED PROBABILITY, 2015, 47 (01) : 106 - 127
  • [10] Continuous-Time Markov Decision Processes with Controlled Observations
    Huang, Yunhan
    Kavitha, Veeraruna
    Zhu, Quanyan
    2019 57TH ANNUAL ALLERTON CONFERENCE ON COMMUNICATION, CONTROL, AND COMPUTING (ALLERTON), 2019, : 32 - 39