Sufficiency of Markov Policies for Continuous-Time Jump Markov Decision Processes

被引：4

作者：

Feinberg, Eugene A. ^{[1
]}

Mandava, Manasa ^{[2
]}

Shiryaev, Albert N. ^{[3
]}

机构：

[1] SUNY Stony Brook, Dept Appl Math & Stat, Stony Brook, NY 11794 USA

[2] Indian Sch Business, Hyderabad 500032, India

[3] Steklov Math Inst, Dept Probabil Theory & Math Stat, Moscow 119991, Russia

来源：

MATHEMATICS OF OPERATIONS RESEARCH | 2022年 / 47卷 / 02期

关键词：

continuous-time jump Markov process; Borel; state; action; Markov policy; COUNTABLE STATE; MODELS;

D O I：

10.1287/moor.2021.1169

中图分类号：

C93 [管理学]; O22 [运筹学];

学科分类号：

070105 ; 12 ; 1201 ; 1202 ; 120202 ;

摘要：

One of the basic facts known for discrete-time Markov decision processes is that, if the probability distribution of an initial state is fixed, then for every policy it is easy to construct a (randomized) Markov policy with the same marginal distributions of state-action pairs as for the original policy. This equality of marginal distributions implies that the values of major objective criteria, including expected discounted total costs and average rewards per unit time, are equal for these two policies. This paper investigates the validity of the similar fact for continuous-time jump Markov decision processes (CTJMDPs). It is shown in this paper that the equality of marginal distributions takes place for a CTJMDP if the corresponding Markov policy defines a nonexplosive jump Markov process. If this Markov process is explosive, then at each time instance, the marginal probability, that a state-action pair belongs to a measurable set of state-action pairs, is not greater for the described Markov policy than the same probability for the original policy. These results are applied in this paper to CTJMDPs with expected discounted total costs and average costs per unit time. It is shown for these criteria that, if the initial state distribution is fixed, then for every policy, there exists a Markov policy with the same or better value of the objective function.

引用

页码：1266 / 1286

页数：21

共 50 条

[1] Sufficiency of Markov Policies for Continuous-Time Markov Decision Processes and Solutions to Kolmogorov's Forward Equation for Jump Markov Processes
Feinberg, Eugene A.
Mandava, Manasa
Shiryaev, Albert N.
2013 IEEE 52ND ANNUAL CONFERENCE ON DECISION AND CONTROL (CDC), 2013, : 5728 - 5732
[2] REALIZABLE STRATEGIES IN CONTINUOUS-TIME MARKOV DECISION PROCESSES
Piunovskiy, Alexey
SIAM JOURNAL ON CONTROL AND OPTIMIZATION, 2018, 56 (01) : 473 - 495
[3] DISCOUNTED CONTINUOUS-TIME CONSTRAINED MARKOV DECISION PROCESSES IN POLISH SPACES
Guo, Xianping
Song, Xinyuan
ANNALS OF APPLIED PROBABILITY, 2011, 21 (05) : 2016 - 2049
[4] The risk probability criterion for discounted continuous-time Markov decision processes
Huo, Haifeng
Zou, Xiaolong
Guo, Xianping
DISCRETE EVENT DYNAMIC SYSTEMS-THEORY AND APPLICATIONS, 2017, 27 (04): : 675 - 699
[5] Tutorial on Structured Continuous-Time Markov Processes
Shelton, Christian R.
Ciardo, Gianfranco
JOURNAL OF ARTIFICIAL INTELLIGENCE RESEARCH, 2014, 51 : 725 - 778
[6] Asynchronous H∞ filtering of continuous-time Markov jump systems
Fang, Mei
Dong, Shanling
Wu, Zheng-Guang
INTERNATIONAL JOURNAL OF ROBUST AND NONLINEAR CONTROL, 2020, 30 (02) : 685 - 698
[7] Finite approximation for finite-horizon continuous-time Markov decision processes
Wei, Qingda
4OR-A QUARTERLY JOURNAL OF OPERATIONS RESEARCH, 2017, 15 (01): : 67 - 84
[8] Finite horizon continuous-time Markov decision processes with mean and variance criteria
Huang, Yonghui
DISCRETE EVENT DYNAMIC SYSTEMS-THEORY AND APPLICATIONS, 2018, 28 (04): : 539 - 564
[9] Square-Root Regret Bounds for Continuous-Time Episodic Markov Decision Processes
Gao, Xuefeng
Zhou, Xunyu
MATHEMATICS OF OPERATIONS RESEARCH, 2025,
[10] Risk Probability Minimization Problems for Continuous-Time Markov Decision Processes on Finite Horizon
Huo, Haifeng
Guo, Xianping
IEEE TRANSACTIONS ON AUTOMATIC CONTROL, 2020, 65 (07) : 3199 - 3206

← 1 2 3 4 5 →