Economic MPC of Markov Decision Processes: Dissipativity in undiscounted infinite-horizon optimal control

被引:10
|
作者
Gros, Sebastien [1 ]
Zanon, Mario [2 ]
机构
[1] NTNU, Fac Informat Technol, Dept Eng Cybernet, Trondheim, Norway
[2] IMT Sch Adv Studies Lucca, Piazza San Francesco 19, I-55100 Lucca, Italy
关键词
Markov Decision Processes; Dissipativity for economic MPC; Storage functions; Economic costs; MODEL-PREDICTIVE CONTROL; SYSTEMS;
D O I
10.1016/j.automatica.2022.110602
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Economic Model Predictive Control (MPC) dissipativity theory is central to discussing the stability of policies resulting from minimizing economic stage costs. In its current form, the dissipativity theory for economic MPC applies to problems based on deterministic dynamics or to very specific classes of stochastic problems, and does not readily extend to generic Markov decision processes. In this paper, we clarify the core reason for this difficulty, and propose a generalization of the economic MPC dissipativity theory that circumvents it. This generalization focuses on undiscounted infinite-horizon problems and is based on nonlinear stage cost functionals, allowing one to discuss the Lyapunov asymptotic stability of policies for Markov decision processes in terms of the probability measures underlying their stochastic dynamics. This theory is illustrated for the stochastic linear quadratic regulator with Gaussian process noise, for which a storage functional can be provided explicitly. For the sake of brevity, we limit our discussion to undiscounted Markov decision processes.(c) 2022 Elsevier Ltd. All rights reserved.
引用
收藏
页数:11
相关论文
共 50 条