Online Solution of Two-Player Zero-Sum Games for Continuous-Time Nonlinear Systems With Completely Unknown Dynamics

被引：56

作者：

Fu, Yue ^{[1
]}

Chai, Tianyou ^{[1
]}

机构：

[1] Northeastern Univ, State Key Lab Synthet Automat Proc Ind, Shenyang 110819, Peoples R China

来源：

IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS | 2016年 / 27卷 / 12期

关键词：

Game algebraic Riccati equation (GARE); Hamilton-Jacobi-Isaacs (HJI); nonlinear systems; policy iteration (PI); two-player zero-sum (ZS) games; H-INFINITY CONTROL; STATE-FEEDBACK CONTROL; POLICY UPDATE ALGORITHM; EQUATION;

D O I：

10.1109/TNNLS.2015.2496299

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Regarding two-player zero-sum games of continuous-time nonlinear systems with completely unknown dynamics, this paper presents an online adaptive algorithm for learning the Nash equilibrium solution, i.e., the optimal policy pair. First, for known systems, the simultaneous policy updating algorithm (SPUA) is reviewed. A new analytical method to prove the convergence is presented. Then, based on the SPUA, without using a priori knowledge of any system dynamics, an online algorithm is proposed to simultaneously learn in real time either the minimal nonnegative solution of the Hamilton-Jacobi-Isaacs (HJI) equation or the generalized algebraic Riccati equation for linear systems as a special case, along with the optimal policy pair. The approximate solution to the HJI equation and the admissible policy pair is reexpressed by the approximation theorem. The unknown constants or weights of each are identified simultaneously by resorting to the recursive least square method. The convergence of the online algorithm to the optimal solutions is provided. A practical online algorithm is also developed. Simulation results illustrate the effectiveness of the proposed method.

引用

页码：2577 / 2587

页数：11

共 29 条

[1] Nearly optimal control laws for nonlinear systems with saturating actuators using a neural network HJB approach [J].

Abu-Khalaf, M ;

Lewis, FL .

AUTOMATICA, 2005, 41 (05) :779-791

[2] Policy iterations on the Hamilton-Jacobi-Isaacs equation for H∞ state feedback control with input saturation [J].

Abu-Khalaf, Murad ;

Lewis, Frank L. ;

Huang, Jie .

IEEE TRANSACTIONS ON AUTOMATIC CONTROL, 2006, 51 (12) :1989-1995

[3] Neurodynamic programming and zero-sum games for constrained control systems [J].

Abu-Khalaf, Murad ;

Lewis, Frank L. ;

Huang, Jie .

IEEE TRANSACTIONS ON NEURAL NETWORKS, 2008, 19 (07) :1243-1252

[4]

[Anonymous], 2002, NONLINEAR SYSTEMS

[5]

Bellman R. E., 1957, Dynamic programming. Princeton landmarks in mathematics

[6]

Bellman R.E., 1962, Applied Dynamic Programming

[7] A game theoretic algorithm to compute local stabilizing solutions to HJBI equations in nonlinear H∞ control [J].

Feng, Yantao ;

Anderson, Brian D. O. ;

Rotkowitz, Michael .

AUTOMATICA, 2009, 45 (04) :881-888

[8]

Goodwin G. C., 1984, Adaptive filtering prediction and control

[9]

Hongliang Li, 2013, Neural Information Processing. 20th International Conference, ICONIP 2013. Proceedings: LNCS 8226, P225, DOI 10.1007/978-3-642-42054-2_29

[10]

JIANG Y., 2013, Journal of Materials Chemistry A, V1, P1, DOI DOI 10.1155/2013/170910

← 1 2 3 →