Adaptive optimal control for a class of continuous-time affine nonlinear systems with unknown internal dynamics

被引：75

作者：

Liu, Derong ^{[1
]}

Yang, Xiong ^{[1
]}

Li, Hongliang ^{[1
]}

机构：

[1] Chinese Acad Sci, Inst Automat, State Key Lab Management & Control Complex Syst, Beijing 100190, Peoples R China

来源：

NEURAL COMPUTING & APPLICATIONS | 2013年 / 23卷 / 7-8期

基金：

中国国家自然科学基金;

关键词：

Adaptive dynamic programming; Reinforcement learning; Policy iteration; Adaptive optimal control; Neural network; Online control; Nonlinear system; APPROXIMATION;

D O I：

10.1007/s00521-012-1249-y

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

This paper develops an online algorithm based on policy iteration for optimal control with infinite horizon cost for continuous-time nonlinear systems. In the present method, a discounted value function is employed, which is considered to be a more general case for optimal control problems. Meanwhile, without knowledge of the internal system dynamics, the algorithm can converge uniformly online to the optimal control, which is the solution of the modified Hamilton-Jacobi-Bellman equation. By means of two neural networks, the algorithm is able to find suitable approximations of both the optimal control and the optimal cost. The uniform convergence to the optimal control is shown, guaranteeing the stability of the nonlinear system. A simulation example is provided to illustrate the effectiveness and applicability of the present approach.

引用

页码：1843 / 1850

页数：8

共 24 条

[1] Nearly optimal control laws for nonlinear systems with saturating actuators using a neural network HJB approach
Abu-Khalaf, M
Lewis, FL
[J]. AUTOMATICA, 2005, 41 (05) : 779 - 791
[2] Discrete-time nonlinear HJB solution using approximate dynamic programming: Convergence proof
Al-Tamimi, Asma
Lewis, Frank L.
Abu-Khalaf, Murad
[J]. IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS PART B-CYBERNETICS, 2008, 38 (04): : 943 - 949
[3] [Anonymous], 1972, The method of weighted residuals and variational principles
[4] Galerkin approximations of the generalized Hamilton-Jacobi-Bellman equation
Beard, RW
Saridis, GN
Wen, JT
[J]. AUTOMATICA, 1997, 33 (12) : 2159 - 2177
[5] Bellman R. E., 1957, Dynamic programming. Princeton landmarks in mathematics
[6] Guo L., 2005, INTRO CONTROL THEORY
[7] UNIVERSAL APPROXIMATION OF AN UNKNOWN MAPPING AND ITS DERIVATIVES USING MULTILAYER FEEDFORWARD NETWORKS
HORNIK, K
STINCHCOMBE, M
WHITE, H
[J]. NEURAL NETWORKS, 1990, 3 (05) : 551 - 560
[8] Howard R. A., 1960, Dynamic programming and Markov processes
[9] SOME NEW ALGORITHMS FOR RECURSIVE ESTIMATION IN CONSTANT LINEAR-SYSTEMS
KAILATH, T
[J]. IEEE TRANSACTIONS ON INFORMATION THEORY, 1973, 19 (06) : 750 - 760
[10] SCHUR METHOD FOR SOLVING ALGEBRAIC RICCATI-EQUATIONS
LAUB, AJ
[J]. IEEE TRANSACTIONS ON AUTOMATIC CONTROL, 1979, 24 (06) : 913 - 921

← 1 2 3 →