Reinforcement Learning with Temporal Logic Constraints

被引:3
作者
Lennartson, Bengt [1 ]
Jia, Qing-Shan [2 ]
机构
[1] Chalmers Univ Technol, Div Syst & Control, SE-41296 Gothenburg, Sweden
[2] Tsinghua Univ, CFINS, Dept Automat, BNRist, Beijing 100084, Peoples R China
来源
IFAC PAPERSONLINE | 2020年 / 53卷 / 04期
关键词
reinforcement learning; adaption; temporal logic specifications; modular systems;
D O I
10.1016/j.ifacol.2021.04.044
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Reinforcement learning (RL) is an agent based AI learning method, where learning and optimization are combined. Dynamic programming is then performed iteratively, based on reward and next state observations from the system to be controlled. A brief survey of RL is given, followed by an evaluation of a recently proposed method to include temporal logic safety and liveness guarantees in RL, here combined with classical performance optimization. RL is based on Markov decision processes (MDPs), and to reduce the number of observations from the system, a modular MDP framework is proposed. In the learning process, it is then assumed that some parts of the system are represented by known MDP models, while other parts can be estimated by observations from the real system. Local information from the modular system may then be used to reduce the computational complexity, especially in the handling of safety properties. Copyright (C) 2020 The Authors.
引用
收藏
页码:485 / 492
页数:8
相关论文
共 22 条
[1]  
Aksaray D, 2016, IEEE DECIS CONTR P, P6565, DOI 10.1109/CDC.2016.7799279
[2]  
Alshiekh M, 2018, AAAI CONF ARTIF INTE, P2669
[3]   Deterministic generators and games for LTL fragments [J].
Alur, R ;
La Torre, S .
16TH ANNUAL IEEE SYMPOSIUM ON LOGIC IN COMPUTER SCIENCE, PROCEEDINGS, 2001, :291-300
[4]  
Astrom K., 2008, ADAPTIVECONTROL, Vsecond
[5]  
Bacci G, 2013, LECT NOTES COMPUT SC, V8087, P74, DOI 10.1007/978-3-642-40313-2_9
[6]  
Baier C, 2008, PRINCIPLES OF MODEL CHECKING, P1
[7]  
Bertsekas D, 2019, REINFORCEMENTLEARNIN
[8]  
Busoniu L, 2010, AUTOM CONTROL ENG SE, P1, DOI 10.1201/9781439821091-f
[9]  
Cassandras C.G., 2008, Introductiontodiscreteeventsystems
[10]  
Gosavi A., 2015, SIMULATION BASEDOPTI, V2