Robust Imitation via Mirror Descent Inverse Reinforcement Learning

被引：0

作者：

Han, Dong-Sig ^{[1
]}

Kim, Hyunseo ^{[1
]}

Lee, Hyundo ^{[1
]}

Ryu, Je-Hwan ^{[1
]}

Zhang, Byoung-Tak ^{[1
]}

机构：

[1] Seoul Natl Univ, Artificial Intelligence Inst, Seoul, South Korea

来源：

ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 35 (NEURIPS 2022) | 2022年

关键词：

D O I：

暂无

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Recently, adversarial imitation learning has shown a scalable reward acquisition method for inverse reinforcement learning (IRL) problems. However, estimated reward signals often become uncertain and fail to train a reliable statistical model since the existing methods tend to solve hard optimization problems directly. Inspired by a first-order optimization method called mirror descent, this paper proposes to predict a sequence of reward functions, which are iterative solutions for a constrained convex problem. IRL solutions derived by mirror descent are tolerant to the uncertainty incurred by target density estimation since the amount of reward learning is regulated with respect to local geometric constraints. We prove that the proposed mirror descent update rule ensures robust minimization of a Bregman divergence in terms of a rigorous regret bound of O(1/T) for step sizes {eta(t)}(t=1)(T). Our IRL method was applied on top of an adversarial framework, and it outperformed existing adversarial methods in an extensive suite of benchmarks.

引用

页数：13

共 56 条

[1] Abadi M, 2016, PROCEEDINGS OF OSDI'16: 12TH USENIX SYMPOSIUM ON OPERATING SYSTEMS DESIGN AND IMPLEMENTATION, P265
[2] Abbeel P., 2004, P 21 INT C MACH LEAR, P1, DOI [10.1145/1015330.1015430, DOI 10.1145/1015330.1015430]
[3] Acharyya S., 2013, Proc. of SDM. SIAM, P476
[4] Natural gradient works efficiently in learning
Amari, S
[J]. NEURAL COMPUTATION, 1998, 10 (02) : 251 - 276
[5] Amari SI, 2016, APPL MATH SCI, V194, P1, DOI 10.1007/978-4-431-55978-8
[6] [Anonymous], OPERATIONS RES LETT
[7] Banerjee A, 2005, J MACH LEARN RES, V6, P1705
[8] Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A
Teboulle, M
[J]. OPERATIONS RESEARCH LETTERS, 2003, 31 (03) : 167 - 175
[9] Boyd S., 2004, CONVEX OPTIMIZATION
[10] Bregman LM., 1967, USSR COMP MATH MATH, V7, P200, DOI DOI 10.1016/0041-5553(67)90040-7

← 1 2 3 4 5 6 →