Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder Dimension

被引：0

作者：

Wang, Ruosong ^{[1
]}

Salakhutdinov, Ruslan ^{[1
]}

Yang, Lin F. ^{[2
]}

机构：

[1] Carnegie Mellon Univ, Pittsburgh, PA 15213 USA

[2] Univ Calif Los Angeles, Los Angeles, CA USA

来源：

ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 33, NEURIPS 2020 | 2020年 / 33卷

关键词：

PAC BOUNDS;

D O I：

暂无

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Value function approximation has demonstrated phenomenal empirical success in reinforcement learning (RL). Nevertheless, despite a handful of recent progress on developing theory for RL with linear function approximation, the understanding of general function approximation schemes largely remains missing. In this paper, we establish the first provably efficient RL algorithm with general value function approximation. We show that if the value functions admit an approximation with a function class F, our algorithm achieves a regret bound of (O) over tilde (poly(dH)root T) where d is a complexity measure of F that depends on the eluder dimension [Russo and Van Roy, 2013] and log-covering numbers, H is the planning horizon, and T is the number interactions with the environment. Our theory generalizes the linear MDP assumption to general function classes. Moreover, our algorithm is model-free and provides a framework to justify the effectiveness of algorithms used in practice.

引用

页数：13

共 69 条

[21] Du Simon S, 2020, Advances in Neural Information Processing Systems
[22] Feldman D, 2013, PROCEEDINGS OF THE TWENTY-FOURTH ANNUAL ACM-SIAM SYMPOSIUM ON DISCRETE ALGORITHMS (SODA 2013), P1434
[23] Feldman D, 2011, ACM S THEORY COMPUT, P569
[24] Filippi S., 2010, ADV NEURAL INFORM PR, P586
[25] Foster DJ, 2018, PR MACH LEARN RES, V80
[26] Jaksch T, 2010, J MACH LEARN RES, V11, P1563
[27] Jia Z., 2019, ARXIV190600423
[28] Jiang N, 2017, PR MACH LEARN RES, V70
[29] Jin C, 2018, Advances in Neural Information Processing Systems, V31, P4863
[30] Jin C., 2020, Provably Efficient Reinforcement Learning with Linear Function Approximation, P2137

← 1 2 3 4 5 6 7 →