Convergence of a Q-learning Variant for Continuous States and Actions

被引:5
|
作者
Carden, Stephen [1 ]
机构
[1] Clemson Univ, Dept Math Sci, Clemson, SC 29631 USA
关键词
D O I
10.1613/jair.4271
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
This paper presents a reinforcement learning algorithm for solving infinite horizon Markov Decision Processes under the expected total discounted reward criterion when both the state and action spaces are continuous. This algorithm is based on Watkins' Q-learning, but uses Nadaraya-Watson kernel smoothing to generalize knowledge to unvisited states. As expected, continuity conditions must be imposed on the mean rewards and transition probabilities. Using results from kernel regression theory, this algorithm is proven capable of producing a Q-value function estimate that is uniformly within an arbitrary tolerance of the true Q-value function with probability one. The algorithm is then applied to an example problem to empirically show convergence as well.
引用
收藏
页码:705 / 731
页数:27
相关论文
共 50 条
  • [1] Convergence of a Q-learning variant for continuous states and actions
    Carden, S., 1600, AI Access Foundation (49):
  • [2] A novel Q-learning approach with continuous states and actions
    Zhou, Yi
    Er, Meng Joo
    PROCEEDINGS OF THE 2007 IEEE CONFERENCE ON CONTROL APPLICATIONS, VOLS 1-3, 2007, : 447 - +
  • [3] Q-learning based on regularization theory to treat the continuous states and actions
    Fukao, T
    Sumitomo, T
    Ineyama, N
    Adachi, N
    IEEE WORLD CONGRESS ON COMPUTATIONAL INTELLIGENCE, 1998, : 1057 - 1062
  • [4] Fuzzy interporation-based Q-learning with continuous states and actions
    Horiuchi, T
    Fujino, A
    Katai, O
    Sawaragi, T
    FUZZ-IEEE '96 - PROCEEDINGS OF THE FIFTH IEEE INTERNATIONAL CONFERENCE ON FUZZY SYSTEMS, VOLS 1-3, 1996, : 594 - 600
  • [5] q-Learning in Continuous Time
    Jia, Yanwei
    Zhou, Xun Yu
    JOURNAL OF MACHINE LEARNING RESEARCH, 2023, 24
  • [6] Convergence of optimistic and incremental Q-learning
    Even-Dar, E
    Mansour, Y
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 14, VOLS 1 AND 2, 2002, 14 : 1499 - 1506
  • [7] Continuous-Action Q-Learning
    José del R. Millán
    Daniele Posenato
    Eric Dedieu
    Machine Learning, 2002, 49 : 247 - 265
  • [8] Continuous-action Q-learning
    Millán, JDR
    Posenato, D
    Dedieu, E
    MACHINE LEARNING, 2002, 49 (2-3) : 247 - 265
  • [9] The asymptotic convergence-rate of Q-learning
    Szepesvari, C
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 10, 1998, 10 : 1064 - 1070
  • [10] Q-learning in continuous state and action spaces
    Gaskett, C
    Wettergreen, D
    Zelinsky, A
    ADVANCED TOPICS IN ARTIFICIAL INTELLIGENCE, 1999, 1747 : 417 - 428