PRIORITIZED SWEEPING - REINFORCEMENT LEARNING WITH LESS DATA AND LESS TIME

被引:257
作者
MOORE, AW
ATKESON, CG
机构
[1] MIT Artificial Intelligence Laboratory, NE43-771, Cambridge, MA, 02139
关键词
MEMORY-BASED LEARNING; LEARNING CONTROL; REINFORCEMENT LEARNING; TEMPORAL DIFFERENCING; ASYNCHRONOUS DYNAMIC PROGRAMMING; HEURISTIC SEARCH; PRIORITIZED SWEEPING;
D O I
10.1023/A:1022635613229
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
We present a new algorithm, prioritized sweeping, for efficient prediction and control of stochastic Markov systems. Incremental learning methods such as temporal differencing and Q-learning have real-time performance. Classical methods are slower, but more accurate, because they make full use of the observations. Prioritized sweeping aims for the best of both worlds. It uses all previous experiences both to prioritize important dynamic programming sweeps and to guide the exploration of state-space. We compare prioritized sweeping with other reinforcement learning schemes for a number of different stochastic optimal control problems. It successfully solves large state-space real-time problems with which other methods have difficulty.
引用
收藏
页码:103 / 130
页数:28
相关论文
共 31 条
  • [21] SAMUEL AL, 1959, IBM J RES DEV, V3, P211, DOI 10.1147/rd.441.0206
  • [22] LEARNING CONTROL OF FINITE MARKOV-CHAINS WITH AN EXPLICIT TRADE-OFF BETWEEN ESTIMATION AND CONTROL
    SATO, M
    ABE, K
    TAKEDA, H
    [J]. IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS, 1988, 18 (05): : 677 - 684
  • [23] SINGH SP, 1991, MACHINE LEARNING, P348
  • [24] PARALLEL FREE-TEXT SEARCH ON THE CONNECTION MACHINE SYSTEM
    STANFILL, C
    KAHLE, B
    [J]. COMMUNICATIONS OF THE ACM, 1986, 29 (12) : 1229 - 1239
  • [25] Sutton R. S., 1988, Machine Learning, V3, P9, DOI 10.1007/BF00115009
  • [26] Sutton R. S., 1990, LEARNING COMPUTATION, P497, DOI DOI 10.1111/J.1748-1716.1960.TB01900.X
  • [27] SUTTON RS, 1990, 7TH P INT C MACH LEA
  • [28] Sutton RS, 1984, THESIS U MASSACHUSET
  • [29] TESAURO GJ, 1991, RC17223 IBM TJ WATS
  • [30] THRUN SB, 1992, ADV NEUR IN, V4, P531