Solving transition independent decentralized Markov decision processes

被引:102
作者
Becker, R [1 ]
Zilberstein, S [1 ]
Lesser, V [1 ]
Goldman, CV [1 ]
机构
[1] Univ Massachusetts, Dept Comp Sci, Amherst, MA 01003 USA
关键词
D O I
10.1613/jair.1497
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Formal treatment of collaborative multi-agent systems has been lagging behind the rapid progress in sequential decision making by individual agents. Recent work in the area of decentralized Markov Decision Processes (MDPs) has contributed to closing this gap, but the computational complexity of these models remains a serious obstacle. To overcome this complexity barrier, we identify a specific class of decentralized MDPs in which the agents' transitions are independent. The class consists of independent collaborating agents that are tied together through a structured global reward function that depends on all of their histories of states and actions. We present a novel algorithm for solving this class of problems and examine its properties, both as an optimal algorithm and as an anytime algorithm. To the best of our knowledge, this is the first algorithm to optimally solve a non-trivial subclass of decentralized MDPs. It lays the foundation for further work in this area on both exact and approximate algorithms.
引用
收藏
页码:423 / 455
页数:33
相关论文
共 31 条
  • [1] [Anonymous], P 2 INT JOINT C AUT
  • [2] [Anonymous], 2001, P 5 INT C AUTONOMOUS
  • [3] BECKER R, 2004, P 3 INT JOINT C AUT, V1, P302
  • [4] Becker R., 2003, P 2 INT JOINT C AUT, P41
  • [5] The complexity of decentralized control of Markov decision processes
    Bernstein, DS
    Givan, R
    Immerman, N
    Zilberstein, S
    [J]. MATHEMATICS OF OPERATIONS RESEARCH, 2002, 27 (04) : 819 - 840
  • [6] Boutilier C, 1999, IJCAI-99: PROCEEDINGS OF THE SIXTEENTH INTERNATIONAL JOINT CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOLS 1 & 2, P478
  • [7] Decker K. S., 1993, International Journal of Intelligent Systems in Accounting, Finance and Management, V2, P215
  • [8] GHAVAMZADEH M, 2002, P 1 INT JOINT C AUT
  • [9] GOLDMAN CV, 2004, IN PRESS J ARTIFICIA
  • [10] GOLDMAN CV, 2003, P 2 INT JOINT C AUT, P137