Improving interactive reinforcement learning: What makes a good teacher?

被引：25

作者：

Cruz, Francisco ^{[1
,2
]}

Magg, Sven ^{[1
]}

Nagai, Yukie ^{[3
]}

Wermter, Stefan ^{[1
]}

机构：

[1] Univ Hamburg, Knowledge Technol Grp, Dept Informat, Hamburg, Germany

[2] Univ Cent Chile, Fac Ingn, Escuela Computac & Informat, Santiago, Chile

[3] Osaka Univ, Grad Sch Engn, Emergent Robot Lab, Osaka, Japan

来源：

CONNECTION SCIENCE | 2018年 / 30卷 / 03期

基金：

欧盟地平线“2020”;

关键词：

Interactive reinforcement learning; policy shape; artificial trainer-agent; cleaning scenario;

D O I：

10.1080/09540091.2018.1443318

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Interactive reinforcement learning (IRL) has become an important apprenticeship approach to speed up convergence in classic reinforcement learning (RL) problems. In this regard, a variant of IRL is policy shaping which uses a parent-like trainer to propose the next action to be performed and by doing so reduces the search space by advice. On some occasions, the trainer may be another artificial agent which in turn was trained using RL methods to afterward becoming an advisor for other learner-agents. In this work, we analyse internal representations and characteristics of artificial agents to determine which agent may outperform others to become a better trainer-agent. Using a polymath agent, as compared to a specialist agent, an advisor leads to a larger reward and faster convergence of the reward signal and also to a more stable behaviour in terms of the state visit frequency of the learner-agents. Moreover, we analyse system interaction parameters in order to determine how influential they are in the apprenticeship process, where the consistency of feedback is much more relevant when dealing with different learner obedience parameters.

引用

页码：306 / 325

页数：20

共 35 条

[1] Expertness based cooperative Q-learning [J].

Ahmadabadi, MN ;

Asadpour, M .

IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS PART B-CYBERNETICS, 2002, 32 (01) :66-76

[2]

Ahmadabadi MN, 2000, 2000 IEEE/RSJ INTERNATIONAL CONFERENCE ON INTELLIGENT ROBOTS AND SYSTEMS (IROS 2000), VOLS 1-3, PROCEEDINGS, P2261, DOI 10.1109/IROS.2000.895305

[3]

Amir O., 2016, the Twenty-Fifth International Joint Conference on Artificial Intelligence, IJCAI, P804

[4]

[Anonymous], 2006, AAAI

[5]

[Anonymous], 1994, ON LINE Q LEARNING U

[6]

[Anonymous], 2015, DEV ROBOTICS BABIES

[7]

[Anonymous], 2015, Reinforcement Learning: An Introduction

[8]

Breazeal C., 1998, Proc. 1998 Simulation ofAdaptive Behavior, P25

[9]

Cederborg T, 2015, PROCEEDINGS OF THE TWENTY-FOURTH INTERNATIONAL JOINT CONFERENCE ON ARTIFICIAL INTELLIGENCE (IJCAI), P3366

[10]

Cruz F, 2015, P INT JOINT C NEUR N, P1341

← 1 2 3 4 →