Enriching behavioral ecology with reinforcement learning methods

被引：35

作者：

Frankenhuis, Willem E. ^{[1
]}

Panchanathan, Karthik ^{[2
]}

Barto, Andrew G. ^{[3
]}

机构：

[1] Radboud Univ Nijmegen, Behav Sci Inst, Montessorilaan 3,POB 9104, NL-6500 HE Nijmegen, Netherlands

[2] Univ Missouri, Dept Anthropol, 200 Swallow Hall, Columbia, MO 65211 USA

[3] Univ Massachusetts, Coll Informat & Comp Sci, Amherst, MA 01003 USA

来源：

BEHAVIOURAL PROCESSES | 2019年 / 161卷

关键词：

Adaptation; Evolution; Development; Learning; Dynamic programming; Reinforcement learning; EVOLUTIONARY PSYCHOLOGY; INFORMATION; ADAPTATION; MODEL; PLASTICITY; GAME; PERSPECTIVE; UNCERTAIN; SELECTION; GENETICS;

D O I：

10.1016/j.beproc.2018.01.008

中图分类号：

B84 [心理学];

学科分类号：

04 ; 0402 ;

摘要：

This article focuses on the division of labor between evolution and development in solving sequential, state-dependent decision problems. Currently, behavioral ecologists tend to use dynamic programming methods to study such problems. These methods are successful at predicting animal behavior in a variety of contexts. However, they depend on a distinct set of assumptions. Here, we argue that behavioral ecology will benefit from drawing more than it currently does on a complementary collection of tools, called reinforcement learning methods. These methods allow for the study of behavior in highly complex environments, which conventional dynamic programming methods do not feasibly address. In addition, reinforcement learning methods are well-suited to studying how biological mechanisms solve developmental and learning problems. For instance, we can use them to study simple rules that perform well in complex environments. Or to investigate under what conditions natural selection favors fixed, non-plastic traits (which do not vary across individuals), cue-driven-switch plasticity (innate instructions for adaptive behavioral development based on experience), or developmental selection (the incremental acquisition of adaptive behavior based on experience). If natural selection favors developmental selection, which includes learning from environmental feedback, we can also make predictions about the design of reward systems. Our paper is written in an accessible manner and for a broad audience, though we believe some novel insights can be drawn from our discussion. We hope our paper will help advance the emerging bridge connecting the fields of behavioral ecology and reinforcement learning.

引用

页码：94 / 100

页数：7

共 100 条

[1] Transgenerational induction of defences in animals and plants [J].

Agrawal, AA ;

Laforsch, C ;

Tollrian, R .

NATURE, 1999, 401 (6748) :60-63

[2]

[Anonymous], 2001, CYCLES CONTINGENCY D

[3]

[Anonymous], 1998, AM J PHYS ANTHR

[4]

[Anonymous], 2016, Developmental Biology

[5]

[Anonymous], 1896, The American Naturalist, DOI DOI 10.1086/276408

[6] EVOLUTION OF A SPECIAL-CLASS OF MODIFIABLE BEHAVIORS IN RELATION TO ENVIRONMENTAL PATTERN [J].

ARNOLD, SJ .

AMERICAN NATURALIST, 1978, 112 (984) :415-427

[7]

Barrett H.C., 2015, SHAPE THOUGHT MENTAL

[8]

Barto A. G., 2013, Intrinsically motivated learning in natural and artificial systems, P17

[9]

Bellman R. E., 2010, Dynamic Programming

[10]

Bertsekas Dimitri P., 1996, Neuro-Dynamic Programming, V5

← 1 2 3 4 5 6 7 8 9 10 →