Toward Training Recurrent Neural Networks for Lifelong Learning

被引:41
作者
Sodhani, Shagun [1 ]
Chandar, Sarath [1 ]
Bengio, Yoshua [1 ,2 ]
机构
[1] Univ Montreal, Mila, Montreal, PQ H3T 1J4, Canada
[2] CIFAR, Toronto, ON, Canada
关键词
Learning systems;
D O I
10.1162/neco_a_01246
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Catastrophic forgetting and capacity saturation are the central challenges of any parametric lifelong learning system. In this work, we study these challenges in the context of sequential supervised learning with an emphasis on recurrent neural networks. To evaluate the models in the lifelong learning setting, we propose a curriculum-based, simple, and intuitive benchmark where the models are trained on tasks with increasing levels of difficulty. To measure the impact of catastrophic forgetting, the model is tested on all the previous tasks as it completes any task. As a step toward developing true lifelong learning systems, we unify gradient episodic memory (a catastrophic forgetting alleviation approach) and Net2Net (a capacity expansion approach). Both models are proposed in the context of feedforward networks, and we evaluate the feasibility of using them for recurrent networks. Evaluation on the proposed benchmark shows that the unified model is more suitable than the constituent models for lifelong learning setting.
引用
收藏
页码:1 / 35
页数:35
相关论文
共 35 条
[21]  
Lopez-Paz D, 2017, ADV NEUR IN, V30
[22]   Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights [J].
Mallya, Arun ;
Davis, Dillon ;
Lazebnik, Svetlana .
COMPUTER VISION - ECCV 2018, PT IV, 2018, 11208 :72-88
[23]  
McCloskey M., 1989, Psychology of learning and motivation, V24, P109, DOI [10.1016/S0079-7421(08)60536-8, 10.1016/S0079-7421]
[24]   Metric Learning for Large Scale Image Classification: Generalizing to New Classes at Near-Zero Cost [J].
Mensink, Thomas ;
Verbeek, Jakob ;
Perronnin, Florent ;
Csurka, Gabriela .
COMPUTER VISION - ECCV 2012, PT II, 2012, 7573 :488-501
[25]  
Rebuffi Sylvestre-Alvise, 2017, P C COMP VIS PATT RE, P3
[26]   CHILD: A first step towards continual learning [J].
Ring, MB .
MACHINE LEARNING, 1997, 28 (01) :77-104
[27]  
Romero A., 2014, P INT C LEARNING REP
[28]  
Rusu A. A., 2016, arXiv:1606.04671
[29]  
Serr`a J., 2018, arXiv:1801.01423
[30]  
Silver D. L., 2002, P 15 C CANADIAN SOC, P90