Attend and Imagine: Multi-Label Image Classification With Visual Attention and Recurrent Neural Networks

被引:51
作者
Lyu, Fan [1 ,2 ]
Wu, Qi [3 ]
Hu, Fuyuan [4 ]
Wu, Qingyao [5 ]
Tan, Mingkui [5 ]
机构
[1] Suzhou Univ Sci & Technol, Suzhou 215009, Peoples R China
[2] Tianjin Univ, Coll Intelligence & Comp, Tianjin 300000, Peoples R China
[3] Univ Adelaide, Sch Comp Sci, Australian Ctr Visual Technol, Adelaide, SA 5005, Australia
[4] Suzhou Univ Sci & Technol, Sch Elect & Informat Engn, Suzhou 215010, Peoples R China
[5] South China Univ Technol, Sch Software Engn, Guangzhou 510640, Guangdong, Peoples R China
关键词
Multi-label classification; visual attention; deep learning; Convolutional Neural Network (CNN); Recurrent Neural Network (RNN); ANNOTATION;
D O I
10.1109/TMM.2019.2894964
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Real images often have multiple labels, i.e., each image is associated with multiple objects or attributes. Compared to single-label image classification, the multilabel classification problem is much more challenging due to several issues. At first, multiple objects can be anywhere in the image. Second, the importance of different regions in an image is different, and the regions of interest in a multilabel image can be very different from another one. Finally, multiple labels of an image can have label dependencies due to complex image structures. To address these challenges, in this paper, we propose to predict the labels sequentially by applying the recurrent neural networks (RNNs), which are used to encode the label dependencies. When predicting a specific label, we introduce a dynamic attention mechanism to enable the model to focus on only regions of interest in the image. Two benchmark datasets (i.e., Pascal VOC and MS-COCO) are adopted to demonstrate the effectiveness of ourwork. Moreover, we construct a new dataset, which includes many semantic dependent labels in each image, to verify the effectiveness of our model. Experimental results show that our method outperforms several state-of-the-arts, especially when predicting some semantic relative labels.
引用
收藏
页码:1971 / 1981
页数:11
相关论文
共 60 条
[1]   Measuring the Objectness of Image Windows [J].
Alexe, Bogdan ;
Deselaers, Thomas ;
Ferrari, Vittorio .
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2012, 34 (11) :2189-2202
[2]  
[Anonymous], PROC CVPR IEEE
[3]  
[Anonymous], IEEE T NEURAL NETW L
[4]  
[Anonymous], 2017, COMMUN ACM, DOI DOI 10.1145/3065386
[5]  
[Anonymous], 2015, P 3 INT C LEARN REPR
[6]  
[Anonymous], 2018, IEEE T NEUR NET LEAR, DOI DOI 10.1109/TNNLS.2018.2827036
[7]  
[Anonymous], 2017, P 31 AAAI C ART INT
[8]  
[Anonymous], 2017, CVPR
[9]  
[Anonymous], J MACH LEARN RES
[10]  
[Anonymous], 2013, ARXIV13124894