Going from Image to Video Saliency: Augmenting Image Salience with Dynamic Attentional Push

被引:42
作者
Gorji, Siavash [1 ]
Clark, James J. [1 ]
机构
[1] McGill Univ, Dept Elect & Comp Engn, Ctr Intelligent Machines, Montreal, PQ, Canada
来源
2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR) | 2018年
关键词
VISUAL-ATTENTION; DETECTION MODEL; SCENE; GAZE; EYES;
D O I
10.1109/CVPR.2018.00783
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
We present a novel method to incorporate the recent advent in static saliency models to predict the saliency in videos. Our model augments the static saliency models with the Attentional Push effect of the photographer and the scene actors in a shared attention setting. We demonstrate that not only it is imperative to use static Attentional Push cues, noticeable performance improvement is achievable by learning the time-varying nature of Attentional Push. We propose a multi-stream Convolutional Long Short-Term Memory network (ConvLSTM) structure which augments state-of-the-art in static saliency models with dynamic Attentional Push. Our network contains four pathways, a saliency pathway and three Attentional Push pathways. The multi-pathway structure is followed by an augmenting convnet that learns to combine the complementary and time-varying outputs of the ConvLSTMs by minimizing the relative entropy between the augmented saliency and viewers fixation patterns on videos. We evaluate our model by comparing the performance of several augmented static saliency models with state-of-the-art in spatiotemporal saliency on three largest dynamic eye tracking datasets, HOLLYWOOD2, UCF-Sport and DIEM. Experimental results illustrates that solid performance gain is achievable using the proposed methodology.
引用
收藏
页码:7501 / 7511
页数:11
相关论文
共 67 条
[1]  
[Anonymous], 2015, PROC 28 ADV NEURAL I, DOI DOI 10.1038/SCIENTIFICAMERICAN0700-38
[2]  
[Anonymous], 2009, BMVC
[3]  
[Anonymous], Mit saliency benchmark
[4]  
[Anonymous], 2017, INT C LEARNING REPRE
[5]  
[Anonymous], 2015, PROC CVPR IEEE, DOI DOI 10.1109/CVPR.2015.7298710
[6]  
[Anonymous], 23 INT C PATT REC IC
[7]  
Bak C., 2016, CoRR
[8]  
Bethge M., 2014, International Conference on Learning Representations (ICLR 2015)
[9]   Complementary effects of gaze direction and early saliency in guiding fixations during free viewing [J].
Borji, Ali ;
Parks, Daniel ;
Itti, Laurent .
JOURNAL OF VISION, 2014, 14 (13)
[10]   What stands out in a scene? A study of human explicit saliency judgment [J].
Borji, Ali ;
Sihite, Dicky N. ;
Itti, Laurent .
VISION RESEARCH, 2013, 91 :62-77