Weakly Supervised Gaussian Networks for Action Detection

被引:0
作者
Fernando, Basura [1 ]
Chet, Cheston Tan Yin [2 ]
Bilen, Hakan [3 ]
机构
[1] ASTAR, A AI, Singapore, Singapore
[2] ASTAR, I2R, Singapore, Singapore
[3] Univ Edinburgh, VICO, Edinburgh, Midlothian, Scotland
来源
2020 IEEE WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV) | 2020年
基金
新加坡国家研究基金会;
关键词
RECOGNITION;
D O I
10.1109/wacv45572.2020.9093263
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Detecting temporal extents of human actions in videos is a challenging computer vision problem that requires detailed manual supervision including frame-level labels. This expensive annotation process limits deploying action detectors to a limited number of categories. We propose a novel method, called WSGN, that learns to detect actions from weak supervision, using only video-level labels. WSGN learns to exploit both video-specific and dataset-wide statistics to predict relevance of each frame to an action category. This strategy leads to significant gains in action detection for two standard benchmarks THUMOS14 and Charades. Our method obtains excellent results compared to state-of-the-art methods that uses similar features and loss functions on THUMOS14 dataset. Similarly, our weakly supervised method is only 0.3% mAP behind a state-of-the-art supervised method on challenging Charades dataset for action localization.
引用
收藏
页码:526 / 535
页数:10
相关论文
共 43 条
[1]  
[Anonymous], 2017, ARXIV170500873
[2]  
[Anonymous], 2015, 3 INT C LEARN REPR I
[3]  
[Anonymous], 2014, ARXIV PREPRINT ARXIV
[4]  
[Anonymous], 2014, ECCV WORKSH
[5]  
[Anonymous], 2019, CVPR
[6]   Action Recognition with Dynamic Image Networks [J].
Bilen, Hakan ;
Fernando, Basura ;
Gavves, Efstratios ;
Vedaldi, Andrea .
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2018, 40 (12) :2799-2813
[7]   Weakly Supervised Deep Detection Networks [J].
Bilen, Hakan ;
Vedaldi, Andrea .
2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, :2846-2854
[8]   The recognition of human movement using temporal templates [J].
Bobick, AF ;
Davis, JW .
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2001, 23 (03) :257-267
[9]   Weakly-Supervised Alignment of Video With Text [J].
Bojanowski, P. ;
Lajugie, R. ;
Grave, E. ;
Bach, F. ;
Laptev, I. ;
Ponce, J. ;
Schmid, C. .
2015 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2015, :4462-4470
[10]  
Bojanowski P, 2014, LECT NOTES COMPUT SC, V8693, P628, DOI 10.1007/978-3-319-10602-1_41