Exploiting spatio-temporal knowledge for video action recognition

被引:3
作者
Zhang, Huigang [1 ]
Wang, Liuan [1 ]
Sun, Jun [1 ]
机构
[1] Fujitsu R&D Ctr, Beijing 100022, Peoples R China
关键词
action recognition; commonsense knowledge; GCN; STKM;
D O I
10.1049/cvi2.12154
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Action recognition has been a popular area of computer vision research in recent years. The goal of this task is to recognise human actions in video frames. Most existing methods often depend on the visual features and their relationships inside the videos. The extracted features only represent the visual information of the current video itself and cannot represent the general knowledge of particular actions beyond the video. Thus, there are some deviations in these features, and the recognition performance still requires improvement. In this sudy, we present a novel spatio-temporal knowledge module (STKM) to endow the current methods with commonsense knowledge. To this end, we first collect hybrid external knowledge from universal fields, which contains both visual and semantic information. Then graph convolution networks (GCN) are used to represent and aggregate this knowledge. The GCNs involve (i) a spatial graph to capture spatial relations and (ii) a temporal graph to capture serial occurrence relations among actions. By integrating knowledge and visual features, we can get better recognition results. Experiments on AVA, UCF101-24 and JHMDB datasets show the robustness and generalisation ability of STKM. The results report a new state-of-the-art 32.0 mAP on AVA v2.1. On UCF101-24 and JHMDB datasets, our method also improves by 1.5 AP and 2.6 AP, respectively, over the baseline method.
引用
收藏
页码:222 / 230
页数:9
相关论文
共 46 条
[31]  
Speer R, 2017, AAAI CONF ARTIF INTE, P4444
[32]   Asynchronous Interaction Aggregation for Action Detection [J].
Tang, Jiajun ;
Xia, Jin ;
Mu, Xinzhi ;
Pang, Bo ;
Lu, Cewu .
COMPUTER VISION - ECCV 2020, PT XV, 2020, 12360 :71-87
[33]  
Velickovic P, 2018, 6 INT C LEARNING REP
[34]   Videos as Space-Time Region Graphs [J].
Wang, Xiaolong ;
Gupta, Abhinav .
COMPUTER VISION - ECCV 2018, PT V, 2018, 11209 :413-431
[35]   Long-Term Feature Banks for Detailed Video Understanding [J].
Wu, Chao-Yuan ;
Feichtenhofer, Christoph ;
Fan, Haoqi ;
He, Kaiming ;
Krahenbuhl, Philipp ;
Girshick, Ross .
2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, :284-293
[36]   Multi-scale Positive Sample Refinement for Few-Shot Object Detection [J].
Wu, Jiaxi ;
Liu, Songtao ;
Huang, Di ;
Wang, Yunhong .
COMPUTER VISION - ECCV 2020, PT XVI, 2020, 12361 :456-472
[37]   A Low-power Pyramid Motion Estimation Engine for 4K@30fps Realtime HEVC Video Encoding [J].
Xu, Ke ;
Huang, Bo ;
Liu, Xiangkai ;
Tu, Xueying ;
Wu, Zhuoyan ;
Yan, Zhanpeng ;
Liu, Peng ;
Han, Bin ;
Li, Yu .
2018 IEEE INTERNATIONAL SYMPOSIUM ON CIRCUITS AND SYSTEMS (ISCAS), 2018,
[38]  
Yan SJ, 2018, AAAI CONF ARTIF INTE, P7444
[39]   Graph R-CNN for Scene Graph Generation [J].
Yang, Jianwei ;
Lu, Jiasen ;
Lee, Stefan ;
Batra, Dhruv ;
Parikh, Devi .
COMPUTER VISION - ECCV 2018, PT I, 2018, 11205 :690-706
[40]   STEP: Spatio-Temporal Progressive Learning for Video Action Detection [J].
Yang, Xitong ;
Yang, Xiaodong ;
Liu, Ming-Yu ;
Xiao, Fanyi ;
Davis, Larry ;
Kautz, Jan .
2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, :264-272