Bottom-Up Temporal Action Localization with Mutual Regularization

被引：145

作者：

Zhao, Peisen ^{[1
]}

Xie, Lingxi ^{[2
]}

Ju, Chen ^{[1
]}

Zhang, Ya ^{[1
]}

Wang, Yanfeng ^{[1
]}

Tian, Qi ^{[2
]}

机构：

[1] Shanghai Jiao Tong Univ, Cooperat Medianet Innovat Ctr, Shanghai, Peoples R China

[2] Huawei Inc, Shenzhen, Peoples R China

来源：

COMPUTER VISION - ECCV 2020, PT VIII | 2020年 / 12353卷

关键词：

Action localization; Action proposals; Mutual regularization;

D O I：

10.1007/978-3-030-58598-3_32

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Recently, temporal action localization (TAL), i.e., finding specific action segments in untrimmed videos, has attracted increasing attentions of the computer vision community. State-of-the-art solutions for TAL involves evaluating the frame-level probabilities of three action-indicating phases, i.e. starting, continuing, and ending; and then post-processing these predictions for the final localization. This paper delves deep into this mechanism, and argues that existing methods, by modeling these phases as individual classification tasks, ignored the potential temporal constraints between them. This can lead to incorrect and/or inconsistent predictions when some frames of the video input lack sufficient discriminative information. To alleviate this problem, we introduce two regularization terms to mutually regularize the learning procedure: the Intra-phase Consistency (IntraC) regularization is proposed to make the predictions verified inside each phase; and the Inter-phase Consistency (InterC) regularization is proposed to keep consistency between these phases. Jointly optimizing these two terms, the entire framework is aware of these potential constraints during an end-to-end optimization process. Experiments are performed on two popular TAL datasets, THUMOS14 and ActivityNet1.3. Our approach clearly outperforms the baseline both quantitatively and qualitatively. The proposed regularization also generalizes to other TAL methods ( e.g., TSA-Net and PGCN). Code: https://github.com/PeisenZhao/Bottom-Up-TAL-with-MR.

引用

页码：539 / 555

页数：17

共 43 条

[1]

Abu-El-Haija S., 2016, arXiv

[2] Soft-NMS - Improving Object Detection With One Line of Code [J].

Bodla, Navaneeth ;

Singh, Bharat ;

Chellappa, Rama ;

Davis, Larry S. .

2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, :5562-5570

[3] SST: Single-Stream Temporal Action Proposals [J].

Buch, Shyamal ;

Escorcia, Victor ;

Shen, Chuanqi ;

Ghanem, Bernard ;

Niebles, Juan Carlos .

30TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2017), 2017, :6373-6382

[4]

Heilbron FC, 2015, PROC CVPR IEEE, P961, DOI 10.1109/CVPR.2015.7298698

[5] Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset [J].

Carreira, Joao ;

Zisserman, Andrew .

30TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2017), 2017, :4724-4733

[6] Rethinking the Faster R-CNN Architecture for Temporal Action Localization [J].

Chao, Yu-Wei ;

Vijayanarasimhan, Sudheendra ;

Seybold, Bryan ;

Ross, David A. ;

Deng, Jia ;

Sukthankar, Rahul .

2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :1130-1139

[7] Temporal Context Network for Activity Localization in Videos [J].

Dai, Xiyang ;

Singh, Bharat ;

Zhang, Guyue ;

Davis, Larry S. ;

Chen, Yan Qiu .

2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, :5727-5736

[8] Learning Spatiotemporal Features with 3D Convolutional Networks [J].

Du Tran ;

Bourdev, Lubomir ;

Fergus, Rob ;

Torresani, Lorenzo ;

Paluri, Manohar .

2015 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2015, :4489-4497

[9] DAPs: Deep Action Proposals for Action Understanding [J].

Escorcia, Victor ;

Heilbron, Fabian Caba ;

Niebles, Juan Carlos ;

Ghanem, Bernard .

COMPUTER VISION - ECCV 2016, PT III, 2016, 9907 :768-784

[10] You Lead, We Exceed: Labor-Free Video Concept Learning by Jointly Exploiting Web Videos and Images [J].

Gan, Chuang ;

Yao, Ting ;

Yang, Kuiyuan ;

Yang, Yi ;

Mei, Tao .

2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, :923-932

← 1 2 3 4 5 →