A Large-scale Robustness Analysis of Video Action Recognition Models

被引：3

作者：

Schiappa, Madeline Chantry ^{[1
]}

Biyani, Naman ^{[2
]}

Kamtam, Prudvi ^{[1
]}

Vyas, Shruti ^{[1
]}

Palangi, Hamid ^{[3
]}

Vineet, Vibhav ^{[3
]}

Rawat, Yogesh ^{[1
]}

机构：

[1] Univ Cent Florida, CRCV, Orlando, FL 32816 USA

[2] IIT Kanpur, Kanpur, Uttar Pradesh, India

[3] Microsoft Res, Redmond, WA USA

来源：

2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR) | 2023年

关键词：

D O I：

10.1109/CVPR52729.2023.01412

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

We have seen a great progress in video action recognition in recent years. There are several models based on convolutional neural network (CNN) and some recent transformer based approaches which provide top performance on existing benchmarks. In this work, we perform a large-scale robustness analysis of these existing models for video action recognition. We focus on robustness against real-world distribution shift perturbations instead of adversarial perturbations. We propose four different benchmark datasets, HMDB51-P, UCF101-P, Kinetics400-P, and SSv2-P to perform this analysis. We study robustness of six state-of-the-art action recognition models against 90 different perturbations. The study reveals some interesting findings, 1) transformer based models are consistently more robust compared to CNN based models, 2) Pretraining improves robustness for Transformer based models more than CNN based models, and 3) All of the studied models are robust to temporal perturbations for all datasets but SSv2; suggesting the importance of temporal information for action recognition varies based on the dataset and activities. Next, we study the role of augmentations in model robustness and present a real-world dataset, UCF101-DS, which contains realistic distribution shifts, to further validate some of these findings. We believe this study will serve as a benchmark for future research in robust video action recognition

引用

页码：14698 / 14708

页数：11

共 69 条

[1] Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video Captioning
Aafaq, Nayyer
Akhtar, Naveed
Liu, Wei
Gilani, Syed Zulqarnain
Mian, Ajmal
[J]. 2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, : 12479 - 12488
[2] Abu-El-Haija Sami, 2016, Youtube-8m: A large-scale video classification benchmark
[3] Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
Akhtar, Naveed
Mian, Ajmal
[J]. IEEE ACCESS, 2018, 6 : 14410 - 14430
[4] [Anonymous], 2019, ICML
[5] Ardulov Victor, 2021, SCI REPORTS, V11, P1
[6] Artificial Intelligence and Human Trust in Healthcare: Focus on Clinicians
Asan, Onur
Bayrak, Alparslan Emrah
Choudhury, Avishek
[J]. JOURNAL OF MEDICAL INTERNET RESEARCH, 2020, 22 (06)
[7] Bertasius G, 2021, PR MACH LEARN RES, V139
[8] Understanding Robustness of Transformers for Image Classification
Bhojanapalli, Srinadh
Chakrabarti, Ayan
Glasner, Daniel
Li, Daliang
Unterthiner, Thomas
Veit, Andreas
[J]. 2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, : 10211 - 10221
[9] Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Carreira, Joao
Zisserman, Andrew
[J]. 30TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2017), 2017, : 4724 - 4733
[10] Carreira Joao, 2018, QUO VADIS ACTION REC

← 1 2 3 4 5 6 7 →