A Large-scale Robustness Analysis of Video Action Recognition Models

被引:3
作者
Schiappa, Madeline Chantry [1 ]
Biyani, Naman [2 ]
Kamtam, Prudvi [1 ]
Vyas, Shruti [1 ]
Palangi, Hamid [3 ]
Vineet, Vibhav [3 ]
Rawat, Yogesh [1 ]
机构
[1] Univ Cent Florida, CRCV, Orlando, FL 32816 USA
[2] IIT Kanpur, Kanpur, Uttar Pradesh, India
[3] Microsoft Res, Redmond, WA USA
来源
2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR) | 2023年
关键词
D O I
10.1109/CVPR52729.2023.01412
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
We have seen a great progress in video action recognition in recent years. There are several models based on convolutional neural network (CNN) and some recent transformer based approaches which provide top performance on existing benchmarks. In this work, we perform a large-scale robustness analysis of these existing models for video action recognition. We focus on robustness against real-world distribution shift perturbations instead of adversarial perturbations. We propose four different benchmark datasets, HMDB51-P, UCF101-P, Kinetics400-P, and SSv2-P to perform this analysis. We study robustness of six state-of-the-art action recognition models against 90 different perturbations. The study reveals some interesting findings, 1) transformer based models are consistently more robust compared to CNN based models, 2) Pretraining improves robustness for Transformer based models more than CNN based models, and 3) All of the studied models are robust to temporal perturbations for all datasets but SSv2; suggesting the importance of temporal information for action recognition varies based on the dataset and activities. Next, we study the role of augmentations in model robustness and present a real-world dataset, UCF101-DS, which contains realistic distribution shifts, to further validate some of these findings. We believe this study will serve as a benchmark for future research in robust video action recognition
引用
收藏
页码:14698 / 14708
页数:11
相关论文
共 69 条
  • [21] X3D: Expanding Architectures for Efficient Video Recognition
    Feichtenhofer, Christoph
    [J]. 2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2020, : 200 - 210
  • [22] Geirhos R., 2019, INT C LEARN REPR
  • [23] Goyal R., 2017, The "Something Something"Video Database for Learning and Evaluating Visual Common Sense
  • [24] The "something something" video database for learning and evaluating visual common sense
    Goyal, Raghav
    Kahou, Samira Ebrahimi
    Michalski, Vincent
    Materzynska, Joanna
    Westphal, Susanne
    Kim, Heuna
    Haenel, Valentin
    Fruend, Ingo
    Yianilos, Peter
    Mueller-Freitag, Moritz
    Hoppe, Florian
    Thurau, Christian
    Bax, Ingo
    Memisevic, Roland
    [J]. 2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, : 5843 - 5851
  • [25] Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition
    Hara, Kensho
    Kataoka, Hirokatsu
    Satoh, Yutaka
    [J]. 2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION WORKSHOPS (ICCVW 2017), 2017, : 3154 - 3160
  • [26] Hendrycks D., 2018, INT C LEARN REPR
  • [27] Hendrycks D., 2019, P 7 INT C LEARN REPR
  • [28] PIXMIX: Dreamlike Pictures Comprehensively Improve Safety Measures
    Hendrycks, Dan
    Zou, Andy
    Mazeika, Mantas
    Tang, Leonard
    Li, Bo
    Song, Dawn
    Steinhardt, Jacob
    [J]. 2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 16762 - 16771
  • [29] The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
    Hendrycks, Dan
    Basart, Steven
    Mu, Norman
    Kadavath, Saurav
    Wang, Frank
    Dorundo, Evan
    Desai, Rahul
    Zhu, Tyler
    Parajuli, Samyak
    Guo, Mike
    Song, Dawn
    Steinhardt, Jacob
    Gilmer, Justin
    [J]. 2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, : 8320 - 8329
  • [30] What Makes a Video a Video: Analyzing Temporal Information in Video Understanding Models and Datasets
    Huang, De-An
    Ramanathan, Vignesh
    Mahajan, Dhruv
    Torresani, Lorenzo
    Paluri, Manohar
    Li Fei-Fei
    Niebles, Juan Carlos
    [J]. 2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 7366 - 7375