Learning Scene-Aware Spatio-Temporal GNNs for Few-Shot Early Action Prediction

被引:8
|
作者
Hu, Yufan [1 ,2 ]
Gao, Junyu [3 ,4 ]
Xu, Changsheng [2 ,3 ,4 ]
机构
[1] Hefei Univ Technol, Hefei 230009, Peoples R China
[2] Peng Cheng Lab, Shenzhen 518055, Peoples R China
[3] Chinese Acad Sci, Inst Automat, Natl Lab Pattern Recognit, Beijing 100190, Peoples R China
[4] Univ Chinese Acad Sci, Sch Artificial Intelligence, Natl Lab Pattern Recognit, Beijing 100190, Peoples R China
基金
北京市自然科学基金; 中国国家自然科学基金;
关键词
Few-shot learning; early action prediction; scene graph; graph neural network; OBJECT AFFORDANCES;
D O I
10.1109/TMM.2022.3142413
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
We aim to address a new task named few-shot early action prediction (FS-EAP) that learns classifiers for novel actions from only a few partially observed videos. We argue that the task is extremely challenging since the partially observed videos do not contain enough action information in a few-shot environment. To tackle this task, in this paper, we propose a scene-aware spatio-temporal graph neural network (SA-STGNN) by leveraging the fine-grained spatio-temporal interactions in the video scenes. Specifically, we first generate a spatio-temporal graph corresponding to the partially observed video to capture comprehensive spatio-temporal correlations. Then we utilize the spatio-temporal graph as the input of our SA-STGNN and predict the augmented video features corresponding to the complete video. The architecture uses several scene-aware learning blocks, which are a combination of edge fusion graph neural layers and temporal gated convolutional layers to jointly model spatial and temporal dependencies. Finally, we employ an early action predictor to exploit the learned video features for predicting actions in the few-shot setting. Extensive experimental results on two widely adopted video datasets demonstrate the effectiveness of our approach and its superior performance over the state-of-the-art approaches.
引用
收藏
页码:2061 / 2073
页数:13
相关论文
共 50 条
  • [1] Spatio-temporal Relation Modeling for Few-shot Action Recognition
    Thatipelli, Anirudh
    Narayan, Sanath
    Khan, Salman
    Anwer, Rao Muhammad
    Khan, Fahad Shahbaz
    Ghanem, Bernard
    2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 19926 - 19935
  • [2] Searching for Better Spatio-temporal Alignment in Few-Shot Action Recognition
    Cao, Yichao
    Su, Xiu
    Tang, Qingfei
    You, Shan
    Lu, Xiaobo
    Xu, Chang
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 35 (NEURIPS 2022), 2022,
  • [3] Spatio-Temporal Self-supervision for Few-Shot Action Recognition
    Yu, Wanchuan
    Guo, Hanyu
    Yan, Yan
    Li, Jie
    Wang, Hanzi
    PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2023, PT I, 2024, 14425 : 84 - 96
  • [4] Semantic-guided spatio-temporal attention for few-shot action recognition
    Jianyu Wang
    Baolin Liu
    Applied Intelligence, 2024, 54 : 2458 - 2471
  • [5] Semantic-guided spatio-temporal attention for few-shot action recognition
    Wang, Jianyu
    Liu, Baolin
    APPLIED INTELLIGENCE, 2024, 54 (03) : 2458 - 2471
  • [6] Spatio-Temporal Graph Few-Shot Learning with Cross-City Knowledge Transfer
    Lu, Bin
    Gan, Xiaoying
    Zhang, Weinan
    Yao, Huaxiu
    Fu, Luoyi
    Wang, Xinbing
    PROCEEDINGS OF THE 28TH ACM SIGKDD CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING, KDD 2022, 2022, : 1162 - 1172
  • [7] Cross-modal guides spatio-temporal enrichment network for few-shot action recognition
    Chen, Zhiwen
    Yang, Yi
    Li, Li
    Li, Min
    APPLIED INTELLIGENCE, 2024, 54 (22) : 11196 - 11211
  • [8] Few-shot human motion prediction using deformable spatio-temporal CNN with parameter generation
    Zang, Chuanqi
    Li, Menghao
    Pei, Mingtao
    NEUROCOMPUTING, 2022, 513 : 46 - 58
  • [9] TAEN: Temporal Aware Embedding Network for Few-Shot Action Recognition
    Ben-Ari, Rami
    Nacson, Mor Shpigel
    Azulai, Ophir
    Barzelay, Udi
    Rotman, Daniel
    2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION WORKSHOPS, CVPRW 2021, 2021, : 2780 - 2788
  • [10] Meta Learning with Attention Based FP-GNNs for Few-Shot Molecular Property Prediction
    Qian, Xiaoliang
    Ju, Bin
    Shen, Ping
    Yang, Keda
    Li, Li
    Liu, Qi
    ACS OMEGA, 2024, 9 (22): : 23940 - 23948