FINE-GRAINED POSE TEMPORAL MEMORY MODULE FOR VIDEO POSE ESTIMATION AND TRACKING

被引:0
|
作者
Wang, Chaoyi [1 ]
Hua, Yang [2 ]
Song, Tao [1 ]
Xue, Zhengui [1 ]
Ma, Ruhui [1 ]
Robertson, Neil [2 ]
Guan, Haibing [1 ]
机构
[1] Shanghai Jiao Tong Univ, Shanghai, Peoples R China
[2] Queens Univ Belfast, Belfast, Antrim, North Ireland
关键词
video pose estimation and tracking; keypoint occlusion;
D O I
10.1109/ICASSP39728.2021.9413650
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
The task of video pose estimation and tracking has been largely improved with the development of image pose estimation recently. However, there are still many challenging cases, such as body part occlusion, fast body motion, camera zooming, and complex background. Most existing methods generally use the temporal information to get more precise human bounding boxes or just use it in the tracking stage, but they fail to improve the accuracy of pose estimation tasks. To better solve these problems and utilize the temporal information efficiently and effectively, we present a novel structure, called pose temporal memory module, which is flexible to be transferred into top-down pose estimation frameworks. The temporal information stored in the pose temporal memory is aggregated into the current frame feature in our proposed module. We also transfer compositional de-attention (CoDA) to solve the unique keypoint occlusion problem in this task and propose a novel keypoint feature replacement to recover the extreme error detection under fine-grained keypoint-level guidance. To verify the generality and effectiveness of our proposed method, we integrate our module into two widely used pose estimation frameworks and obtain notable improvement on the PoseTrack dataset with only a few extra computing resources.
引用
收藏
页码:2205 / 2209
页数:5
相关论文
共 50 条
  • [21] Face analysis in video: face detection and tracking with pose estimation
    Mliki, Hazar
    Hammami, Mohamed
    INTERNATIONAL JOURNAL OF BIOMETRICS, 2018, 10 (02) : 121 - 141
  • [22] A Particle Filtering Framework for Joint Video Tracking and Pose Estimation
    Chen, Chong
    Schonfeld, Dan
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2010, 19 (06) : 1625 - 1634
  • [23] Fine-Grained Motion Estimation for Video Frame Interpolation
    Yan, Bo
    Tan, Weimin
    Lin, Chuming
    Shen, Liquan
    IEEE TRANSACTIONS ON BROADCASTING, 2021, 67 (01) : 174 - 184
  • [24] FSA-Net: Learning Fine-Grained Structure Aggregation for Head Pose Estimation from a Single Image
    Yang, Tsun-Yi
    Chen, Yi-Ting
    Lin, Yen-Yu
    Chuang, Yung-Yu
    2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, : 1087 - 1096
  • [25] Spatio-Temporal Matching for Human Pose Estimation in Video
    Zhou, Feng
    De la Torre, Fernando
    IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2016, 38 (08) : 1492 - 1504
  • [26] GAPNET: GENERIC-ATTRIBUTE-POSE NETWORK FOR FINE-GRAINED VISUAL CATEGORIZATION USING MULTI-ATTRIBUTE ATTENTION MODULE
    Ju, Minjeong
    Ryu, Hobin
    Moon, Sangkeun
    Yoo, Chang D.
    2020 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2020, : 703 - 707
  • [27] Aligned to the Object, not to the Image: A Unified Pose-aligned Representation for Fine-grained Recognition
    Guo, Pei
    Farrell, Ryan
    2019 IEEE WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV), 2019, : 1876 - 1885
  • [28] FineAction: A Fine-Grained Video Dataset for Temporal Action Localization
    Liu, Yi
    Wang, Limin
    Wang, Yali
    Ma, Xiao
    Qiao, Yu
    IEEE Transactions on Image Processing, 2022, 31 : 6937 - 6950
  • [29] Player Pose Analysis in Tennis Video based on Pose Estimation
    Kurose, Ryunosuke
    Hayashi, Masaki
    Ishii, Takeo
    Aoki, Yoshimitsu
    2018 INTERNATIONAL WORKSHOP ON ADVANCED IMAGE TECHNOLOGY (IWAIT), 2018,
  • [30] SyDog-Video: A Synthetic Dog Video Dataset for Temporal Pose Estimation
    Shooter, Moira
    Malleson, Charles
    Hilton, Adrian
    INTERNATIONAL JOURNAL OF COMPUTER VISION, 2024, 132 (06) : 1986 - 2002