FINE-GRAINED POSE TEMPORAL MEMORY MODULE FOR VIDEO POSE ESTIMATION AND TRACKING

被引：0

作者：

Wang, Chaoyi ^{[1
]}

Hua, Yang ^{[2
]}

Song, Tao ^{[1
]}

Xue, Zhengui ^{[1
]}

Ma, Ruhui ^{[1
]}

Robertson, Neil ^{[2
]}

Guan, Haibing ^{[1
]}

机构：

[1] Shanghai Jiao Tong Univ, Shanghai, Peoples R China

[2] Queens Univ Belfast, Belfast, Antrim, North Ireland

来源：

2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021) | 2021年

关键词：

video pose estimation and tracking; keypoint occlusion;

D O I：

10.1109/ICASSP39728.2021.9413650

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

The task of video pose estimation and tracking has been largely improved with the development of image pose estimation recently. However, there are still many challenging cases, such as body part occlusion, fast body motion, camera zooming, and complex background. Most existing methods generally use the temporal information to get more precise human bounding boxes or just use it in the tracking stage, but they fail to improve the accuracy of pose estimation tasks. To better solve these problems and utilize the temporal information efficiently and effectively, we present a novel structure, called pose temporal memory module, which is flexible to be transferred into top-down pose estimation frameworks. The temporal information stored in the pose temporal memory is aggregated into the current frame feature in our proposed module. We also transfer compositional de-attention (CoDA) to solve the unique keypoint occlusion problem in this task and propose a novel keypoint feature replacement to recover the extreme error detection under fine-grained keypoint-level guidance. To verify the generality and effectiveness of our proposed method, we integrate our module into two widely used pose estimation frameworks and obtain notable improvement on the PoseTrack dataset with only a few extra computing resources.

引用

页码：2205 / 2209

页数：5

共 50 条

[21] Face analysis in video: face detection and tracking with pose estimation
Mliki, Hazar
Hammami, Mohamed
INTERNATIONAL JOURNAL OF BIOMETRICS, 2018, 10 (02) : 121 - 141
[22] A Particle Filtering Framework for Joint Video Tracking and Pose Estimation
Chen, Chong
Schonfeld, Dan
IEEE TRANSACTIONS ON IMAGE PROCESSING, 2010, 19 (06) : 1625 - 1634
[23] Fine-Grained Motion Estimation for Video Frame Interpolation
Yan, Bo
Tan, Weimin
Lin, Chuming
Shen, Liquan
IEEE TRANSACTIONS ON BROADCASTING, 2021, 67 (01) : 174 - 184
[24] FSA-Net: Learning Fine-Grained Structure Aggregation for Head Pose Estimation from a Single Image
Yang, Tsun-Yi
Chen, Yi-Ting
Lin, Yen-Yu
Chuang, Yung-Yu
2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, : 1087 - 1096
[25] Spatio-Temporal Matching for Human Pose Estimation in Video
Zhou, Feng
De la Torre, Fernando
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2016, 38 (08) : 1492 - 1504
[26] GAPNET: GENERIC-ATTRIBUTE-POSE NETWORK FOR FINE-GRAINED VISUAL CATEGORIZATION USING MULTI-ATTRIBUTE ATTENTION MODULE
Ju, Minjeong
Ryu, Hobin
Moon, Sangkeun
Yoo, Chang D.
2020 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2020, : 703 - 707
[27] Aligned to the Object, not to the Image: A Unified Pose-aligned Representation for Fine-grained Recognition
Guo, Pei
Farrell, Ryan
2019 IEEE WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV), 2019, : 1876 - 1885
[28] FineAction: A Fine-Grained Video Dataset for Temporal Action Localization
Liu, Yi
Wang, Limin
Wang, Yali
Ma, Xiao
Qiao, Yu
IEEE Transactions on Image Processing, 2022, 31 : 6937 - 6950
[29] Player Pose Analysis in Tennis Video based on Pose Estimation
Kurose, Ryunosuke
Hayashi, Masaki
Ishii, Takeo
Aoki, Yoshimitsu
2018 INTERNATIONAL WORKSHOP ON ADVANCED IMAGE TECHNOLOGY (IWAIT), 2018,
[30] SyDog-Video: A Synthetic Dog Video Dataset for Temporal Pose Estimation
Shooter, Moira
Malleson, Charles
Hilton, Adrian
INTERNATIONAL JOURNAL OF COMPUTER VISION, 2024, 132 (06) : 1986 - 2002

← 1 2 3 4 5 →