Integrating Human Parsing and Pose Network for Human Action Recognition

被引：1

作者：

Ding, Runwei ^{[1
]}

Wen, Yuhang ^{[2
]}

Liu, Jinfu ^{[2
]}

Dai, Nan ^{[3
]}

Meng, Fanyang ^{[4
]}

Liu, Mengyuan ^{[1
]}

机构：

[1] Peking Univ, Shenzhen Grad Sch, Shenzhen, Peoples R China

[2] Sun Yat Sen Univ, Shenzhen, Peoples R China

[3] Changchun Univ Sci & Technol, Changchun, Peoples R China

[4] Peng Cheng Lab, Shenzhen, Peoples R China

来源：

ARTIFICIAL INTELLIGENCE, CICAI 2023, PT I | 2024年 / 14473卷

基金：

中国国家自然科学基金;

关键词：

Action recognition; Human parsing; Human skeletons;

D O I：

10.1007/978-981-99-8850-1_15

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Human skeletons and RGB sequences are both widelyadopted input modalities for human action recognition. However, skeletons lack appearance features and color data suffer large amount of irrelevant depiction. To address this, we introduce human parsing feature map as a novel modality, since it can selectively retain spatiotemporal features of the body parts, while filtering out noises regarding outfits, backgrounds, etc. We propose an Integrating Human Parsing and Pose Network (IPP-Net) for action recognition, which is the first to leverage both skeletons and human parsing feature maps in dual-branch approach. The human pose branch feeds compact skeletal representations of different modalities in graph convolutional network to model pose features. In human parsing branch, multi-frame body-part parsing features are extracted with human detector and parser, which is later learnt using a convolutional backbone. A late ensemble of two branches is adopted to get final predictions, considering both robust keypoints and rich semantic body-part features. Extensive experiments on NTU RGB+D and NTU RGB+D 120 benchmarks consistently verify the effectiveness of the proposed IPP-Net, which outperforms the existing action recognition methods. Our code is publicly available at https://github.com/liujf69/IPPNet-Parsing.

引用

页码：182 / 194

页数：13

共 33 条

[1] Beyond Appearance: a Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual Tasks [J].

Chen, Weihua ;

Xu, Xianzhe ;

Jia, Jian ;

Luo, Hao ;

Wang, Yaohua ;

Wang, Fan ;

Jin, Rong ;

Sun, Xiuyu .

2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, :15050-15061

[2] Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition [J].

Chen, Yuxin ;

Zhang, Ziqi ;

Yuan, Chunfeng ;

Li, Bing ;

Deng, Ying ;

Hu, Weiming .

2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, :13339-13348

[3] Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action Recognition [J].

Chen, Tailin ;

Zhou, Desen ;

Wang, Jian ;

Wang, Shidong ;

Guan, Yu ;

He, Xuming ;

Ding, Errui .

PROCEEDINGS OF THE 29TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2021, 2021, :4334-4342

[4] Skeleton-Based Action Recognition with Shift Graph Convolutional Network [J].

Cheng, Ke ;

Zhang, Yifan ;

He, Xiangyu ;

Chen, Weihan ;

Cheng, Jian ;

Lu, Hanqing .

2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2020, :180-189

[5] InfoGCN: Representation Learning for Human Skeleton-based Action Recognition [J].

Chi, Hyung-gun ;

Ha, Myoung Hoon ;

Chi, Seunggeun ;

Lee, Sang Wan ;

Huang, Qixing ;

Ramani, Karthik .

2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, :20154-20164

[6] VPN: Learning Video-Pose Embedding for Activities of Daily Living [J].

Das, Srijan ;

Sharma, Saurav ;

Dai, Rui ;

Bremond, Francois ;

Thonnat, Monique .

COMPUTER VISION - ECCV 2020, PT IX, 2020, 12354 :72-90

[7] Deep Residual Learning for Image Recognition [J].

He, Kaiming ;

Zhang, Xiangyu ;

Ren, Shaoqing ;

Sun, Jian .

2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, :770-778

[8] Self-Correction for Human Parsing [J].

Li, Peike ;

Xu, Yunqiu ;

Wei, Yunchao ;

Yang, Yi .

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2022, 44 (06) :3260-3271

[9] Look into Person: Joint Body Parsing & Pose Estimation Network and a New Benchmark [J].

Liang, Xiaodan ;

Gong, Ke ;

Shen, Xiaohui ;

Lin, Liang .

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2019, 41 (04) :871-885

[10] Deep Human Parsing with Active Template Regression [J].

Liang, Xiaodan ;

Liu, Si ;

Shen, Xiaohui ;

Yang, Jianchao ;

Liu, Luoqi ;

Dong, Jian ;

Lin, Liang ;

Yan, Shuicheng .

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2015, 37 (12) :2402-2414

← 1 2 3 4 →