Computer Science > Computer Vision and Pattern Recognition

arXiv:2307.07977 (cs)

[Submitted on 16 Jul 2023]

Title:Integrating Human Parsing and Pose Network for Human Action Recognition

Authors:Runwei Ding, Yuhang Wen, Jinfu Liu, Nan Dai, Fanyang Meng, Mengyuan Liu

View PDF

Abstract:Human skeletons and RGB sequences are both widely-adopted input modalities for human action recognition. However, skeletons lack appearance features and color data suffer large amount of irrelevant depiction. To address this, we introduce human parsing feature map as a novel modality, since it can selectively retain spatiotemporal features of the body parts, while filtering out noises regarding outfits, backgrounds, etc. We propose an Integrating Human Parsing and Pose Network (IPP-Net) for action recognition, which is the first to leverage both skeletons and human parsing feature maps in dual-branch approach. The human pose branch feeds compact skeletal representations of different modalities in graph convolutional network to model pose features. In human parsing branch, multi-frame body-part parsing features are extracted with human detector and parser, which is later learnt using a convolutional backbone. A late ensemble of two branches is adopted to get final predictions, considering both robust keypoints and rich semantic body-part features. Extensive experiments on NTU RGB+D and NTU RGB+D 120 benchmarks consistently verify the effectiveness of the proposed IPP-Net, which outperforms the existing action recognition methods. Our code is publicly available at this https URL .

Comments:	CICAI 2023 Camera-ready Version
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2307.07977 [cs.CV]
	(or arXiv:2307.07977v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2307.07977

Submission history

From: Yuhang Wen [view email]
[v1] Sun, 16 Jul 2023 07:58:29 UTC (1,092 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Integrating Human Parsing and Pose Network for Human Action Recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Integrating Human Parsing and Pose Network for Human Action Recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators