ML-HDP: A Hierarchical Bayesian Nonparametric Model for Recognizing Human Actions in Video

被引：38

作者：

Nguyen Anh Tu ^{[1
]}

Thien Huynh-The ^{[1
]}

Khan, Kifayat Ullah ^{[1
]}

Lee, Young-Koo ^{[1
]}

机构：

[1] Kyung Hee Univ, Dept Comp Sci & Engn, Global Campus, Seoul 17104, South Korea

来源：

IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY | 2019年 / 29卷 / 03期

关键词：

Action recognition; action segmentation; action localization; topic modeling; Dirichlet process; ACTION RECOGNITION; VECTOR; CLASSIFICATION; DESCRIPTORS;

D O I：

10.1109/TCSVT.2018.2816960

中图分类号：

TM [电工技术]; TN [电子技术、通信技术];

学科分类号：

0808 ; 0809 ;

摘要：

Action recognition from videos is an important area of computer vision research due to its various applications, ranging from visual surveillance to human-computer interaction. To address action recognition problems, this paper presents a framework that jointly models multiple complex actions and motion units at different hierarchical levels. We achieve this by proposing a generative topic model, namely, multi-label hierarchical Dirichlet process (ML-HDP). The ML-HDP model formulates the co-occurrence relationship of actions and motion units, and enables highly accurate recognition. In particular, our topic model possesses the three-level representation in action understanding, where low-level local features are connected to high-level actions via mid-level atomic actions. This allows the recognition model to work discriminatively. In our ML-HDP, atomic actions are treated as latent topics and automatically discovered from data. In addition, we incorporate the notion of class labels into our model in a semi-supervised fashion to effectively learn and infer multi-labeled videos. Using discovered topics and inferred labels, which are jointly assigned to local features, we present the straightforward methods to perform three recognition tasks including action classification, joint classification and segmentation of continuous actions, and spatiotemporal action localization. In experiments, we explore the use of three different features and demonstrate the effectiveness of our proposed approach for these tasks on four public datasets: KTH, MSR-II, Hollywood2, and UCF101.

引用

页码：800 / 814

页数：15