Towards understanding action recognition

被引：594

作者：

Jhuang, Hueihan ^{[1
]}

Gall, Juergen ^{[2
]}

Zuffi, Silvia ^{[3
]}

Schmid, Cordelia ^{[4
]}

Black, Michael J. ^{[1
]}

机构：

[1] MPI Intelligent Syst, Stuttgart, Germany

[2] Univ Bonn, Bonn, Germany

[3] Brown Univ, Providence, RI 02912 USA

[4] INRIA, LEAR, Paris, France

来源：

2013 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV) | 2013年

关键词：

HISTOGRAMS;

D O I：

10.1109/ICCV.2013.396

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Although action recognition in videos is widely studied, current methods often fail on real-world datasets. Many recent approaches improve accuracy and robustness to cope with challenging video sequences, but it is often unclear what affects the results most. This paper attempts to provide insights based on a systematic performance evaluation using thoroughly-annotated data of human actions. We annotate human Joints for the HMDB dataset (J-HMDB). This annotation can be used to derive ground truth optical flow and segmentation. We evaluate current methods using this dataset and systematically replace the output of various algorithms with ground truth. This enables us to discover what is important - for example, should we work on improving flow algorithms, estimating human bounding boxes, or enabling pose estimation? In summary, we find that high-level pose features greatly outperform low/mid level features; in particular, pose over time is critical. While current pose estimation algorithms are far from perfect, features extracted from estimated pose on a subset of J-HMDB, in which the full body is visible, outperform low/mid-level features. We also find that the accuracy of the action recognition framework can be greatly increased by refining the underlying low/mid level features; this suggests it is important to improve optical flow and human detection algorithms. Our analysis and J-HMDB dataset should facilitate a deeper understanding of action recognition algorithms.

引用

页码：3192 / 3199

页数：8

共 36 条

[1]

[Anonymous], 2005, PROC CVPR IEEE

[2]

[Anonymous], COMPUT VIS PATT RECO, DOI DOI 10.1109/CVPR.2009.5206557

[3]

[Anonymous], 2013, TRISMPI007

[4] Poselets: Body Part Detectors Trained Using 3D Human Pose Annotations [J].

Bourdev, Lubomir ;

Malik, Jitendra .

2009 IEEE 12TH INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2009, :1365-1372

[5]

Bourdev L, 2010, LECT NOTES COMPUT SC, V6316, P168, DOI 10.1007/978-3-642-15567-3_13

[6]

CAMPBELL LW, 1995, FIFTH INTERNATIONAL CONFERENCE ON COMPUTER VISION, PROCEEDINGS, P624, DOI 10.1109/ICCV.1995.466880

[7] LIBSVM: A Library for Support Vector Machines [J].

Chang, Chih-Chung ;

Lin, Chih-Jen .

ACM TRANSACTIONS ON INTELLIGENT SYSTEMS AND TECHNOLOGY, 2011, 2 (03)

[8] Human detection using oriented histograms of flow and appearance [J].

Dalal, Navneet ;

Triggs, Bill ;

Schmid, Cordelia .

COMPUTER VISION - ECCV 2006, PT 2, PROCEEDINGS, 2006, 3952 :428-441

[9] Human Pose Estimation using Body Parts Dependent Joint Regressors [J].

Dantone, Matthias ;

Gall, Juergen ;

Leistner, Christian ;

Van Gool, Luc .

2013 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2013, :3041-3048

[10]

Eichner M., 2009, BETTER APPEARANCE MO

← 1 2 3 4 →