MULTI-MODAL FEATURE FUSION FOR ACTION RECOGNITION IN RGB-D SEQUENCES

被引:0
作者
Shahroudy, Amir [1 ,2 ]
Wang, Gang [1 ]
Ng, Tian-Tsong [2 ]
机构
[1] Nanyang Technol Univ, Sch Elect & Elect Engn, Singapore 639798, Singapore
[2] ASTAR, Inst Infocomm Res, Singapore, Singapore
来源
2014 6TH INTERNATIONAL SYMPOSIUM ON COMMUNICATIONS, CONTROL AND SIGNAL PROCESSING (ISCCSP) | 2014年
关键词
Action Recognition; Kinect; Feature Fusion; Structured Sparsity;
D O I
暂无
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Microsoft Kinect's output is a multi-modal signal which gives RGB videos, depth sequences and skeleton information simultaneously. Various action recognition techniques focused on different single modalities of the signals and built their classifiers over the features extracted from one of these channels. For better recognition performance, it's desirable to fuse these multi-modal information into an integrated set of discriminative features. Most of current fusion methods merged heterogeneous features in a holistic manner and ignored the complementary properties of these modalities in finer levels. In this paper, we proposed a new hierarchical bag-of-words feature fusion technique based on multi-view structured sparsity learning to fuse atomic features from RGB and skeletons for the task of action recognition.
引用
收藏
页码:73 / 76
页数:4
相关论文
共 12 条
[1]  
[Anonymous], 2010, CVPR
[2]  
[Anonymous], 2011, CVPR
[3]  
[Anonymous], PAMI
[4]  
[Anonymous], 2013, ICCVW
[5]  
[Anonymous], 2011, CVPR
[6]  
[Anonymous], IROS
[7]  
[Anonymous], ECCVW
[8]   Histograms of oriented gradients for human detection [J].
Dalal, N ;
Triggs, B .
2005 IEEE COMPUTER SOCIETY CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, VOL 1, PROCEEDINGS, 2005, :886-893
[9]  
Dalal N., 2006, ECCV
[10]  
Wang H., 2013, ICML