Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition

被引:116
作者
Wang, Yongjin [1 ]
Guan, Ling [2 ]
Venetsanopoulos, Anastasios N. [2 ]
机构
[1] Hisense Co Ltd, State Key Lab Digital Multimedia Technol, Qingdao, Peoples R China
[2] Ryerson Univ, Dept Elect & Comp Engn, Toronto, ON M5B 2K3, Canada
关键词
Cross-modal association; emotion recognition; information fusion; kernel machine technique; FACE; MODELS;
D O I
10.1109/TMM.2012.2189550
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
In this paper, we investigate kernel based methods for multimodal information analysis and fusion. We introduce a novel approach, kernel cross-modal factor analysis, which identifies the optimal transformations that are capable of representing the coupled patterns between two different subsets of features by minimizing the Frobenius norm in the transformed domain. The kernel trick is utilized for modeling the nonlinear relationship between two multidimensional variables. We examine and compare with kernel canonical correlation analysis which finds projection directions that maximize the correlation between two modalities, and kernel matrix fusion which integrates the kernel matrices of respective modalities through algebraic operations. The performance of the introduced method is evaluated on an audiovisual based bimodal emotion recognition problem. We first perform feature extraction from the audio and visual channels respectively. The presented approaches are then utilized to analyze the cross-modal relationship between audio and visual features. A hidden Markov model is subsequently applied for characterizing the statistical dependence across successive time segments, and identifying the inherent temporal structure of the features in the transformed domain. The effectiveness of the proposed solution is demonstrated through extensive experimentation.
引用
收藏
页码:597 / 607
页数:11
相关论文
共 39 条
[1]  
Aldea E, 2007, LECT NOTES COMPUT SC, V4842, P307
[2]  
[Anonymous], 2003, P ACM INT C MULT ACM
[3]  
[Anonymous], 2008, P 2008 16 EUR SIGN P
[4]  
[Anonymous], P 3 INT C AN MOD FAC
[5]   Multimodal fusion for multimedia analysis: a survey [J].
Atrey, Pradeep K. ;
Hossain, M. Anwar ;
El Saddik, Abdulmotaleb ;
Kankanhalli, Mohan S. .
MULTIMEDIA SYSTEMS, 2010, 16 (06) :345-379
[6]   Generalized discriminant analysis using a kernel approach [J].
Baudat, G ;
Anouar, FE .
NEURAL COMPUTATION, 2000, 12 (10) :2385-2404
[7]  
Blaschko M. B., 2008, P IEEE CVPR, P1
[8]  
Bredin H, 2007, INT CONF ACOUST SPEE, P233
[9]  
Chan CH, 2010, LECT NOTES COMPUT SC, V6218, P718, DOI 10.1007/978-3-642-14980-1_71
[10]   Canonical Correlation Analysis for Data Fusion and Group Inferences [J].
Correa, Nicolle M. ;
Adali, Tulay ;
Li, Yi-Ou ;
Calhoun, Vince D. .
IEEE SIGNAL PROCESSING MAGAZINE, 2010, 27 (04) :39-50