Multimodal 2D+3D Facial Expression Recognition With Deep Fusion Convolutional Neural Network

被引:149
作者
Li, Huibin [1 ]
Sun, Jian [1 ]
Xu, Zongben [1 ]
Chen, Liming [2 ]
机构
[1] Xi An Jiao Tong Univ, Inst Informat & Syst Sci, Sch Math & Stat, Xian 710049, Shaanxi, Peoples R China
[2] Ecole Cent Lyon, Dept Math & Informat, LIRIS UMR 5205, F-69134 Lyon, France
关键词
Deep fusion convolutional neural network (DF-CNN); facial expression recognition (FER); multimodal; textured three-dimensional (3D) face scan; EMOTION RECOGNITION; 3D; FACE; FRAMEWORK; DATABASE;
D O I
10.1109/TMM.2017.2713408
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
This paper presents a novel and efficient deep fusion convolutional neural network (DF-CNN) for multimodal 2D+3D facial expression recognition (FER). DF-CNN comprises a feature extraction subnet, a feature fusion subnet, and a softmax layer. In particular, each textured three-dimensional (3D) face scan is represented as six types of 2D facial attribute maps (i.e., geometry map, three normal maps, curvature map, and texture map), all of which are jointly fed into DF-CNN for feature learning and fusion learning, resulting in a highly concentrated facial representation (32-dimensional). Expression prediction is performed by two ways: 1) learning linear support vector machine classifiers using the 32-dimensional fused deep features, or 2) directly performing softmax prediction using the six-dimensional expression probability vectors. Different from existing 3D FER methods, DF-CNN combines feature learning and fusion learning into a single end-to-end training framework. To demonstrate the effectiveness of DF-CNN, we conducted comprehensive experiments to compare the performance of DF-CNN with handcrafted features, pre-trained deep features, fine-tuned deep features, and state-of-the-art methods on three 3D face datasets (i.e., BU-3DFE Subset I, BU-3DFE Subset II, and Bosphorus Subset). In all cases, DF-CNN consistently achieved the best results. To the best of our knowledge, this is the first work of introducing deep CNN to 3D FER and deep learning-based feature level fusion for multimodal 2D+3D FER.
引用
收藏
页码:2816 / 2831
页数:16
相关论文
共 69 条
[1]   Survey on RGB, 3D, Thermal, and Multimodal Approaches for Facial Expression Recognition: History, Trends, and Affect-Related Applications [J].
Adrian Corneanu, Ciprian ;
Oliu Simon, Marc ;
Cohn, Jeffrey F. ;
Escalera Guerrero, Sergio .
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2016, 38 (08) :1548-1568
[2]  
[Anonymous], 2010, IEEE CVPR 10 WORKSHO
[3]  
[Anonymous], 2010, P 18 ACM INT C MULT, DOI [10.1145/1873951.1874249, 10.1145/1873951.1874249.2]
[4]  
[Anonymous], 2008 23 INT S COMP I
[5]  
[Anonymous], 2015, CORR
[6]  
[Anonymous], 2015, DEEPLY LEARNING DEFO
[7]  
[Anonymous], 2014, P 31 INT C INT C MAC
[8]  
[Anonymous], 2013, 2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG)
[9]  
Berretti S., 2010, Proceedings of the 2010 20th International Conference on Pattern Recognition (ICPR 2010), P4125, DOI 10.1109/ICPR.2010.1002
[10]   Emotion Recognition in Text for 3-D Facial Expression Rendering [J].
Calix, Ricardo A. ;
Mallepudi, Sri Abhishikth ;
Chen, Bin ;
Knapp, Gerald M. .
IEEE TRANSACTIONS ON MULTIMEDIA, 2010, 12 (06) :544-551