Cross-Corpus Acoustic Emotion Recognition: Variances and Strategies

被引:267
作者
Schuller, Bjoern [1 ]
Vlasenko, Bogdan [2 ]
Eyben, Florian [1 ]
Woellmer, Martin [1 ]
Stuhlsatz, Andre [3 ]
Wendemuth, Andreas [2 ]
Rigoll, Gerhard [1 ]
机构
[1] Tech Univ Munich, Inst Human Machine Commun, D-80333 Munich, Germany
[2] OVGU, Cognit Syst Grp, IESK, D-39106 Magdeburg, Germany
[3] Univ Appl Sci Dusseldorf, Dept Elect Engn, Lab Pattern Recognit, Dusseldorf, Germany
关键词
Affective computing; speech emotion recognition; cross-corpus evaluation; normalization; SPEECH; CLASSIFICATION; ANNOTATION; EXPRESSION; PITCH;
D O I
10.1109/T-AFFC.2010.8
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
As the recognition of emotion from speech has matured to a degree where it becomes applicable in real-life settings, it is time for a realistic view on obtainable performances. Most studies tend to overestimation in this respect: Acted data is often used rather than spontaneous data, results are reported on preselected prototypical data, and true speaker disjunctive partitioning is still less common than simple cross-validation. Even speaker disjunctive evaluation can give only a little insight into the generalization ability of today's emotion recognition engines since training and test data used for system development usually tend to be similar as far as recording conditions, noise overlay, language, and types of emotions are concerned. A considerably more realistic impression can be gathered by interset evaluation: We therefore show results employing six standard databases in a cross-corpora evaluation experiment which could also be helpful for learning about chances to add resources for training and overcoming the typical sparseness in the field. To better cope with the observed high variances, different types of normalization are investigated. 1.8 k individual evaluations in total indicate the crucial performance inferiority of inter to intracorpus testing.
引用
收藏
页码:119 / 131
页数:13
相关论文
共 105 条
  • [1] [Anonymous], P INT C MULT EXP
  • [2] [Anonymous], 2003, P HUMAN COMPUTER INT
  • [3] [Anonymous], P INT C AC SPEECH SI
  • [4] [Anonymous], ACOUST SPEECH SIG PR
  • [5] [Anonymous], 2009, P INT C AFF COMP INT
  • [6] [Anonymous], P 4 IEEE TUT RES WOR
  • [7] [Anonymous], 2008, Proceedings of the 33rd World Small Animal Veterinary Congress WSAVA
  • [8] [Anonymous], 1997, P 5 EUROPEAN C SPEEC, DOI DOI 10.21437/EUROSPEECH.1997-494
  • [9] [Anonymous], P IEEE CS C COMP VIS
  • [10] [Anonymous], TECHNICAL REPORT