Improving speech embedding using crossmodal transfer learning with audio-visual data

被引:4
作者
Nam Le [1 ,2 ]
Odobez, Jean-Marc [1 ,3 ]
机构
[1] Idiap Res Inst, Martigny, Switzerland
[2] Ecole Polytech Fed Lausanne, Lausanne, Switzerland
[3] Ecole Polytech Fed Lausanne, MER, Percept & Activ Understanding Grp, Lausanne, Switzerland
基金
欧盟地平线“2020”;
关键词
Speaker diariazation; Multimodal identification; Metric learning; Transfer learning; Deep learning; SPEAKER; FACE;
D O I
10.1007/s11042-018-6992-3
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Learning a discriminative voice embedding allows speaker turns to be compared directly and efficiently, which is crucial for tasks such as diarization and verification. This paper investigates several transfer learning approaches to improve a voice embedding using knowledge transferred from a face representation. The main idea of our crossmodal approaches is to constrain the target voice embedding space to share latent attributes with the source face embedding space. The shared latent attributes can be formalized as geometric properties or distribution characterics between these embedding spaces. We propose four transfer learning approaches belonging to two categories: the first category relies on the structure of the source face embedding space to regularize at different granularities the speaker turn embedding space. The second category -a domain adaptation approach- improves the embedding space of speaker turns by applying a maximum mean discrepancy loss to minimize the disparity between the distributions of the embedded features. Experiments are conducted on TV news datasets, REPERE and ETAPE, to demonstrate our methods. Quantitative results in verification and clustering tasks show promising improvement, especially in cases where speaker turns are short or the training data size is limited. The analysis also gives insights the embedding spaces and shows their potential applications.
引用
收藏
页码:15681 / 15704
页数:24
相关论文
共 52 条
[1]  
[Anonymous], 2012, P INT
[2]  
[Anonymous], LREC
[3]  
[Anonymous], 2016, ARXIV160306432
[4]  
[Anonymous], CVPR
[5]  
[Anonymous], 2016, ARXIV160200955
[6]  
[Anonymous], 2016, CVPR
[7]  
[Anonymous], 2016, CVPR IEEE
[8]  
Baktashmotlagh M, 2016, J MACH LEARN RES, V17
[9]   Multistage speaker diarization of broadcast news [J].
Barras, Claude ;
Zhu, Xuan ;
Meignier, Sylvain ;
Gauvain, Jean-Luc .
IEEE TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2006, 14 (05) :1505-1512
[10]  
Bendris Meriem, 2014, 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), P494, DOI 10.1109/ICASSP.2014.6853645