Detecting Converted Speech and Natural Speech for anti-Spoofing Attack in Speaker Recognition

被引:0
作者
Wu, Zhizheng [1 ]
Chng, Eng Siong [1 ]
Li, Haizhou [1 ]
机构
[1] Nanyang Technol Univ, Sch Comp Engn, Singapore 639798, Singapore
来源
13TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION 2012 (INTERSPEECH 2012), VOLS 1-3 | 2012年
关键词
Speaker verification; voice conversion; anti-spoofing attack; synthetic speech detection; phase spectrum; HUMAN LISTENING TESTS; TIME PHASE SPECTRUM;
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Voice conversion techniques present a threat to speaker verification systems. To enhance the security of speaker verification systems, We study how to automatically distinguish natural speech and synthetic/converted speech. Motivated by the research on phase spectrum in speech perception, in this study, we propose to use features derived from phase spectrum to detect converted speech. The features are tested under three different training situations of the converted speech detector: a) only Gaussian mixture model (GMM) based converted speech data are available; b) only unit-selection based converted speech data are available; c) no converted speech data are available for training converted speech model. Experiments conducted on the National Institute of Standards and Technology (NIST) 2006 speaker recognition evaluation (SRE) corpus show that the performance of the features derived from phase spectrum outperform the mel-frequency cepstral coefficients (MFCCs) tremendously: even without converted speech for training, the equal error rate (EER) is reduced from 20.20% of MFCCs to 2.35%.
引用
收藏
页码:1698 / 1701
页数:4
相关论文
共 21 条
[1]   Short-time phase spectrum in speech processing: A review and some experimental results [J].
Alsteris, Leigh D. ;
Paliwal, Kuldip K. .
DIGITAL SIGNAL PROCESSING, 2007, 17 (03) :578-616
[2]   Further intelligibility results from human listening tests using the short-time phase spectrum [J].
Alsteris, Leigh D. ;
Paliwal, Kuldip K. .
SPEECH COMMUNICATION, 2006, 48 (06) :727-736
[3]  
Bonastre J.-F., 2007, INTERSPEECH
[4]   Speaker recognition: A tutorial [J].
Campbell, JP .
PROCEEDINGS OF THE IEEE, 1997, 85 (09) :1437-1462
[5]  
De Leon P.L., ICASSP 2011
[6]  
DeLeon P., ODYSSEY 2010
[7]  
Fukada T., ICASSP 1992
[8]   Significance of the modified group delay feature in speech recognition [J].
Hegde, Rajesh M. ;
Murthy, Hema A. ;
Gadde, Venkata Ramana Rao .
IEEE TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2007, 15 (01) :190-202
[9]  
Jin Q., ICASSP 2008
[10]  
Jin Q., ICASSP 2007