Modulation Spectral Features for Robust Far-Field Speaker Identification

被引：75

作者：

Falk, Tiago H. ^{[1
]}

Chan, Wai-Yip ^{[1
]}

机构：

[1] Queens Univ, Dept Elect & Comp Engn, Kingston, ON K7L 3N6, Canada

来源：

IEEE TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING | 2010年 / 18卷 / 01期

关键词：

Gaussian mixture model (GMM); modulation spectrum; reverberation; reverberation time; speaker identification; RECOGNITION;

D O I：

10.1109/TASL.2009.2023679

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

In this paper, auditory inspired modulation spectral features are used to improve automatic speaker identification (ASI) performance in the presence of room reverberation. The modulation spectral signal representation is obtained by first filtering the speech signal with a 23-channel gammatone filterbank. An eight-channel modulation filterbank is then applied to the temporal envelope of each gammatone filter output. Features are extracted from modulation frequency bands ranging from 3-15 Hz and are shown to be robust to mismatch between training and testing conditions and to increasing reverberation levels. To demonstrate the gains obtained with the proposed features, experiments are performed with clean speech, artificially generated reverberant speech, and reverberant speech recorded in a meeting room. Simulation results show that a Gaussian mixture model based ASI system, trained on the proposed features, consistently outperforms a baseline system trained on mel-frequency cepstral coefficients. For multimicrophone ASI applications, three multichannel score combination and adaptive channel selection techniques are investigated and shown to further improve ASI performance.

引用

页码：90 / 100

页数：11

共 41 条

[1]

*3GPP2, 1999, CS00140 3GPP2

[2]

ABUELQURAN A, 2007, P IEEE C SENS OCT, P970

[3]

AKULA A, 2008, P EUR SIGN PROC C AU

[4]

[Anonymous], P IEEE C INSTR MEAS

[5]

[Anonymous], P ICSLP

[6]

[Anonymous], P56 ITUT

[7]

Arai T, 1996, ICSLP 96 - FOURTH INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE PROCESSING, PROCEEDINGS, VOLS 1-4, P2490, DOI 10.1109/ICSLP.1996.607318

[8] Analysis of feature extraction and channel compensation in a GMM speaker recognition system [J].

Burget, Lukas ;

Matejka, Pavel ;

Schwarz, Petr ;

Glembek, Ondfei ;

Cernocky, Jan 'Honza' .

IEEE TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2007, 15 (07) :1979-1986

[9]

Castellano PJ, 1996, INT CONF ACOUST SPEE, P117, DOI 10.1109/ICASSP.1996.540304

[10]

CUMMINS F, 2006, P INT C SPEECH COMP

← 1 2 3 4 5 →