Short Utterance-based Video Aided Speaker Recognition

被引：0

作者：

Larcher, Anthony ^{[1
]}

Bonastre, Jean-Francois ^{[2
]}

Mason, John S. D.

机构：

[1] Univ Avignon, LIA, 339 Ch Meinajaries,BP1228, F-84911 Avignon, France

[2] Swansea Univ, Speech & Image Grp, Swansea SA2 8PP, W Glam, Wales

来源：

2008 IEEE 10TH WORKSHOP ON MULTIMEDIA SIGNAL PROCESSING, VOLS 1 AND 2 | 2008年

关键词：

D O I：

暂无

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Embedded speaker recognition in mobile devices could involve several ergonomic constraints and a limited amount of computing resources. Even if they have proved their efficiency in more classical contexts, GMM/UBM based systems show their limits in such situations, with good accuracy demanding a relatively large quantity of speech data, but with negligible harnessing of linguistic content The proposed approach addresses these limitations and takes advantage of the linguistic nature of the speech material into the GMM/UBM framework by using client-customised utterances. Furthermore, the acoustic structure is then reinforced with video information. Experiments on the MyIdea database are performed when impostors know the client utterance and also when they do not, highlighting the potential of this new approach. A relative gain up to 47% in terms of EER is achieved when impostors do not know the client utterance and performance is equivalent to the GMM/UBM baseline system in other configurations.

引用

页码：901 / +

页数：2

共 15 条

[1] [Anonymous], 2005, P INTERSPEECH
[2] [Anonymous], 1992, P ICASSP
[3] BENGIO S, 2003, AUDIOAND VIDEOBASED, V2688, P1056
[4] User-customized password speaker verification using multiple reference and background models
BenZeghiba, Mohamed Faouzi
Bourlard, Herve
[J]. SPEECH COMMUNICATION, 2006, 48 (09) : 1200 - 1213
[5] A tutorial on text-independent speaker verification
Bimbot, F
Bonastre, JF
Fredouille, C
Gravier, G
Magrin-Chagnolleau, I
Meignier, S
Merlin, T
Ortega-García, J
Petrovska-Delacrétaz, D
Reynolds, DA
[J]. EURASIP JOURNAL ON APPLIED SIGNAL PROCESSING, 2004, 2004 (04) : 430 - 451
[6] BONASTRE JF, 2003, EUR C SPEECH COMM TE
[7] Multimodal speaker/speech recognition using lip motion, lip texture and audio
Cetingul, H. E.
Erzin, E.
Yemez, Y.
Tekalp, A. M.
[J]. SIGNAL PROCESSING, 2006, 86 (12) : 3549 - 3558
[8] Chibelushi CC, 1997, IEE CONF PUBL, P399, DOI 10.1049/cp:19970924
[9] MAXIMUM LIKELIHOOD FROM INCOMPLETE DATA VIA EM ALGORITHM
DEMPSTER, AP
LAIRD, NM
RUBIN, DB
[J]. JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-METHODOLOGICAL, 1977, 39 (01): : 1 - 38
[10] Audio-visual person authentication using lip-motion from orientation maps
Faraj, Maycel-Isaac
Bigun, Josef
[J]. PATTERN RECOGNITION LETTERS, 2007, 28 (11) : 1368 - 1382

← 1 2 →