Short Utterance-based Video Aided Speaker Recognition

被引:0
作者
Larcher, Anthony [1 ]
Bonastre, Jean-Francois [2 ]
Mason, John S. D.
机构
[1] Univ Avignon, LIA, 339 Ch Meinajaries,BP1228, F-84911 Avignon, France
[2] Swansea Univ, Speech & Image Grp, Swansea SA2 8PP, W Glam, Wales
来源
2008 IEEE 10TH WORKSHOP ON MULTIMEDIA SIGNAL PROCESSING, VOLS 1 AND 2 | 2008年
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Embedded speaker recognition in mobile devices could involve several ergonomic constraints and a limited amount of computing resources. Even if they have proved their efficiency in more classical contexts, GMM/UBM based systems show their limits in such situations, with good accuracy demanding a relatively large quantity of speech data, but with negligible harnessing of linguistic content The proposed approach addresses these limitations and takes advantage of the linguistic nature of the speech material into the GMM/UBM framework by using client-customised utterances. Furthermore, the acoustic structure is then reinforced with video information. Experiments on the MyIdea database are performed when impostors know the client utterance and also when they do not, highlighting the potential of this new approach. A relative gain up to 47% in terms of EER is achieved when impostors do not know the client utterance and performance is equivalent to the GMM/UBM baseline system in other configurations.
引用
收藏
页码:901 / +
页数:2
相关论文
共 15 条
  • [1] [Anonymous], 2005, P INTERSPEECH
  • [2] [Anonymous], 1992, P ICASSP
  • [3] BENGIO S, 2003, AUDIOAND VIDEOBASED, V2688, P1056
  • [4] User-customized password speaker verification using multiple reference and background models
    BenZeghiba, Mohamed Faouzi
    Bourlard, Herve
    [J]. SPEECH COMMUNICATION, 2006, 48 (09) : 1200 - 1213
  • [5] A tutorial on text-independent speaker verification
    Bimbot, F
    Bonastre, JF
    Fredouille, C
    Gravier, G
    Magrin-Chagnolleau, I
    Meignier, S
    Merlin, T
    Ortega-García, J
    Petrovska-Delacrétaz, D
    Reynolds, DA
    [J]. EURASIP JOURNAL ON APPLIED SIGNAL PROCESSING, 2004, 2004 (04) : 430 - 451
  • [6] BONASTRE JF, 2003, EUR C SPEECH COMM TE
  • [7] Multimodal speaker/speech recognition using lip motion, lip texture and audio
    Cetingul, H. E.
    Erzin, E.
    Yemez, Y.
    Tekalp, A. M.
    [J]. SIGNAL PROCESSING, 2006, 86 (12) : 3549 - 3558
  • [8] Chibelushi CC, 1997, IEE CONF PUBL, P399, DOI 10.1049/cp:19970924
  • [9] MAXIMUM LIKELIHOOD FROM INCOMPLETE DATA VIA EM ALGORITHM
    DEMPSTER, AP
    LAIRD, NM
    RUBIN, DB
    [J]. JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-METHODOLOGICAL, 1977, 39 (01): : 1 - 38
  • [10] Audio-visual person authentication using lip-motion from orientation maps
    Faraj, Maycel-Isaac
    Bigun, Josef
    [J]. PATTERN RECOGNITION LETTERS, 2007, 28 (11) : 1368 - 1382