Hidden-articulator Markov models for speech recognition

被引：33

作者：

Richardson, M ^{[1
]}

Bilmes, J ^{[1
]}

Diorio, C ^{[1
]}

机构：

[1] Univ Washington, Dept Comp Sci & Engn, Seattle, WA 98195 USA

来源：

SPEECH COMMUNICATION | 2003年 / 41卷 / 2-3期

关键词：

speech recognition; articulatory models; noise robustness; factorial HMM;

D O I：

10.1016/S0167-6393(03)00031-1

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

Most existing automatic speech recognition systems today do not explicitly use knowledge about human speech production. We show that the incorporation of articulatory knowledge into these systems is a promising direction for speech recognition, with the potential for lower error rates and more robust performance. To this end, we introduce the Hidden-Articulator Markov model (HAMM), a model which directly integrates articulatory information into speech recognition. The HAMM is an extension of the articulatory-feature model introduced by Erler in 1996. We extend the model by using diphone units, developing a new technique for model initialization, and constructing a novel articulatory feature mapping. We also introduce a method to decrease the number of parameters, making the HAMM comparable in size to standard HMMs. We demonstrate that the HAMM can reasonably predict the movement of articulators, which results in a decreased word error rate (WER). The articulatory knowledge also proves useful in noisy acoustic conditions. When combined with a standard model, the HAMM reduces WER 28-35% relative to the standard model alone. (C) 2003 Elsevier B.V. All rights reserved.

引用

页码：511 / 529

页数：19

共 34 条

[1]

BAILLY G, 1992, SIGNAL PROCESSING 6, V1, P159

[2]

Bilmes J., 2000, P 16 C UNC ART INT, P38

[3]

BILMES JA, 1999, ICASSP, V2, P713

[4]

Bishop C. M., 1995, NEURAL NETWORKS PATT

[5]

BLACKBURN C, 1995, P EUR, V2, P1623

[6]

BLOMBERG M, 1991, P EUR

[7] Production models as a structural basis for automatic speech recognition [J].

Deng, L ;

Ramsay, G ;

Sun, D .

SPEECH COMMUNICATION, 1997, 22 (2-3) :93-111

[8] A dynamic, feature-based approach to the interface between phonology and phonetics for speech modeling and recognition [J].

Deng, L .

SPEECH COMMUNICATION, 1998, 24 (04) :299-323

[9] A STATISTICAL APPROACH TO AUTOMATIC SPEECH RECOGNITION USING THE ATOMIC SPEECH UNITS CONSTRUCTED FROM OVERLAPPING ARTICULATORY FEATURES [J].

DENG, L ;

SUN, DX .

JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1994, 95 (05) :2702-2719

[10]

DENG L, 1994, INT CONF ACOUST SPEE, P45

← 1 2 3 4 →