Improved Class Definition in Two Dimensional Linear Discriminant Analysis of Speech

被引：0

作者：

Conka, David ^{[1
]}

Viszlay, Peter ^{[1
]}

Juhar, Jozef ^{[1
]}

机构：

[1] Tech Univ Kosice, Dept Elect & Multimedia Commun, Kosice, Slovakia

来源：

2015 25TH INTERNATIONAL CONFERENCE RADIOELEKTRONIKA (RADIOELEKTRONIKA) | 2015年

关键词：

cluster analysis; discriminant analysis; scatter matrix; triphone-level class;

D O I：

暂无

中图分类号：

TM [电工技术]; TN [电子技术、通信技术];

学科分类号：

0808 ; 0809 ;

摘要：

Two-dimensional linear discriminant analysis (2DLDA) is a popular feature transformation being applied in current automatic speech recognition (ASR). The parameters of 2DLDA are usually computed on labelled training data partitioned into phonetic classes. It is generally known that one phonetic class contains speech data collected from different speakers with different speech variability and context for the same phonetic unit. Therefore, many clusters exist in each phonetic class. The mentioned effects are not taken into account in the conventional 2DLDA. In this paper, we present an efficient improvement of 2DLDA, which involves the well-known K-means clustering technique to modify the standard class definition. The clustering algorithm is used to identify the existing clusters in the basic classes, which are treated as the new classes for the subsequent 2DLDA estimation. The proposed method is thoroughly evaluated in Slovak triphone-based large vocabulary continuous speech recognition (LVCSR) task. The modified 2DLDA is compared to the state-of-the-art Mel-frequency cepstral coefficients (MFCCs) and to conventional LDA. The results show that the modified 2DLDA features outperform the MFCCs, LDA and also lead to improvement over the conventional 2DLDA.

引用

页码：261 / 263

页数：3

共 10 条

[1]

[Anonymous], 2005, Advances in neural information processing systems. p, DOI DOI 10.5555/2976040.2976237

[2]

[Anonymous], NEW TECHNOLOGIES TRE

[3]

Chen SB, 2008, INT CONF ACOUST SPEE, P4701

[4]

Darjaa S, 2011, 12TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION 2011 (INTERSPEECH 2011), VOLS 1-5, P1728

[5]

Dijun Luo, 2009, 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), P2820, DOI 10.1109/CVPRW.2009.5206635

[6]

Juhar J., 2012, MODERN SPEECH RECOGN, P131

[7]

Jun-Ying Gan, 2009, 2009 1st International Conference on Information Science and Engineering (ICISE 2009), P852, DOI 10.1109/ICISE.2009.582

[8]

Li X. B., 2007, P ANN C INT SPEECH C, P1126

[9]

Viszlay P., 2014, P 22 EUR SIGN PROC C, P1

[10]

Young S., 2006, HTK BOOK HTK VERSION

← 1 →