Dithering techniques in automatic recognition of speech corrupted by MP3 compression: Analysis, solutions and experiments

被引：4

作者：

Borsky, Michal ^{[1
]}

Mizera, Petr ^{[1
]}

Pollak, Petr ^{[1
]}

Nouza, Jan ^{[2
]}

机构：

[1] Czech Tech Univ, Fac Elect Engn, Prague, Czech Republic

[2] TUL, Inst Informat Technol & Elect, Liberec, Czech Republic

来源：

SPEECH COMMUNICATION | 2017年 / 86卷

关键词：

Uniform dithering; Spectrally selective dithering; MP3; compression; GMM-HMM; DNN-HMM;

D O I：

10.1016/j.specom.2016.11.007

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

A large portion of the audio files distributed over the Internet or those stored in personal and corporate media archives are in a compressed form. There exist several compression techniques and algorithms but it is the MPEG Layer-3 (known as MP3) that has achieved a really wide popularity in general audio coding, and in speech, too. However, the algorithm is lossy in nature and introduces distortion into spectral and temporal characteristics of a signal. In this paper we study its impact on automatic speech recognition (ASR). We show that with decreasing MP3 bitrates the major source of ASR performance degradation is deep spectral valleys (i.e. bins with almost zero energy) caused by the masking effect of the MP3 algorithm. We demonstrate that these unnatural gaps in spectrum can be effectively compensated by adding a certain amount of noise to the distorted signal. We provide theoretical background for this approach where we show that the added noise affects mainly the spectral valleys. They are filled by the noise while the spectral bins with speech remain almost unchanged. This helps to restore a more natural shape of log spectrum and cepstrum, and consequently has a positive impact on ASR performance. In our previous work, we have proposed two types of the signal dithering (noise addition) technique, one applied globally, the other in a more selective way. In this paper, we offer a more detailed insight into their performance. We provide results from many experiments where we test them in various scenarios, using a large vocabulary continuous speech recognition (LVCSR) system, acoustic models based on gaussian-mixture model (GMM) as well as on deep-neural network (DNN), and multiple speech databases in three languages (Czech, English and German). Our results prove that both the proposed techniques, and the selective dithering method, in particular, yield consistent compensation of the negative impact of the MP3 compressed speech on ASR performance. (C) 2016 Elsevier B.V. All rights reserved.

引用

页码：75 / 84

页数：10

共 2 条

[1] Advanced acoustic modelling techniques in MP3 speech recognition
Borsky, Michal
Pollak, Petr
Mizera, Petr
EURASIP JOURNAL ON AUDIO SPEECH AND MUSIC PROCESSING, 2015,
[2] Advanced acoustic modelling techniques in MP3 speech recognition
Michal Borsky
Petr Pollak
Petr Mizera
EURASIP Journal on Audio, Speech, and Music Processing, 2015

← 1 →