VARIATIONAL AUTOENCODER FOR SPEECH ENHANCEMENT WITH A NOISE-AWARE ENCODER

被引:39
作者
Fang, Huajian [1 ,2 ]
Carbajal, Guillaume [1 ]
Wermter, Stefan [2 ]
Gerkmann, Timo [1 ]
机构
[1] Univ Hamburg, Signal Proc SP, Hamburg, Germany
[2] Univ Hamburg, Knowledge Technol WTM, Hamburg, Germany
来源
2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021) | 2021年
关键词
speech enhancement; generative model; variational autoencoder; semi-supervised learning;
D O I
10.1109/ICASSP39728.2021.9414060
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
Recently, a generative variational autoencoder (VAE) has been proposed for speech enhancement to model speech statistics. However, this approach only uses clean speech in the training phase, making the estimation particularly sensitive to noise presence, especially in low signal-to-noise ratios (SNRs). To increase the robustness of the VAE, we propose to include noise information in the training phase by using a noise-aware encoder trained on noisy-clean speech pairs. We evaluate our approach on real recordings of different noisy environments and acoustic conditions using two different noise datasets. We show that our proposed noise-aware VAE outperforms the standard VAE in terms of overall distortion without increasing the number of model parameters. At the same time, we demonstrate that our model is capable of generalizing to unseen noise conditions better than a supervised feedforward deep neural network (DNN). Furthermore, we demonstrate the robustness of the model performance to a reduction of the noisy-clean speech training data size.
引用
收藏
页码:676 / 680
页数:5
相关论文
共 26 条
[1]  
Bando Y, 2018, 2018 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), P716, DOI 10.1109/ICASSP.2018.8461530
[2]  
Dean D, 2010, 11TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION 2010 (INTERSPEECH 2010), VOLS 3 AND 4, P3110
[3]   Nonnegative Matrix Factorization with the Itakura-Saito Divergence: With Application to Music Analysis [J].
Fevotte, Cedric ;
Bertin, Nancy ;
Durrieu, Jean-Louis .
NEURAL COMPUTATION, 2009, 21 (03) :793-830
[4]  
Garofolo J., 1993, CSR-I (WSJ0) Complete
[5]   Unbiased MMSE-Based Noise Power Estimation With Low Complexity and Low Tracking Delay [J].
Gerkmann, Timo ;
Hendriks, Richard C. .
IEEE TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2012, 20 (04) :1383-1393
[6]  
Goodfellow IJ, 2014, ADV NEUR IN, V27, P2672
[7]  
Hendriks R. C., 2013, DFT-Domain Based SingleMicrophone Noise Reduction for Speech Enhancement: A Survey of the State-of-the-Art
[8]  
Kingma D. P., 2014, 2 INT C LEARN REPR
[9]  
Kingma DP, 2014, ADV NEUR IN, V27
[10]  
Kohl SAA, 2018, ADV NEUR IN, V31