Speech enhancement using progressive learning-based convolutional recurrent neural network

被引：57

作者：

Li, Andong ^{[1
,2
]}

Yuan, Minmin ^{[3
]}

Zheng, Chengshi ^{[1
,2
]}

Li, Xiaodong ^{[1
,2
]}

机构：

[1] Chinese Acad Sci, Inst Acoust, Key Lab Noise & Vibrat Res, Beijing 100190, Peoples R China

[2] Univ Chinese Acad Sci, Beijing 100049, Peoples R China

[3] Highway Minist Transport, Res Inst, Beijing 100088, Peoples R China

来源：

APPLIED ACOUSTICS | 2020年 / 166卷

关键词：

Speech enhancement; Deep learning; Progressive learning; Convolutional neural network; Long short-term memory; NOISE; RECOGNITION; REDUCTION; ALGORITHM;

D O I：

10.1016/j.apacoust.2020.107347

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

Recently, progressive learning has shown its capacity to improve speech quality and speech intelligibility when it is combined with deep neural network (DNN) and long short-term memory (LSTM) based monaural speech enhancement algorithms, especially in low signal-to-noise ratio (SNR) conditions. Nevertheless, due to a large number of parameters and high computational complexity, it is hard to implement in current resource-limited micro-controllers and thus, it is essential to significantly reduce both the number of parameters and the computational load for practical applications. For this purpose, we propose a novel progressive learning framework with causal convolutional recurrent neural networks called PL-CRNN, which takes advantage of both convolutional neural networks and recurrent neural networks to drastically reduce the number of parameters and simultaneously improve speech quality and speech intelligibility. Numerous experiments verify the effectiveness of the proposed PL-CRNN model and indicate that it yields consistent better performance than the PL-DNN and PL-LSTM algorithms and also it gets results close even better than the CRNN in terms of objective measurements. Compared with PL-DNN, PL-LSTM, and CRNN, the proposed PL-CRNN algorithm can reduce the number of parameters up to 93%, 97%, and 92%, respectively. (C) 2020 Elsevier Ltd. All rights reserved.

引用

页数：9

共 42 条

[1]

[Anonymous], ARXIV151107289

[2] SUPPRESSION OF ACOUSTIC NOISE IN SPEECH USING SPECTRAL SUBTRACTION [J].

BOLL, SF .

IEEE TRANSACTIONS ON ACOUSTICS SPEECH AND SIGNAL PROCESSING, 1979, 27 (02) :113-120

[3] New insights into the noise reduction Wiener filter [J].

Chen, Jingdong ;

Benesty, Jacob ;

Huang, Yiteng ;

Doclo, Simon .

IEEE TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2006, 14 (04) :1218-1234

[4] Long short-term memory for speaker generalization in supervised speech separation [J].

Chen, Jitong ;

Wang, DeLiang .

JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 2017, 141 (06) :4705-4714

[5]

Erdogan H, 2015, INT CONF ACOUST SPEE, P708, DOI 10.1109/ICASSP.2015.7178061

[6]

Fakoor R, 2018, 2018 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), P3011, DOI 10.1109/ICASSP.2018.8462042

[7]

Fu S.-W., 2017, Machine Learning for Signal Processing (MLSP), 2017 IEEE 27th International Workshop on, P1

[8] SNR-Aware Convolutional Neural Network Modeling for Speech Enhancement [J].

Fu, Szu-Wei ;

Tsao, Yu ;

Lu, Xugang .

17TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION (INTERSPEECH 2016), VOLS 1-5: UNDERSTANDING SPEECH PROCESSING IN HUMANS AND MACHINES, 2016, :3768-3772

[9]

Gao Q, 2015, 2015 IEEE INTERNATIONAL CONFERENCE ON INFORMATION AND AUTOMATION, P1371, DOI 10.1109/ICInfA.2015.7279500

[10]

Gao T, 2018, 2018 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), P5054, DOI 10.1109/ICASSP.2018.8461861

← 1 2 3 4 5 →