Monaural speech separation using GA-DNN integration scheme

被引:10
作者
Sivapatham, Shoba [1 ]
Ramadoss, Rajavel [1 ]
Kar, Asutosh [2 ]
Majhi, Banshidhar [3 ]
机构
[1] SSN Coll Engn, Dept Elect & Commun Engn, Kalavakkam, India
[2] Indian Inst Informat Technol Design & Mfg, Dept Elect & Commun Engn, Chennai, Tamil Nadu, India
[3] Indian Inst Informat Technol Design & Mfg, Dept Comp Sci & Engn, Chennai, Tamil Nadu, India
关键词
Genetic Algorithm; Deep Neural Network; Monaural Speech Separation; Segmentation; Voiced Speech; Unvoiced Speech; RECURRENT NEURAL-NETWORKS; PITCH TRACKING; ENHANCEMENT; INTELLIGIBILITY; SYSTEM; NOISE;
D O I
10.1016/j.apacoust.2019.107140
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
In this research work, we propose the model based on the Genetic Algorithm (GA) and Deep Neural Network (DNN) to enhance the quality and intelligibility of the noisy speech. In this proposed model, the Voiced Speech (VS) T-F mask is computed using correlogram, frame energy and cross-channel correlogram and Unvoiced Speech (UVS) T-F mask is computed using speech onset/offset. The T-F mask obtained using speech onset and offset represents both voiced and unvoiced segment of the noisy speech signal. The UVS T-F mask is obtained by subtracting the VS from the T-F mask obtained earlier using speech onset/offset. Next, the GA is used to find the optimum weight to combine the T-F mask of VS and UVS to improve speech quality and intelligibility. The weight obtained using GA may not be an optimum one for all sets of speech and noise. This research work focuses on this issue and proposes a DNN model to estimate the optimum weight for all sets of speech and noise. The DNN model is trained using features and optimum weight obtained using GA. Later, the trained DNN model is used to estimate the optimum weight for the testing speech and noise samples. The performance of the proposed GA-DNN based model is evaluated using objective and subjective quality and intelligibility measures. The results of the proposed model shows a prompt improvement in the speech quality and intelligibility with average of 0.73, 4.07, 0.17, 0.26 and 0.22 for PESQ SNR, STOI, CSII and NCM when compared with the existing speech separation systems. (C) 2019 Elsevier Ltd. All rights reserved.
引用
收藏
页数:11
相关论文
共 50 条
  • [1] Abien Fred M., 2019, ARXIV180308375V2CSNE
  • [2] Alamdari N., 2019, ARXIV190412069
  • [3] [Anonymous], 1996, ITU T RECOMMENDATION, P830
  • [4] [Anonymous], 2011, PERCEPTUAL EVALUATIO
  • [5] [Anonymous], THESIS
  • [6] Anwar MU, INT C EL COMP COMM E
  • [7] SUPPRESSION OF ACOUSTIC NOISE IN SPEECH USING SPECTRAL SUBTRACTION
    BOLL, SF
    [J]. IEEE TRANSACTIONS ON ACOUSTICS SPEECH AND SIGNAL PROCESSING, 1979, 27 (02): : 113 - 120
  • [8] Brown G., 2005, SPEECH ENHANCEMENT
  • [9] COMPUTATIONAL AUDITORY SCENE ANALYSIS
    BROWN, GJ
    COOKE, M
    [J]. COMPUTER SPEECH AND LANGUAGE, 1994, 8 (04) : 297 - 336
  • [10] Duchi J, 2011, J MACH LEARN RES, V12, P2121