Genetic Algorithm-Based Adaptive Wiener Gain for Speech Enhancement Using an Iterative Posterior NMF

被引：3

作者：

Yechuri, Sivaramakrishna ^{[1
]}

Vanabathina, Sunny Dayal ^{[1
]}

机构：

[1] VIT AP, Sch Elect Engn, Amaravati, Andhra Pradesh, India

来源：

INTERNATIONAL JOURNAL OF IMAGE AND GRAPHICS | 2023年 / 23卷 / 06期

关键词：

NMF; adaptive Wiener gain; inverse gamma; Students-t; SDR; PESQ; STOI; NONNEGATIVE MATRIX FACTORIZATION; QUALITY; NOISE; MACHINE;

D O I：

10.1142/S0219467823500547

中图分类号：

TP31 [计算机软件];

学科分类号：

081202 ; 0835 ;

摘要：

In this paper, we propose a genetic algorithm-based adaptive Wiener gain for speech enhancement using an iterative posterior non-negative matrix factorization (NMF). In the recent past, NMF-based Wiener filtering methods were used to improve the performance of speech enhancement, which has shown that they provide better performance when compared with conventional NMF methods. But performance degrades in non-stationary noise environments. Template-based approaches are more robust and perform better in non-stationary noise environments compared to statistical model-based approaches but are dependent on a priori information. Combining the approaches avoids the drawbacks of both. To improve the performance further, speech and noise bases are adapted simultaneously in the NMF approach. The usage of Super-Gaussian constraints in iterative NMF still improves the performance in non-stationary noise. The silence frame is a challenging task in the case of NMF; still there will be some amount of noise present in those frames. For further enhancement, we have combined with a genetic algorithm (GA)-based adaptive Wiener filter which performs well in denoising and also the GA search the adaptive alpha allows us to control the trade-off between fitting the observed spectrogram of mixed speech and noise achieving high likelihood under our prior model. The proposed method outperforms other benchmark algorithms in terms of the source to distortion ratio (SDR), short-time objective intelligibility (STOI), and perceptual evaluation of speech quality (PESQ).

引用

页数：20

共 40 条

[1] Speech enhancement with an adaptive Wiener filter [J].

Abd El-Fattah, Marwa ;

Dessouky, Moawad ;

Abbas, Alaa ;

Diab, Salaheldin ;

El-Rabaie, El-Sayed ;

Al-Nuaimy, Waleed ;

Alshebeili, Saleh ;

Abd El-Samie, Fathi .

INTERNATIONAL JOURNAL OF SPEECH TECHNOLOGY, 2014, 17 (01) :53-64

[2]

[Anonymous], 2014, P ICASSP

[3] Immersive visualization of visual data using nonnegative matrix factorization [J].

Babaee, Mohammadreza ;

Tsoukalas, Stefanos ;

Rigoll, Gerhard ;

Datcu, Mihai .

NEUROCOMPUTING, 2016, 173 :245-255

[4] Algorithms and applications for approximate nonnegative matrix factorization [J].

Berry, Michael W. ;

Browne, Murray ;

Langville, Amy N. ;

Pauca, V. Paul ;

Plemmons, Robert J. .

COMPUTATIONAL STATISTICS & DATA ANALYSIS, 2007, 52 (01) :155-173

[5] SUPPRESSION OF ACOUSTIC NOISE IN SPEECH USING SPECTRAL SUBTRACTION [J].

BOLL, SF .

IEEE TRANSACTIONS ON ACOUSTICS SPEECH AND SIGNAL PROCESSING, 1979, 27 (02) :113-120

[6]

Bryan Nicholas., 2013, International Conference on Machine Learning, P208

[7] Speech enhancement strategy for speech recognition microcontroller under noisy environments [J].

Chan, Kit Yan ;

Nordholm, Sven ;

Yiu, Ka Fai Cedric ;

Togneri, Roberto .

NEUROCOMPUTING, 2013, 118 :279-288

[8] Generalized Alpha-Beta Divergences and Their Application to Robust Nonnegative Matrix Factorization [J].

Cichocki, Andrzej ;

Cruces, Sergio ;

Amari, Shun-ichi .

ENTROPY, 2011, 13 (01) :134-170

[9] From blind signal extraction to blind instantaneous signal separation: Criteria, algorithms, and stability [J].

Cruces-Alvarez, SA ;

Cichocki, A ;

Amari, SI .

IEEE TRANSACTIONS ON NEURAL NETWORKS, 2004, 15 (04) :859-873

[10]

Duan ZY, 2012, 13TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION 2012 (INTERSPEECH 2012), VOLS 1-3, P594

← 1 2 3 4 →