Direction-aware target speaker extraction with a dual-channel system based on conditional variational autoencoders under underdetermined conditions

被引：0

作者：

Wang, Rui ^{[1
]}

Li, Li ^{[1
]}

Toda, Tomoki ^{[1
]}

机构：

[1] Nagoya Univ, Nagoya, Aichi, Japan

来源：

PROCEEDINGS OF 2022 ASIA-PACIFIC SIGNAL AND INFORMATION PROCESSING ASSOCIATION ANNUAL SUMMIT AND CONFERENCE (APSIPA ASC) | 2022年

关键词：

multichannel source separation; target speaker extraction; multichannel variational autoencoder (MVAE); INDEPENDENT VECTOR ANALYSIS; AUDIO SOURCE SEPARATION; INFORMATION;

D O I：

暂无

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

In this paper, we deal with a dual-channel target speaker extraction (TSE) problem under underdetermined conditions. For the dual-channel system, the generalized sidelobe canceller (GSC) is a commonly used structure for estimating a blocking matrix (BM) to generate interference, and geometric source separation (GSS) can be used as an implementation of BM estimation utilizing directional information. However, the performance of the conventional GSS methods is limited under underdetermined conditions because of the lack of a powerful source model. In this paper, we propose a dual-channel TSE method that combines the ability of target selection based on geometric constraints, more powerful source modeling, and nonlinear postprocessing. The target directional information is used as a geometric constraint, and two conditional variational autoencoders (CVAEs) are used to model a single speaker's speech and interference mixture speech. For the postprocessing, an ideal ratio Time-Frequency (T-F) mask estimated from the separated interference mixture speech is used to extract the target speaker's speech. The experimental results demonstrate that the proposed method achieves 6.24 dB and 8.37 dB improvements compared with the baseline method in terms of signal-to-distortions ratio (SDR) and source-to-interferences ratio (SIR) respectively under strong reverberation for 470 ms.

引用

页码：347 / 353

页数：7

共 27 条

[1] IMAGE METHOD FOR EFFICIENTLY SIMULATING SMALL-ROOM ACOUSTICS
ALLEN, JB
BERKLEY, DA
[J]. JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1979, 65 (04) : 943 - 950
[2] Barfuss H, 2018, SIGNALS COMMUN TECHN, P237, DOI 10.1007/978-3-319-73031-8_10
[3] A Unified Probabilistic View on Spatially Informed Source Separation and Extraction Based on Independent Vector Analysis
Brendel, Andreas
Haubner, Thomas
Kellermann, Walter
[J]. IEEE TRANSACTIONS ON SIGNAL PROCESSING, 2020, 68 : 3545 - 3558
[4] AN ADAPTIVE GENERALIZED SIDELOBE CANCELER WITH DERIVATIVE CONSTRAINTS
BUCKLEY, KM
GRIFFITHS, LJ
[J]. IEEE TRANSACTIONS ON ANTENNAS AND PROPAGATION, 1986, 34 (03) : 311 - 319
[5] Maximum likelihood approach for blind audio source separation using time-frequency Gaussian source models
Févotte, C
Cardoso, JF
[J]. 2005 WORKSHOP ON APPLICATIONS OF SIGNAL PROCESSING TO AUDIO AND ACOUSTICS (WASPAA), 2005, : 78 - 81
[6] Signal enhancement using beamforming and nonstationarity with applications to speech
Gannot, S
Burshtein, D
Weinstein, E
[J]. IEEE TRANSACTIONS ON SIGNAL PROCESSING, 2001, 49 (08) : 1614 - 1626
[7] AN ALTERNATIVE APPROACH TO LINEARLY CONSTRAINED ADAPTIVE BEAMFORMING
GRIFFITHS, LJ
JIM, CW
[J]. IEEE TRANSACTIONS ON ANTENNAS AND PROPAGATION, 1982, 30 (01) : 27 - 34
[8] Hiroe A, 2006, LECT NOTES COMPUT SC, V3889, P601
[9] HyvEarinen A., 2001, INDEPENDENT COMPONEN
[10] Supervised Determined Source Separation with Multichannel Variational Autoencoder
Kameoka, Hirokazu
Li, Li
Inoue, Shota
Makino, Shoji
[J]. NEURAL COMPUTATION, 2019, 31 (09) : 1891 - 1914

← 1 2 3 →