ENHANCING INTO THE CODEC: NOISE ROBUST SPEECH CODING WITH VECTOR-QUANTIZED AUTOENCODERS

被引：12

作者：

Casebeer, Jonah ^{[1
]}

Vale, Vinjai ^{[2
]}

Isik, Umut ^{[3
]}

Valin, Jean-Marc ^{[3
]}

Giri, Ritwik ^{[3
]}

Krishnaswamy, Arvindh ^{[3
]}

机构：

[1] Univ Illinois, Champaign, IL 61820 USA

[2] Stanford Univ, Stanford, CA 94305 USA

[3] Amazon Web Serv, Seattle, WA USA

来源：

2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021) | 2021年

关键词：

speech enhancement; speech coding; audio compression; ENHANCEMENT;

D O I：

10.1109/ICASSP39728.2021.9414605

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

Audio codecs based on discretized neural autoencoders have recently been developed and shown to provide significantly higher compression levels for comparable quality speech output. However, these models are tightly coupled with speech content, and produce unintended outputs in noisy conditions. Based on VQ-VAE autoencoders with WaveRNN decoders, we develop compressor-enhancer encoders and accompanying decoders, and show that they operate well in noisy conditions. We also observe that a compressor-enhancer model performs better on clean speech inputs than a compressor model trained only on clean speech.

引用

页码：711 / 715

页数：5

共 21 条

[1]

[Anonymous], 2017, NOISY SPEECH DATABAS

[2] Unsupervised Acoustic Unit Representation Learning for Voice Conversion using WaveNet Auto-encoders [J].

Chen, Mingjie ;

Hain, Thomas .

INTERSPEECH 2020, 2020, :4866-4870

[3] Unsupervised Speech Representation Learning Using WaveNet Autoencoders [J].

Chorowski, Jan ;

Weiss, Ron J. ;

Bengio, Samy ;

van den Oord, Aaron .

IEEE-ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2019, 27 (12) :2041-2053

[4]

Gârbacea C, 2019, INT CONF ACOUST SPEE, P735, DOI [10.1109/icassp.2019.8683277, 10.1109/ICASSP.2019.8683277]

[5]

Gemmeke JF, 2017, INT CONF ACOUST SPEE, P776, DOI 10.1109/ICASSP.2017.7952261

[6] PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased Loss [J].

Isik, Umut ;

Giri, Ritwik ;

Phansalkar, Neerad ;

Valin, Jean-Marc ;

Helwani, Karim ;

Krishnaswamy, Arvindh .

INTERSPEECH 2020, 2020, :2487-2491

[7]

ITU-T, 2018, SUBJ EV SPEECH QUAL

[8]

Kleijn WB, 2018, 2018 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), P676, DOI 10.1109/ICASSP.2018.8462529

[9]

Lim FSC, 2020, INT CONF ACOUST SPEE, P6769, DOI 10.1109/ICASSP40776.2020.9053657

[10]

Lorenzo-Trueba J., 2018, ARXIV181106292

← 1 2 3 →