A Context Encoder For Audio Inpainting

被引：48

作者：

Marafioti, Andres ^{[1
]}

Perraudin, Nathanael ^{[2
]}

Holighaus, Nicki ^{[1
]}

Majdak, Piotr ^{[1
]}

机构：

[1] Austrian Acad Sci, Acoust Res Inst, A-1040 Vienna, Austria

[2] Swiss Fed Inst Technol, Swiss Data Sci Ctr, CH-8006 Zurich, Switzerland

来源：

IEEE-ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING | 2019年 / 27卷 / 12期

基金：

奥地利科学基金会;

关键词：

Instruments; Time-domain analysis; Music; Image reconstruction; Reliability; Prediction algorithms; Psychoacoustic models; machine learning; frequency-domain analysis; signal processing algorithms; RECONSTRUCTION; EXTRAPOLATION; INTERPOLATION; RESTORATION; SEGMENTS; PHASE;

D O I：

10.1109/TASLP.2019.2947232

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

In this article, we study the ability of deep neural networks (DNNs) to restore missing audio content based on its context, i.e., inpaint audio gaps. We focus on a condition which has not received much attention yet: gaps in the range of tens of milliseconds. We propose a DNN structure that is provided with the signal surrounding the gap in the form of time-frequency (TF) coefficients. Two DNNs with either complex-valued TF coefficient output or magnitude TF coefficient output were studied by separately training them on inpainting two types of audio signals (music and musical instruments) having 64-ms long gaps. The magnitude DNN outperformed the complex-valued DNN in terms of signal-to-noise ratios and objective difference grades. Although, for instruments, a reference inpainting obtained through linear predictive coding performed better in both metrics, it performed worse than the magnitude DNN for music. This demonstrates the potential of the magnitude DNN, in particular for inpainting signals that are more complex than single instrument sounds.

引用

页码：2362 / 2372

页数：11

共 69 条

[1] Abadi M., 2015, TENSORFLOW LARGESCAL
[2] Audio Inpainting
Adler, Amir
Emiya, Valentin
Jafari, Maria G.
Elad, Michael
Gribonval, Remi
Plumbley, Mark D.
[J]. IEEE TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2012, 20 (03): : 922 - 932
[3] Adler A, 2011, INT CONF ACOUST SPEE, P329
[4] [Anonymous], 2015, ARXIV PREPRINT ARXIV
[5] [Anonymous], P 7 INT C LEARN REPR
[6] [Anonymous], CORR
[7] [Anonymous], 1992, NIPS 91 P 4 INT C NE
[8] [Anonymous], 2016, DEEP LEARNING
[9] [Anonymous], 2001, FDN TIME FREQUENCY A
[10] [Anonymous], 2017, PROC INT C LEARN REP

← 1 2 3 4 5 6 7 →