SIGNAL RECONSTRUCTION FROM MEL-SPECTROGRAM BASED ON BI-LEVEL CONSISTENCY OF FULL-BAND MAGNITUDE AND PHASE

被引：0

作者：

Masuyama, Yoshiki ^{[1
]}

Ueno, Natsuki ^{[1
]}

Ono, Nobutaka ^{[1
]}

机构：

[1] Tokyo Metropolitan Univ, Tokyo, Japan

来源：

2023 IEEE WORKSHOP ON APPLICATIONS OF SIGNAL PROCESSING TO AUDIO AND ACOUSTICS, WASPAA | 2023年

关键词：

Phase reconstruction; waveform synthesis; mel-spectrogram; bi-level consistency; proximal splitting methods; ALTERNATING LINEARIZED MINIMIZATION; ALGORITHM; NONCONVEX;

D O I：

10.1109/WASPAA58266.2023.10248111

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

We propose an optimization-based method for reconstructing a time-domain signal from a low-dimensional spectral representation such as a mel-spectrogram. Phase reconstruction has been studied to reconstruct a time-domain signal from the full-band short-time Fourier transform (STFT) magnitude. The Griffin-Lim algorithm (GLA) has been widely used because it relies only on the redundancy of STFT and is applicable to various audio signals. In this paper, we jointly reconstruct the full-band magnitude and phase by considering the bi-level relationships among the time-domain signal, its STFT coefficients, and its mel-spectrogram. The proposed method is formulated as a rigorous optimization problem and estimates the full-band magnitude based on the criterion used in GLA. Our experiments demonstrate the effectiveness of the proposed method on speech, music, and environmental signals.

引用

页数：5

共 30 条

[1]

[Anonymous], 2007, ITU-T Recommendation P.805

[2]

Arias-Castro E, 2017, J MACH LEARN RES, V18, P1

[3] Proximal alternating linearized minimization for nonconvex and nonsmooth problems [J].