SIGNAL RECONSTRUCTION FROM MEL-SPECTROGRAM BASED ON BI-LEVEL CONSISTENCY OF FULL-BAND MAGNITUDE AND PHASE

被引:0
作者
Masuyama, Yoshiki [1 ]
Ueno, Natsuki [1 ]
Ono, Nobutaka [1 ]
机构
[1] Tokyo Metropolitan Univ, Tokyo, Japan
来源
2023 IEEE WORKSHOP ON APPLICATIONS OF SIGNAL PROCESSING TO AUDIO AND ACOUSTICS, WASPAA | 2023年
关键词
Phase reconstruction; waveform synthesis; mel-spectrogram; bi-level consistency; proximal splitting methods; ALTERNATING LINEARIZED MINIMIZATION; ALGORITHM; NONCONVEX;
D O I
10.1109/WASPAA58266.2023.10248111
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
We propose an optimization-based method for reconstructing a time-domain signal from a low-dimensional spectral representation such as a mel-spectrogram. Phase reconstruction has been studied to reconstruct a time-domain signal from the full-band short-time Fourier transform (STFT) magnitude. The Griffin-Lim algorithm (GLA) has been widely used because it relies only on the redundancy of STFT and is applicable to various audio signals. In this paper, we jointly reconstruct the full-band magnitude and phase by considering the bi-level relationships among the time-domain signal, its STFT coefficients, and its mel-spectrogram. The proposed method is formulated as a rigorous optimization problem and estimates the full-band magnitude based on the criterion used in GLA. Our experiments demonstrate the effectiveness of the proposed method on speech, music, and environmental signals.
引用
收藏
页数:5
相关论文
共 30 条
[11]   CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-spectrogram Conversion [J].
Kaneko, Takuhiro ;
Kameoka, Hirokazu ;
Tanaka, Kou ;
Hojo, Nobukatsu .
INTERSPEECH 2020, 2020, :2017-2021
[12]  
Kong Jungil, 2020, ADV NEUR IN, V33
[13]  
Kumar K, 2019, ADV NEUR IN, V32
[14]   Deep Griffin-Lim Iteration: Trainable Iterative Phase Reconstruction Using Neural Network [J].
Masuyama, Yoshiki ;
Yatabe, Kohei ;
Koizumi, Yuma ;
Oikawa, Yasuhiro ;
Harada, Noboru .
IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, 2021, 15 (01) :37-50
[15]  
Masuyama Y, 2019, INT CONF ACOUST SPEE, P61, DOI 10.1109/ICASSP.2019.8682744
[16]   Griffin-Lim Like Phase Recovery via Alternating Direction Method of Multipliers [J].
Masuyama, Yoshiki ;
Yatabe, Kohei ;
Oikawa, Yasuhiro .
IEEE SIGNAL PROCESSING LETTERS, 2019, 26 (01) :184-188
[17]  
McFee B., 2015, P SCIPY AUST TX US 6, P18, DOI 10.25080/majora-7b98-3ed-003
[18]  
Mowlaee P., 2016, Single Channel Phase-Aware Signal Processing in Speech Communication: Theory and Practice
[19]   Onoma-to-wave: Environmental Sound Synthesis from Onomatopoeic Words [J].
Okamoto, Yuki ;
Imoto, Keisuke ;
Takamichi, Shinnosuke ;
Yamanishi, Ryosuke ;
Fukumori, Takahiro ;
Yamashita, Yoichi .
APSIPA TRANSACTIONS ON SIGNAL AND INFORMATION PROCESSING, 2022, 11 (01)
[20]  
Parikh N., 2014, Found. Trends Optim, V1, P127, DOI [10.1561/2400000003, DOI 10.1561/2400000003]