AN ITERATIVE FRAMEWORK FOR SELF-SUPERVISED DEEP SPEAKER REPRESENTATION LEARNING

被引:26
作者
Cai, Danwei [1 ]
Wang, Weiqing [1 ]
Li, Ming [2 ]
机构
[1] Duke Univ, Dept Elect & Comp Engn, Durham, NC USA
[2] Duke Kunshan Univ, Data Sci Res Ctr, Kunshan, Peoples R China
来源
2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021) | 2021年
关键词
speaker recognition; speaker embedding; self-supervised learning; contrastive learning; clustering;
D O I
10.1109/ICASSP39728.2021.9414713
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
In this paper, we propose an iterative framework for self-supervised speaker representation learning based on a deep neural network (DNN). The framework starts with training a self-supervision speaker embedding network by maximizing agreement between different segments within an utterance via a contrastive loss. Taking advantage of DNN's ability to learn from data with label noise, we propose to cluster the speaker embedding obtained from the previous speaker network and use the subsequent class assignments as pseudo labels to train a new DNN. Moreover, we iteratively train the speaker network with pseudo labels generated from the previous step to bootstrap the discriminative power of a DNN. Speaker verification experiments are conducted on the VoxCeleb dataset. The results show that our proposed iterative self-supervised learning framework outperformed previous works using self-supervision. The speaker network after 5 iterations obtains a 61% performance gain over the speaker embedding model trained with contrastive loss.
引用
收藏
页码:6728 / 6732
页数:5
相关论文
共 50 条
[41]   Self-Supervised EEG Representation Learning for Robust Emotion Recognition [J].
Liu, Huan ;
Zhang, Yuzhe ;
Chen, Xuxu ;
Zhang, Dalin ;
Li, Rui ;
Qin, Tao .
ACM TRANSACTIONS ON SENSOR NETWORKS, 2024, 20 (05)
[42]   Collaboratively Self-Supervised Video Representation Learning for Action Recognition [J].
Zhang, Jie ;
Wan, Zhifan ;
Hu, Lanqing ;
Lin, Stephen ;
Wu, Shuzhe ;
Shan, Shiguang .
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, 2025, 20 :1895-1907
[43]   Self-Supervised Visual Representation Learning via Residual Momentum [J].
Pham, Trung Xuan ;
Niu, Axi ;
Zhang, Kang ;
Jin, Tee Joshua Tian ;
Hong, Ji Woo ;
Yoo, Chang D. .
IEEE ACCESS, 2023, 11 :116706-116720
[44]   Contrastive Self-Supervised Learning With Smoothed Representation for Remote Sensing [J].
Jung, Heechul ;
Oh, Yoonju ;
Jeong, Seongho ;
Lee, Chaehyeon ;
Jeon, Taegyun .
IEEE GEOSCIENCE AND REMOTE SENSING LETTERS, 2022, 19
[45]   Contrastive Self-supervised Representation Learning Using Synthetic Data [J].
Dong-Yu She ;
Kun Xu .
International Journal of Automation and Computing, 2021, 18 (04) :556-567
[46]   Contrastive Self-supervised Representation Learning Using Synthetic Data [J].
Dong-Yu She ;
Kun Xu .
International Journal of Automation and Computing, 2021, 18 :556-567
[47]   Mitigating background bias in self-supervised video representation learning [J].
Akar, Arif ;
Senturk, Ufuk Umut ;
Ikizler-Cinbis, Nazli .
SIGNAL IMAGE AND VIDEO PROCESSING, 2025, 19 (01)
[48]   A Self-Supervised Deep Learning Framework for Unsupervised Few-Shot Learning and Clustering [J].
Zhang, Hongjing ;
Zhan, Tianyang ;
Davidson, Ian .
PATTERN RECOGNITION LETTERS, 2021, 148 :75-81
[49]   SELF-SUPERVISED TEXT-INDEPENDENT SPEAKER VERIFICATION USING PROTOTYPICAL MOMENTUM CONTRASTIVE LEARNING [J].
Xia, Wei ;
Zhang, Chunlei ;
Weng, Chao ;
Yu, Meng ;
Yu, Dong .
2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021), 2021, :6723-6727
[50]   SELF-SUPERVISED LEARNING FOR AUDIO-VISUAL SPEAKER DIARIZATION [J].
Ding, Yifan ;
Xu, Yong ;
Zhang, Shi-Xiong ;
Cong, Yahuan ;
Wang, Liqiang .
2020 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, 2020, :4367-4371