Self-supervised speaker embeddings

被引:28
|
作者
Stafylakis, Themos [1 ]
Rohdin, Johan [2 ]
Plchot, Oldrich [2 ]
Mizera, Petr [1 ]
Burget, Lukas [2 ]
机构
[1] Omilia Conversat Intelligence, Athens, Greece
[2] Brno Univ Technol, Fac Informat Technol, Ctr Excellence IT4I, Brno, Czech Republic
来源
INTERSPEECH 2019 | 2019年
基金
美国国家科学基金会;
关键词
speaker recognition; self-supervised learning; deep learning; VERIFICATION; RECOGNITION;
D O I
10.21437/Interspeech.2019-2842
中图分类号
R36 [病理学]; R76 [耳鼻咽喉科学];
学科分类号
100104 ; 100213 ;
摘要
Contrary to i-vectors, speaker embeddings such as x-vectors are incapable of leveraging unlabelled utterances, due to the classification loss over training speakers. In this paper, we explore an alternative training strategy to enable the use of unlabelled utterances in training. We propose to train speaker embedding extractors via reconstructing the frames of a target speech segment, given the inferred embedding of another speech segment of the same utterance. We do this by attaching to the standard speaker embedding extractor a decoder network, which we feed not merely with the speaker embedding, but also with the estimated phone sequence of the target frame sequence. The reconstruction loss can be used either as a single objective, or be combined with the standard speaker classification loss. In the latter case, it acts as a regularizer, encouraging generalizability to speakers unseen during training. In all cases, the proposed architectures are trained from scratch and in an end-to-end fashion. We demonstrate the benefits from the proposed approach on the VoxCeleb and Speakers in the Wild Databases, and we report notable improvements over the baseline.
引用
收藏
页码:2863 / 2867
页数:5
相关论文
共 50 条
  • [1] Stuttering detection using speaker representations and self-supervised contextual embeddings
    Sheikh S.A.
    Sahidullah M.
    Hirsch F.
    Ouni S.
    International Journal of Speech Technology, 2023, 26 (02) : 521 - 530
  • [2] Self-supervised Speaker Diarization
    Dissen, Yehoshua
    Kreuk, Felix
    Keshet, Joseph
    INTERSPEECH 2022, 2022, : 4013 - 4017
  • [3] SELF-SUPERVISED SPEAKER VERIFICATION WITH SIMPLE SIAMESE NETWORK AND SELF-SUPERVISED REGULARIZATION
    Sang, Mufan
    Li, Haoqi
    Liu, Fang
    Arnold, Andrew O.
    Wan, Li
    2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2022, : 6127 - 6131
  • [4] Boosting Self-Supervised Embeddings for Speech Enhancement
    Hung, Kuo-Hsuan
    Fu, Szu-Wei
    Tseng, Huan-Hsin
    Chiang, Hsin-Tien
    Tsao, Yu
    Lin, Chii-Wann
    INTERSPEECH 2022, 2022, : 186 - 190
  • [5] Self-Supervised Learning for Online Speaker Diarization
    Chien, Jen-Tzung
    Luo, Sixun
    2021 ASIA-PACIFIC SIGNAL AND INFORMATION PROCESSING ASSOCIATION ANNUAL SUMMIT AND CONFERENCE (APSIPA ASC), 2021, : 2036 - 2042
  • [6] Curriculum learning for self-supervised speaker verification
    Heo, Hee-Soo
    Jung, Jee-weon
    Kang, Jingu
    Kwon, Youngki
    Kim, You Jin
    Lee, Bong-Jin
    Chung, Joon Son
    INTERSPEECH 2023, 2023, : 4693 - 4697
  • [7] Prototype Division for Self-Supervised Speaker Verification
    Zhao, Zhenduo
    Li, Zhuo
    Zhang, Xueshuai
    Wang, Wenchao
    Zhang, Pengyuan
    IEEE SIGNAL PROCESSING LETTERS, 2024, 31 : 880 - 884
  • [8] ROBUST SPEAKER VERIFICATION WITH JOINT SELF-SUPERVISED AND SUPERVISED LEARNING
    Wang, Kai
    Zhang, Xiaolei
    Zhang, Miao
    Li, Yuguang
    Lee, Jaeyun
    Cho, Kiho
    Park, Sung-UN
    2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2022, : 7637 - 7641
  • [9] Self-supervised Learning of Contextualized Local Visual Embeddings
    Silva, Thalles
    Pedrini, Helio
    Rivera, Adin Ramirez
    2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION WORKSHOPS, ICCVW, 2023, : 177 - 186
  • [10] Self-supervised learning of class embeddings from video
    Wiles, Olivia
    Koepke, A. Sophia
    Zisserman, Andrew
    2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION WORKSHOPS (ICCVW), 2019, : 3019 - 3027