LANGUAGE DIARIZATION FOR SEMI-SUPERVISED BILINGUAL ACOUSTIC MODEL TRAINING

被引：0

作者：

Yilmaz, Emre ^{[1
,2
]}

McLaren, Mitchell ^{[2
]}

van den Heuvel, Henk ^{[1
]}

van Leeuwen, David A. ^{[1
]}

机构：

[1] Radboud Univ Nijmegen, CLS CLST, Nijmegen, Netherlands

[2] SRI Int, Speech Technol & Res Lab, 333 Ravenswood Ave, Menlo Pk, CA 94025 USA

来源：

2017 IEEE AUTOMATIC SPEECH RECOGNITION AND UNDERSTANDING WORKSHOP (ASRU) | 2017年

关键词：

Language diarization; code-switching; bilingual acoustic modeling; semi-supervised training; Frisian language; TRANSCRIPTION;

D O I：

暂无

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

In this paper, we investigate several automatic transcription schemes for using raw bilingual broadcast news data in semi-supervised bilingual acoustic model training. Specifically, we compare the transcription quality provided by a bilingual ASR system with another system performing language diarization at the front-end followed by two monolingual ASR systems chosen based on the assigned language label. Our research focuses on the Frisian-Dutch code-switching (CS) speech that is extracted from the archives of a local radio broadcaster. Using 11 hours of manually transcribed Frisian speech as a reference, we aim to increase the amount of available training data by using these automatic transcription techniques. By merging the manually and automatically transcribed data, we learn bilingual acoustic models and run ASR experiments on the development and test data of the FAME! speech corpus to quantify the quality of the automatic transcriptions. Using these acoustic models, we present speech recognition and CS detection accuracies. The results demonstrate that applying language diarization to the raw speech data to enable using the monolingual resources improves the automatic transcription quality compared to a baseline system using a bilingual ASR system.

引用

页码：91 / 96

页数：6

共 42 条

[1] Adel H., 2014, P INT WORKSH SPOK LA, P32
[2] Adel H, 2013, INT CONF ACOUST SPEE, P8411, DOI 10.1109/ICASSP.2013.6639306
[3] [Anonymous], 2003, PATT REC ASS S AFR
[4] [Anonymous], 2014, PROC 4 WORKSHOP SPOK
[5] [Anonymous], C ACOUSTICS SPEECH S, DOI DOI 10.1109/ICASSP.1984.1172555
[6] [Anonymous], 2006, Proc. Odyssey: Speaker and Language Recognition Workshop
[7] [Anonymous], 2011, IEEE 2011 WORKSHOP
[8] [Anonymous], 2014, PROC INTERSPEECH 201
[9] [Anonymous], 2001, P EUR
[10] [Anonymous], 2015, NAT LANG ENG

← 1 2 3 4 5 →