Speech recognition engineering issues in speech to speech translation system design for low resource languages and domains

被引：0

作者：

Narayanan, Shrikanth ^{[1
]}

Georgiou, Panayiotis G. ^{[1
]}

Sethy, Abhinav ^{[1
]}

Wang, Dagen ^{[1
]}

Bulut, Murtaza ^{[1
]}

Sundaram, Shiva ^{[1
]}

Ettelaie, Emil ^{[1
]}

Ananthakrishnan, Sankaranarayanan ^{[1
]}

Franco, Horacio ^{[1
]}

Precoda, Kristin ^{[1
]}

Vergyri, Dimitra ^{[1
]}

Zheng, Jing ^{[1
]}

Wang, Wen ^{[1
]}

Gadde, Ramana Rao ^{[1
]}

Graciarena, Martin ^{[1
]}

Abrash, Victor ^{[1
]}

Frandsen, Michael ^{[1
]}

Richey, Colleen ^{[1
]}

机构：

[1] Univ So Calif, Viterbi Sch Engn, Los Angeles, CA 90089 USA

来源：

2006 IEEE International Conference on Acoustics, Speech and Signal Processing, Vols 1-13 | 2006年

关键词：

D O I：

暂无

中图分类号：

O42 [声学];

学科分类号：

070206 ; 082403 ;

摘要：

Engineering automatic speech recognition (ASR) for speech to speech (S2S) translation systems, especially targeting languages and domains that do not have readily available spoken language resources, is immensely challenging due to a number of reasons. In addition to contending with the conventional data-hungry speech acoustic and language modeling needs, these designs have to accommodate varying requirements imposed by the domain needs and characteristics, target device and usage modality (such as phrase-based, or spontaneous free form interactions, with or without visual feedback) and huge spoken language variability arising due to socio-linguistic and cultural differences of the users. This paper, using case studies of creating speech translation systems between English and languages such as Pashto and Farsi, describes some of the practical issues and the solutions that were developed for multilingual ASR development. These include novel acoustic and language modeling strategies such as language adaptive recognition, active-learning based language modeling, class-based language models that can better exploit resource poor language data, efficient search strategies, including N-best and confidence generation to aid multiple hypotheses translation, use of dialog information and clever interface choices to facilitate ASR, and audio interface design for meeting both usability and robustness requirements.

引用

页码：6067 / 6070

页数：4

共 15 条

[1]

Belvin R., 2004, P LREC LISB PORT

[2]

BELVIN R, 2005, P ASS COMP LING U MI

[3]

FRANCO H, 2003, P EUROSPEECH

[4]

FRANCO H, 2002, P HUM LANG TECHN C

[5]

GANJAVI S, 2004, P WORKSH COMP APPR A

[6]

GEORGIOU PG, 2004, P INT C SPOK LANG PR

[7]

GRACIARENA M, 2005, P EUROSPEECH

[8]

KATHOL A, 2005, EUROSPEECH

[9]

MOHRI M, 2000, ISCA ITRW ASR

[10]

NARAYANAN S, 2003, IEEE AUT SPEECH REC

← 1 2 →