Conversational speech synthesis and the need for some laughter

被引:22
作者
Campbell, Nick [1 ]
机构
[1] Natl Inst Informat & Commun Technol, Keihanna Sci City, Kyoto 6190288, Japan
[2] ATR Spoken Language Commun Lab, Speech & Acoust Proc Dept, Keihanna Sci City, Kyoto 6190288, Japan
来源
IEEE TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING | 2006年 / 14卷 / 04期
关键词
affect; conversation; emotion; expression; laughter; nonverbal; social interaction; speech synthesis;
D O I
10.1109/TASL.2006.876131
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
This paper reports progress in the synthesis of conversational speech, from the viewpoint of work carried out on the analysis of a very large corpus of expressive speech in normal everyday situations. With recent developments in concatenative techniques, speech synthesis has overcome the barrier of realistically portraying extra-linguistic information by using the actual voice of a recognizable person as a source for units, combined with minimal use of signal processing. However, the technology still faces the problem of expressing paralinguistic information, i.e., the variety in the types of speech and laughter that a person might use in everyday social interactions. Paralinguistic modification of an utterance portrays the speaker's affective states and shows his or her relationships with the speaker through variations in the manner of speaking, by means of prosody and voice quality. These inflections are carried on the propositional content of an utterance, and can perhaps be modeled by rule, but they are also expresssed through nonverbal utterances, the complexity of which may be beyond the capabilities of many current synthesis methods. We suggest that this problem may be solved by the use of phrase-sized utterance units taken intact from a large corpus.
引用
收藏
页码:1171 / 1178
页数:8
相关论文
共 39 条