The impact of speaking rate on acoustic-to-articulatory inversion

被引：16

作者：

Illa, Aravind ^{[1
]}

Ghosh, Prasanta Kumar ^{[1
]}

机构：

[1] Indian Inst Sci, Dept Elect Engn, Bangalore 560012, Karnataka, India

来源：

COMPUTER SPEECH AND LANGUAGE | 2020年 / 59卷

关键词：

Acoustic-to-articulatory inversion; Speaking rate; Electromagnetic articulograph; NEURAL-NETWORK MODEL; SPEECH; VELOCITY; MOVEMENT; JAW; LIP; COARTICULATION; CONSTRAINTS; VARIABILITY; ACQUISITION;

D O I：

10.1016/j.csl.2019.05.004

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Acoustic characteristics and articulatory movements are known to vary with speaking rates. This study investigates the role of speaking rate on acoustic-to-articulatory inversion (AAI) performance using deep neural networks (DNNs). Since fast speaking rate causes fast articulatory motion as well as changes in spectro-temporal characteristics of the speech signal, the articulatory-acoustic map in a fast speaking rate could be different from that in a slow speaking rate. We examine how these differences alter the accuracy with which different articulatory positions could be recovered from the acoustics. AAI experiments are performed in both matched and mismatched train-test conditions using data of five subjects, in three different rates - normal, fast and slow (fast and slow rates are at least 1.3 times faster and slower than the normal rate). Experiments in matched cases reveal that, the errors in estimating vertical motion of sensors on the tongue articulators from acoustics with fast speaking rate, is significantly higher than those with slow speaking rate. Experiments in mis-matched conditions reveal that there is consistent drop in AAI performance compared to the matched condition. Further experiments performed by training AAI with acoustic-articulatory data pooled from different speaking rates reveal that a single DNN based AAI model is capable of learning multiple rate-specific mapping. (C) 2019 Elsevier Ltd. All rights reserved.

引用

页码：75 / 90

页数：16

共 69 条

[1] SPEAKING RATE AND SPEECH MOVEMENT VELOCITY PROFILES [J].

ADAMS, SG ;

WEISMER, G ;

KENT, RD .

JOURNAL OF SPEECH AND HEARING RESEARCH, 1993, 36 (01) :41-54

[2] Improved subject-independent acoustic-to-articulatory inversion [J].

Afshan, Amber ;

Ghosh, Prasanta Kumar .

SPEECH COMMUNICATION, 2015, 66 :1-16

[3]

Agwuele A., 2009, PHONETICA, V65, P194

[4]

[Anonymous], Robust speech recognition using articulatory information

[5]

[Anonymous], 2014, ABS14126980 CORR

[6]

[Anonymous], 2009, NEURAL NETWORKS LEAR

[7]

[Anonymous], P NIPS 2011 WORKSH D

[8]

Berry J., 2011, SIG 5 PERSPECTIVES S, V21, P15, DOI [DOI 10.1044/SSOD21.1.15, 10.1044/ssod21.1.15]

[9]

Chollet F., 2015, KERAS LIB

[10]

Cover TM, 2012, Elements of information theory

← 1 2 3 4 5 6 7 →