An investigation into the correlation and prediction of acoustic speech features from MFCC vectors

被引:0
作者
Darch, Jonathan [1 ]
Milner, Ben [1 ]
Almajai, Ibrahim [1 ]
Vaseghi, Saeed [2 ]
机构
[1] Univ East Anglia, Sch Comp Sci, Norwich NR4 7TJ, Norfolk, England
[2] Brunel Univ, Dept Elect & Comp Engn, Uxbridge UB8 3PH, Middx, England
来源
2007 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, VOL IV, PTS 1-3 | 2007年
基金
英国工程与自然科学研究理事会;
关键词
formants; fundamental frequency; voicing; GMM; HMM;
D O I
暂无
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
This work develops a statistical framework to predict acoustic features (fundamental frequency, formant frequencies and voicing) from MFCC vectors. An analysis of correlation between acoustic features and MFCCs is made both globally across all speech and within phoneme classes, and also from speaker-independent and speaker-dependent speech. This leads to the development of both a global prediction method, using a Gaussian mixture model (GMM) to model the joint density of acoustic features and MFCCs, and a phoneme-specific prediction method using a combined hidden Markov model (HMM)-GMM. Prediction accuracy measurements show the phoneme-dependent HMM-GMM system to be more accurate which agrees with the correlation analysis. Results a so show prediction to be more accurate from speaker-dependent speech which also agrees with the correlation analysis.
引用
收藏
页码:465 / +
页数:2
相关论文
共 6 条
[1]  
Chatterjee S., 2006, Regression Analysis by Example, V4th, P317
[2]  
Darch J, 2006, INTERSPEECH 2006 AND 9TH INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE PROCESSING, VOLS 1-5, P1005
[3]  
HIRAHARA T, 1988, 2 JOINT M ASA ASJ
[4]   Predicting fundamental frequency from mel-frequency cepstral coefficients to enable speech reconstruction [J].
Shao, X ;
Milner, B .
JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 2005, 118 (02) :1134-1143
[5]  
SORIN A, 2003, 202212 ES
[6]  
YAN Q, 2004, ICSLP, P2409