Gender and Age Estimation Methods Based on Speech Using Deep Neural Networks

被引:31
作者
Kwasny, Damian [1 ]
Hemmerling, Daria [1 ]
机构
[1] AGH Univ Sci & Technol, Dept Measurement & Elect, PL-30059 Krakow, Poland
关键词
speech processing; neural networks; gender classification; age estimation; x-vector; RECOGNITION;
D O I
10.3390/s21144785
中图分类号
O65 [分析化学];
学科分类号
070302 ; 081704 ;
摘要
The speech signal contains a vast spectrum of information about the speaker such as speakers' gender, age, accent, or health state. In this paper, we explored different approaches to automatic speaker's gender classification and age estimation system using speech signals. We applied various Deep Neural Network-based embedder architectures such as x-vector and d-vector to age estimation and gender classification tasks. Furthermore, we have applied a transfer learning-based training scheme with pre-training the embedder network for a speaker recognition task using the Vox-Celeb1 dataset and then fine-tuning it for the joint age estimation and gender classification task. The best performing system achieves new state-of-the-art results on the age estimation task using popular TIMIT dataset with a mean absolute error (MAE) of 5.12 years for male and 5.29 years for female speakers and a root-mean square error (RMSE) of 7.24 and 8.12 years for male and female speakers, respectively, and an overall gender recognition accuracy of 99.60%.
引用
收藏
页数:18
相关论文
共 50 条
[31]   Automatic Speech Recognition with Deep Neural Networks for Impaired Speech [J].
Espana-Bonet, Cristina ;
Fonollosa, Jose A. R. .
ADVANCES IN SPEECH AND LANGUAGE TECHNOLOGIES FOR IBERIAN LANGUAGES, IBERSPEECH 2016, 2016, 10077 :97-107
[32]   Predicting speech intelligibility with deep neural networks [J].
Spille, Constantin ;
Ewert, Stephan D. ;
Kollmeier, Birger ;
Meyer, Bernd T. .
COMPUTER SPEECH AND LANGUAGE, 2018, 48 :51-66
[33]   DEEP NEURAL NETWORKS FOR ESTIMATION AND INFERENCE [J].
Farrell, Max H. ;
Liang, Tengyuan ;
Misra, Sanjog .
ECONOMETRICA, 2021, 89 (01) :181-213
[34]   Development of methods based on neural networks in the estimation of mineral resources [J].
Alberdi, Elisabete ;
Hernandez, Heber ;
Goti, Aitor .
DYNA, 2024, 99 (03) :303-310
[35]   Non-Intrusive Speech Quality Assessment Based on Deep Neural Networks for Speech Communication [J].
Liu, Miao ;
Wang, Jing ;
Wang, Fei ;
Xiang, Fei ;
Chen, Jingdong .
IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2025, 36 (01) :174-187
[36]   Learning Waveform-Based Acoustic Models Using Deep Variational Convolutional Neural Networks [J].
Oglic, Dino ;
Cvetkovic, Zoran ;
Sollich, Peter .
IEEE-ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2021, 29 :2850-2863
[37]   Human Age Estimation Using Deep Convolutional Neural Network based on Dental Images (Orthopantomogram) [J].
Sathyavathi, S. ;
Baskaran, K. R. .
IETE JOURNAL OF RESEARCH, 2024, 70 (02) :1585-1592
[38]   Methods for Pruning Deep Neural Networks [J].
Vadera, Sunil ;
Ameen, Salem .
IEEE ACCESS, 2022, 10 :63280-63300
[39]   Deep Convolutional Neural Networks for Large-scale Speech Tasks [J].
Sainath, Tara N. ;
Kingsbury, Brian ;
Saon, George ;
Soltau, Hagen ;
Mohamed, Abdel-rahman ;
Dahl, George ;
Ramabhadran, Bhuvana .
NEURAL NETWORKS, 2015, 64 :39-48
[40]   PHYSIOLOGICALLY-BASED SPEECH SYNTHESIS USING NEURAL NETWORKS [J].
HIRAYAMA, M ;
VATIKIOTISBATESON, E ;
KAWATO, M .
IEICE TRANSACTIONS ON FUNDAMENTALS OF ELECTRONICS COMMUNICATIONS AND COMPUTER SCIENCES, 1993, E76A (11) :1898-1910