Time-Frequency Localization Using Deep Convolutional Maxout Neural Network in Persian Speech Recognition

被引:2
|
作者
Dehghani, Arash [1 ]
Seyyedsalehi, Seyyed Ali [1 ]
机构
[1] Amirkabir Univ Technol, Fac Biomed Engn, Hafez Ave, Tehran, Iran
关键词
Time-Frequency Localization; Deep Neural Networks; Convolutional Neural Networks; Speech Recognition; Maxout; Dropout; SPECTROTEMPORAL RECEPTIVE-FIELDS; TASK-RELATED PLASTICITY; FILTER BANK FEATURES; OPTIMIZATION; NEURONS; LAYER; NETS;
D O I
10.1007/s11063-022-11006-1
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
In this paper, a CNN-based structure for the time-frequency localization of information is proposed for Persian speech recognition. Research has shown that the receptive fields' spectrotemporal plasticity of some neurons in mammals' primary auditory cortex and midbrain makes localization facilities improve recognition performance. Over the past few years, much work has been done to localize time-frequency information in ASR systems, using the spatial or temporal immutability properties of methods such as HMMs, TDNNs, CNNs, and LSTM-RNNs. However, most of these models have large parameter volumes and are challenging to train. For this purpose, we have presented a structure called Time-Frequency Convolutional Maxout Neural Network (TFCMNN) in which parallel time-domain and frequency-domain 1D-CMNNs are applied simultaneously and independently to the spectrogram, and then their outputs are concatenated and applied jointly to a fully connected Maxout network for classification. To improve the performance of this structure, we have used newly developed methods and models such as Dropout, maxout, and weight normalization. Two sets of experiments were designed and implemented on the FARSDAT dataset to evaluate the performance of this model compared to conventional 1D-CMNN models. According to the experimental results, the average recognition score of TFCMNN models is about 1.6% higher than the average of conventional 1D-CMNN models. In addition, the average training time of the TFCMNN models is about 17 h lower than the average training time of traditional models. Therefore, as proven in other sources, time-frequency localization in ASR systems increases system accuracy and speeds up the training process.
引用
收藏
页码:3205 / 3224
页数:20
相关论文
共 50 条
  • [1] Time-Frequency Localization Using Deep Convolutional Maxout Neural Network in Persian Speech Recognition
    Arash Dehghani
    Seyyed Ali Seyyedsalehi
    Neural Processing Letters, 2023, 55 : 3205 - 3224
  • [2] Performance Evaluation of Deep Convolutional Maxout Neural Network in Speech Recognition
    Dehghani, Arash
    Seyyedsalehi, Seyyed Ali
    2018 25TH IRANIAN CONFERENCE ON BIOMEDICAL ENGINEERING AND 2018 3RD INTERNATIONAL IRANIAN CONFERENCE ON BIOMEDICAL ENGINEERING (ICBME), 2018, : 240 - 245
  • [3] Maxout neurons for deep convolutional and LSTM neural networks in speech recognition
    Cai, Meng
    Liu, Jia
    SPEECH COMMUNICATION, 2016, 77 : 53 - 64
  • [4] Hand Gesture Recognition with Ensemble Time-Frequency Signatures Using Enhanced Deep Convolutional Neural Network
    Feng, Xiang
    Song, Qun
    Guo, Qingfang
    Liu, Duo
    Zhao, Zhanfeng
    Zhao, Yinan
    2019 ASIA-PACIFIC SIGNAL AND INFORMATION PROCESSING ASSOCIATION ANNUAL SUMMIT AND CONFERENCE (APSIPA ASC), 2019, : 1602 - 1605
  • [5] DEEP MAXOUT NEURAL NETWORKS FOR SPEECH RECOGNITION
    Cai, Meng
    Shi, Yongzhe
    Liu, Jia
    2013 IEEE WORKSHOP ON AUTOMATIC SPEECH RECOGNITION AND UNDERSTANDING (ASRU), 2013, : 291 - 296
  • [6] TIME-FREQUENCY CONVOLUTIONAL NETWORKS FOR ROBUST SPEECH RECOGNITION
    Mitra, Vikramjit
    Franco, Horacio
    2015 IEEE WORKSHOP ON AUTOMATIC SPEECH RECOGNITION AND UNDERSTANDING (ASRU), 2015, : 317 - 323
  • [7] Speech Emotion Recognition via an Attentive Time-Frequency Neural Network
    Lu, Cheng
    Zheng, Wenming
    Lian, Hailun
    Zong, Yuan
    Tang, Chuangao
    Li, Sunan
    Zhao, Yan
    IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, 2023, 10 (06) : 3159 - 3168
  • [8] Frequency hopping modulation recognition of convolutional neural network based on time-frequency characteristics
    Li H.-G.
    Guo Y.
    Sui P.
    Qi Z.-S.
    Zhejiang Daxue Xuebao (Gongxue Ban)/Journal of Zhejiang University (Engineering Science), 2020, 54 (10): : 1945 - 1954
  • [9] Time-Frequency Representation and Convolutional Neural Network-Based Emotion Recognition
    Khare, Smith K.
    Bajaj, Varun
    IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2021, 32 (07) : 2901 - 2909
  • [10] Deep Convolutional Neural Network for Arabic Speech Recognition
    Amari, Rafik
    Noubigh, Zouhaira
    Zrigui, Salah
    Berchech, Dhaou
    Nicolas, Henri
    Zrigui, Mounir
    COMPUTATIONAL COLLECTIVE INTELLIGENCE, ICCCI 2022, 2022, 13501 : 120 - 134