A novel approach to HMM-based speech recognition systems using particle swarm optimization

被引：28

作者：

Najkar, Negin ^{[1
]}

Razzazi, Farbod ^{[1
]}

Sameti, Hossein ^{[2
]}

机构：

[1] Islamic Azad Univ, Sci & Res Branch, Dept Elect Engn, Fac Engn, Tehran, Iran

[2] Sharif Univ Technol, Dept Comp Engn, Tehran, Iran

来源：

MATHEMATICAL AND COMPUTER MODELLING | 2010年 / 52卷 / 11-12期

关键词：

Hidden Markov model (HMM); Particle swarm optimization (PSO); HMM-based speech recognition; Viterbi algorithm;

D O I：

10.1016/j.mcm.2010.03.041

中图分类号：

TP39 [计算机的应用];

学科分类号：

081203 ; 0835 ;

摘要：

The main core of HMM-based speech recognition systems is Viterbial gorithm. Viterbi algorithm uses dynamic programming to find out the best alignment between the input speech and a given speech model. In this paper, dynamic programming is replaced by a search method which is based on particle swarm optimization algorithm. The major idea is focused on generating an initial population of segmentation vectors in the solution search space and improving the location of segments by an updating algorithm. Several methods are introduced and evaluated for the representation of particles and their corresponding movement structures. In addition, two segmentation strategies are explored. The first method is the standard segmentation which tries to maximize the likelihood function for each competing acoustic model separately. In the next method, a global segmentation tied between several models and the system tries to optimize the likelihood using a common tied segmentation. The results show that the effect of these factors is noticeable in finding the global optimum while maintaining the system accuracy. The idea was tested on an isolated word recognition and phone classification tasks and shows its significant performance in both accuracy and computational complexity aspects. (C) 2010 Elsevier Ltd. All rights reserved.

引用

页码：1910 / 1920

页数：11

共 19 条

[1] A genetic algorithm-aided Hidden Markov Model topology estimation for phoneme recognition of Thai continuous speech [J].

Bhuriyakorn, Pattana ;

Punyabukkana, Proadpran ;

Suchato, Atiwong .

PROCEEDINGS OF NINTH ACIS INTERNATIONAL CONFERENCE ON SOFTWARE ENGINEERING, ARTIFICIAL INTELLIGENCE, NETWORKING AND PARALLEL/DISTRIBUTED COMPUTING, 2008, :475-480

[2]

Chau CW, 1997, INT CONF ACOUST SPEE, P1727, DOI 10.1109/ICASSP.1997.598857

[3] Closure duration analysis of incomplete stop consonants due to stop-stop interaction [J].

Ghosh, Prasanta Kumar ;

Narayanan, Shrikanth S. .

JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 2009, 126 (01) :EL1-EL7

[4]

Hong QY, 2003, PROCEEDINGS OF 2003 INTERNATIONAL CONFERENCE ON NEURAL NETWORKS & SIGNAL PROCESSING, PROCEEDINGS, VOLS 1 AND 2, P465

[5]

Kennedy J, 1995, 1995 IEEE INTERNATIONAL CONFERENCE ON NEURAL NETWORKS PROCEEDINGS, VOLS 1-6, P1942, DOI 10.1109/icnn.1995.488968

[6] A genetic classification error method for speech recognition [J].

Kwong, S ;

He, QH ;

Ku, KW ;

Chan, TM ;

Man, KF ;

Tang, KS .

SIGNAL PROCESSING, 2002, 82 (05) :737-748

[7] Optimisation of HMM topology and its model parameters by genetic algorithms [J].

Kwong, S ;

Chau, CW ;

Man, KF ;

Tang, KS .

PATTERN RECOGNITION, 2001, 34 (02) :509-522

[8] Genetic algorithm for optimizing the nonlinear time alignment of automatic speech recognition systems [J].

Kwong, S ;

Chau, CW ;

Halang, WA .

IEEE TRANSACTIONS ON INDUSTRIAL ELECTRONICS, 1996, 43 (05) :559-566

[9]

Mizuta S., 1992, Journal of the Acoustical Society of Japan (E), V13, P389, DOI 10.1250/ast.13.389

[10]

NAJKAR N, 2009, P IEEE INT C BIOINSP, P1

← 1 2 →