A new hybrid approach for automatic speech signal segmentation using silence signal detection, energy convex hull, and spectral variation

被引：0

作者：

Zhao, Xufang ^{[1
]}

O'Shaughnessy, Douglas ^{[1
]}

机构：

[1] Univ Quebec, Inst Natl Rech Sci, Ste Foy, PQ G1V 2M3, Canada

来源：

2008 CANADIAN CONFERENCE ON ELECTRICAL AND COMPUTER ENGINEERING, VOLS 1-4 | 2008年

关键词：

acoustic signal analysis; acoustic signal processing; speech processing;

D O I：

暂无

中图分类号：

TP3 [计算技术、计算机技术];

学科分类号：

0812 ;

摘要：

This paper proposes a new approach for automatic syllable segmentation of Mandarin spontaneous speech. Automatic speech segmentation is important for continuous speech recognition because it reduces the search space effectively in automatic speech recognition. Moreover, the signal segmentation technique is useful in automatic speech marks and labels. However, for automatic speech recognition (ASR), it is difficult to segment the speech input reliably into useful sub-units because (1) syllable units can often be located roughly via intensity changes, but exact boundary positions are elusive in successive vowels, (2) energy changes in speech spectrum or amplitude help to estimate unit boundaries, but these cues are often unreliable due to co-articulation, and (3) finding boundaries for units bigger than phonemes combines the difficulties of detecting phoneme edges and of deciding which phonemes group to form the bigger units. In this paper, we present a hybrid segmentation method that utilizes silence detection, convex hull energy analysis, and spectral variation analysis. Furthermore, Hamming short-time sliding-windows were applied twice on audio signals to get more obvious convex hull valleys. Mandarin speech segmentation was used as a testing case, and the effectiveness of the proposed segmentation system was confirmed by the experimental results.

引用

页码：139 / 142

页数：4

共 9 条

[1] Hsieh CT, 1999, J INF SCI ENG, V15, P615
[2] Thai syllable segmentation for connected speech based on energy
Jittiwarangkul, N
Jitapunkul, S
Luksaneeyanawin, S
Ahkuputra, V
Wutiwiwatchai, C
[J]. APCCAS '98 - IEEE ASIA-PACIFIC CONFERENCE ON CIRCUITS AND SYSTEMS: MICROELECTRONICS AND INTEGRATING SYSTEMS, 1998, : 169 - 172
[3] *LING DAT CONS, 1998, 1997 MAND BROADC NEW
[4] MAKASHAY MJ, 2000, P INT C SPOK LANG PR, P431
[5] AUTOMATIC SEGMENTATION OF SPEECH INTO SYLLABIC UNITS
MERMELSTEIN, P
[J]. JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1975, 58 (04) : 880 - 883
[6] O'Shaughnessy D., 2000, SPEECH COMMUN
[7] Petrillo M., 2003, P 8 EUR C SPEECH COM
[8] Pfitzinger HR, 1996, ICSLP 96 - FOURTH INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE PROCESSING, PROCEEDINGS, VOLS 1-4, P1261, DOI 10.1109/ICSLP.1996.607838
[9] Sharma M, 1996, ICSLP 96 - FOURTH INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE PROCESSING, PROCEEDINGS, VOLS 1-4, P1237, DOI 10.1109/ICSLP.1996.607832

← 1 →