Transformer-Based Joint Learning Approach for Text Normalization in Vietnamese Automatic Speech Recognition Systems

被引：0

作者：

Viet The Bui ^{[1
]}

Tho Chi Luong ^{[2
]}

Oanh Thi Tran ^{[3
]}

机构：

[1] Singapore Management Univ, Sch Comp & Informat Syst, Singapore, Singapore

[2] FPT Univ, FPT Technol Res Inst, Hanoi, Vietnam

[3] Vietnam Natl Univ Hanoi, Int Sch, Hanoi, Vietnam

来源：

CYBERNETICS AND SYSTEMS | 2024年 / 55卷 / 07期

关键词：

ASR; named entity recognition; post-processing; punctuator; text normalization; transformer-based joint learning models;

D O I：

10.1080/01969722.2022.2145654

中图分类号：

TP3 [计算技术、计算机技术];

学科分类号：

0812 ;

摘要：

In this article, we investigate the task of normalizing transcribed texts in Vietnamese Automatic Speech Recognition (ASR) systems in order to improve user readability and the performance of downstream tasks. This task usually consists of two main sub-tasks: predicting and inserting punctuation (i.e., period, comma); and detecting and standardizing named entities (i.e., numbers, person names) from spoken forms to their appropriate written forms. To achieve these goals, we introduce a complete corpus including of 87,700 sentences and investigate conditional joint learning approaches which globally optimize two sub-tasks simultaneously. The experimental results are quite promising. Overall, the proposed architecture outperformed the conventional architecture which trains individual models on the two sub-tasks separately. The joint models are furthered improved when integrated with the surrounding contexts (SCs). Specifically, we obtained 81.13% for the first sub-task and 94.41% for the second sub-task in the F1 scores using the best model.

引用

页码：1614 / 1630

页数：17