In the present study, we propose novel sequence-to-sequence pre-training objectives for low-resource machine translation (NMT): Japanese-specific sequence to sequence (JASS) for language pairs involving Japanese as the source or target language, and English-specific sequence to sequence (ENSS) for language pairs involving English. JASS focuses on masking and reordering Japanese linguistic units known as bunsetsu, whereas ENSS is proposed based on phrase structure masking and reordering tasks. Experiments on ASPEC Japanese-English & Japanese-Chinese, Wikipedia Japanese-Chinese, News English-Korean corpora demonstrate that JASS and ENSS outperform MASS and other existing language-agnostic pre-training methods by up to +2.9 BLEU points for the Japanese-English tasks, up to +7.0 BLEU points for the Japanese-Chinese tasks and up to +1.3 BLEU points for English-Korean tasks. Empirical analysis, which focuses on the relationship between individual parts in JASS and ENSS, reveals the complementary nature of the subtasks of JASS and ENSS. Adequacy evaluation using LASER, human evaluation, and case studies reveals that our proposed methods significantly outperform pre-training methods without injected linguistic knowledge and they have a larger positive impact on the adequacy as compared to the fluency.
机构:
Tianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R ChinaTianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Xie, Yuting
Wang, Kun
论文数: 0引用数: 0
h-index: 0
机构:
Tianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Haihe Lab Brain Comp Interact & Human Machine Inte, Tianjin, Peoples R ChinaTianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Wang, Kun
Meng, Jiayuan
论文数: 0引用数: 0
h-index: 0
机构:
Tianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Tianjin Univ, Coll Precis Instruments & Optoelect Engn, Tianjin, Peoples R China
Haihe Lab Brain Comp Interact & Human Machine Inte, Tianjin, Peoples R ChinaTianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Meng, Jiayuan
Yue, Jin
论文数: 0引用数: 0
h-index: 0
机构:
Tianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R ChinaTianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Yue, Jin
Meng, Lin
论文数: 0引用数: 0
h-index: 0
机构:
Tianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Haihe Lab Brain Comp Interact & Human Machine Inte, Tianjin, Peoples R ChinaTianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Meng, Lin
Yi, Weibo
论文数: 0引用数: 0
h-index: 0
机构:
Haihe Lab Brain Comp Interact & Human Machine Inte, Tianjin, Peoples R China
Beijing Inst Mech Equipment, Beijing, Peoples R ChinaTianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Yi, Weibo
Jung, Tzyy-Ping
论文数: 0引用数: 0
h-index: 0
机构:
Tianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Tianjin Univ, Coll Precis Instruments & Optoelect Engn, Tianjin, Peoples R ChinaTianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Jung, Tzyy-Ping
Xu, Minpeng
论文数: 0引用数: 0
h-index: 0
机构:
Tianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Tianjin Univ, Coll Precis Instruments & Optoelect Engn, Tianjin, Peoples R China
Haihe Lab Brain Comp Interact & Human Machine Inte, Tianjin, Peoples R ChinaTianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Xu, Minpeng
Ming, Dong
论文数: 0引用数: 0
h-index: 0
机构:
Tianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China
Tianjin Univ, Coll Precis Instruments & Optoelect Engn, Tianjin, Peoples R China
Haihe Lab Brain Comp Interact & Human Machine Inte, Tianjin, Peoples R ChinaTianjin Univ, Acad Med Engn & Translat Med, Tianjin, Peoples R China