xcomet: Transparent Machine Translation Evaluation through Fine-grained Error Detection

被引:1
|
作者
Guerreiro, Nuno M. [1 ,3 ,4 ,5 ]
Rei, Ricardo [1 ,2 ,5 ]
van Stigt, Daan [1 ]
Coheur, Luisa [2 ,5 ]
Colombo, Pierre [4 ]
Martins, Andre F. T. [1 ,3 ,5 ]
机构
[1] Unbabel Lisbon, Lisbon, Portugal
[2] INESC ID, Lisbon, Portugal
[3] Inst Telecomunicacoes, Lisbon, Portugal
[4] Univ Paris Saclay, MICS, Cent Supelec, Paris, France
[5] Univ Lisbon, Inst Super Tecn, Lisbon, Portugal
基金
欧洲研究理事会;
关键词
Compendex;
D O I
10.1162/tacl_a_00683
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Widely used learned metrics for machine translation evaluation, such as Comet and Bleurt, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation errors (e.g., what are the errors and what is their severity). On the other hand, generative large language models (LLMs) are amplifying the adoption of more granular strategies to evaluation, attempting to detail and categorize translation errors. In this work, we introduce xcomet, an open-source learned metric designed to bridge the gap between these approaches. xcomet integrates both sentence-level evaluation and error span detection capabilities, exhibiting state-of-the-art performance across all types of evaluation (sentence-level, system-level, and error span detection). Moreover, it does so while highlighting and categorizing error spans, thus enriching the quality assessment. We also provide a robustness analysis with stress tests, and show that xcomet is largely capable of identifying localized critical errors and hallucinations.
引用
收藏
页码:979 / 995
页数:17
相关论文
共 4 条
  • [1] Power system terminal continuous trust evaluation model based on fine-grained data flow analysis
    Xie, Ming
    Proceedings of SPIE - The International Society for Optical Engineering, 2022, 12158
  • [2] Multi-scale Cross-attention Network for Multi-family Fine-grained Malicious Domain Name Detection
    Zhang, Qing
    Zhang, Wen-Chuan
    International Journal of Network Security, 2024, 26 (06): : 1082 - 1091
  • [3] RETRACTED ARTICLE: A Method of Tracking Visual Targets in Fine-Grained Image Using Machine Learning(IETE JOURNAL OF RESEARCH, (2023), 69, (10), (c–cix))
    Ma, Xiao
    Ye, Yufei
    Chen, Leihang
    Tao, Haibo
    Liao, Cancan
    IETE Journal of Research, 2023, 69 (10)
  • [4] Gas recovery enhancement from fine-grained hydrate reservoirs through positive inter-branch interference and optimized spiral multilateral well network
    Mao, Peixiao
    Wu, Nengyou
    Wan, Yizhao
    Ning, Fulong
    Sun, Jiaxin
    Wang, Xingxing
    Hu, Gaowei
    Journal of Natural Gas Science and Engineering, 2022, 107