Duplicate Bug Report Detection: How Far Are We?

被引：6

作者：

Zhang, Ting ^{[1
]}

Han, Donggyun ^{[2
]}

Vinayakarao, Venkatesh ^{[3
]}

Irsan, Ivana Clairine ^{[1
]}

Xu, Bowen ^{[1
]}

Thung, Ferdian ^{[1
]}

Lo, David ^{[1
]}

Jiang, Lingxiao ^{[1
]}

机构：

[1] Singapore Management Univ, Singapore, Singapore

[2] Royal Holloway Univ London, London, England

[3] Chennai Math Inst, Chennai, Tamil Nadu, India

来源：

ACM TRANSACTIONS ON SOFTWARE ENGINEERING AND METHODOLOGY | 2023年 / 32卷 / 04期

基金：

新加坡国家研究基金会;

关键词：

Bug reports; duplicate bug report detection; deep learning; empirical study; PERFORMANCE;

D O I：

10.1145/3576042

中图分类号：

TP31 [计算机软件];

学科分类号：

081202 ; 0835 ;

摘要：

Many Duplicate Bug Report Detection (DBRD) techniques have been proposed in the research literature. The industry uses some other techniques. Unfortunately, there is insufficient comparison among them, and it is unclear how far we have been. This work fills this gap by comparing the aforementioned techniques. To compare them, we first need a benchmark that can estimate how a tool would perform if applied in a realistic setting today. Thus, we first investigated potential biases that affect the fair comparison of the accuracy of DBRD techniques. Our experiments suggest that data age and issue tracking system (ITS) choice cause a significant difference. Based on these findings, we prepared a new benchmark. We then used it to evaluate DBRD techniques to estimate better how far we have been. Surprisingly, a simpler technique outperforms recently proposed sophisticated techniques on most projects in our benchmark. In addition, we compared the DBRD techniques proposed in research with those used in Mozilla and VSCode. Surprisingly, we observe that a simple technique already adopted in practice can achieve comparable results as a recently proposed research tool. Our study gives reflections on the current state of DBRD, and we share our insights to benefit future DBRD research.

引用

页数：32

共 50 条

[31] Improving Performance of Automatic Duplicate Bug Reports Detection using Longest Common Sequence Introducing New Textual Features for Textual Similarity Detection
Neysiani, Behzad Soleimani
Babamir, Seyed Morteza
2019 IEEE 5TH CONFERENCE ON KNOWLEDGE BASED ENGINEERING AND INNOVATION (KBEI 2019), 2019, : 378 - 383
[32] Towards Accurate Duplicate Bug Retrieval using Deep Learning Techniques
Deshmukh, Jayati
Annervaz, K. M.
Podder, Sanjay
Sengupta, Shubhashis
Dubash, Neville
2017 IEEE INTERNATIONAL CONFERENCE ON SOFTWARE MAINTENANCE AND EVOLUTION (ICSME), 2017, : 115 - 124
[33] Deep Just-in-Time Defect Prediction: How Far Are We?
Zeng, Zhengran
Zhang, Yuqun
Zhang, Haotian
Zhang, Lingming
ISSTA '21: PROCEEDINGS OF THE 30TH ACM SIGSOFT INTERNATIONAL SYMPOSIUM ON SOFTWARE TESTING AND ANALYSIS, 2021, : 427 - 438
[34] Pharmaceutical services and health promotion: how far have we gone and how are we faring? Scientific output in pharmaceutical studies
Nakamura, Carina Akemi
Soares, Luciano
Farias, Mareni Rocha
Leite, Silvana Nair
BRAZILIAN JOURNAL OF PHARMACEUTICAL SCIENCES, 2014, 50 (04) : 773 - 782
[35] A first look at bug report templates on GitHub
Li, Hongyan
Yan, Meng
Sun, Weifeng
Liu, Xiao
Wu, Yunsong
JOURNAL OF SYSTEMS AND SOFTWARE, 2023, 202
[36] Challenges and Barriers Facing Women in the IS Workforce: How Far Have We Come?
Armstrong, Deborah J.
Riemenschneider, Cynthia K.
Reid, Margaret F.
Nelms, Jason E.
SIGMIS CPR 2011: PROCEEDINGS OF THE 2011 ACM SIGMIS COMPUTER PERSONNEL RESEARCH CONFERENCE, 2011, : 107 - 112
[37] A method of non-bug report identification from bug report repository
Polpinij, Jantima
ARTIFICIAL LIFE AND ROBOTICS, 2021, 26 (03) : 318 - 328
[38] Deep learning based identification of inconsistent method names: How far are we?
Wang, Taiming
Zhang, Yuxia
Jiang, Lin
Tang, Yi
Li, Guangjie
Liu, Hui
EMPIRICAL SOFTWARE ENGINEERING, 2025, 30 (01)
[39] A method of non-bug report identification from bug report repository
Jantima Polpinij
Artificial Life and Robotics, 2021, 26 : 318 - 328
[40] Duplicate Detection Exploiting Data Relationships
Herschel, Melanie
IT-INFORMATION TECHNOLOGY, 2009, 51 (04): : 231 - 234

← 1 2 3 4 5 →