Automated Question Answering for Improved Understanding of Compliance Requirements: A Multi-Document Study

被引:19
作者
Abualhaija, Sallam [1 ]
Arora, Chetan [1 ,2 ]
Sleimi, Amin [1 ]
Briand, Lionel C. [1 ,3 ]
机构
[1] Univ Luxembourg, SnT Ctr Secur Reliabil & Trust, Luxembourg, Luxembourg
[2] Deakin Univ, Geelong, Vic, Australia
[3] Univ Ottawa, Sch Elect Engn & Comp Sci, Ottawa, ON, Canada
来源
2022 30TH IEEE INTERNATIONAL REQUIREMENTS ENGINEERING CONFERENCE (RE 2022) | 2022年
基金
加拿大自然科学与工程研究理事会;
关键词
Requirements Engineering; Regulatory Compliance; Natural Language Processing (NLP); Question Answering; Language Models (LMs); BERT;
D O I
10.1109/RE54965.2022.00011
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
Software systems are increasingly subject to regulatory compliance. Extracting compliance requirements from regulations is challenging. Ideally, locating compliance-related information in a regulation requires a joint effort from requirements engineers and legal experts, whose availability is limited. However, regulations are typically long documents spanning hundreds of pages, containing legal jargon, applying complicated natural language structures, and including cross-references, thus making their analysis effort-intensive. In this paper, we propose an automated question-answering (QA) approach that assists requirements engineers in finding the legal text passages relevant to compliance requirements. Our approach utilizes large-scale language models fine-tuned for QA, including BERT and three variants. We evaluate our approach on 107 question-answer pairs, manually curated by subject-matter experts, for four different European regulatory documents. Among these documents is the general data protection regulation (GDPR) - a major source for privacy-related requirements. Our empirical results show that, in similar to 94% of the cases, our approach finds the text passage containing the answer to a given question among the top five passages that our approach marks as most relevant. Further, our approach successfully demarcates, in the selected passage, the right answer with an average accuracy of similar to 91%.
引用
收藏
页码:39 / 50
页数:12
相关论文
共 54 条
[1]  
Abualhaija S., 2022, ONLINE ANNEX
[2]   An information-theoretic perspective of tf-idf measures [J].
Aizawa, A .
INFORMATION PROCESSING & MANAGEMENT, 2003, 39 (01) :45-65
[3]  
Alexandrescu A., 2006, Proceedings of the human language technology conference of the naacl, companion volume: Short papers, P1, DOI DOI 10.3115/1614049.1614050
[4]  
[Anonymous], 2020, Law of 25 March 2020 (coordinated version) establishing a central electronic data retrieval system related to IBAN accounts and safe-deposit boxes
[5]  
[Anonymous], 2002, P ACL 02 WORKSH EFF, DOI 10.3115/1225403.1225421
[6]  
[Anonymous], 2019, OJL, V136, P1
[7]  
[Anonymous], 2019, DIR EU 2019 771 EUR, P28
[8]   NARCIA: An Automated Tool for Change Impact Analysis in Natural Language Requirements [J].
Arora, Chetan ;
Sabetzadeh, Mehrdad ;
Goknil, Arda ;
Briand, Lionel C. ;
Zimmer, Frank .
2015 10TH JOINT MEETING OF THE EUROPEAN SOFTWARE ENGINEERING CONFERENCE AND THE ACM SIGSOFT SYMPOSIUM ON THE FOUNDATIONS OF SOFTWARE ENGINEERING (ESEC/FSE 2015) PROCEEDINGS, 2015, :962-965
[9]  
Arora C, 2015, INT REQUIR ENG CONF, P6, DOI 10.1109/RE.2015.7320403
[10]  
Clark K, 2020, Arxiv, DOI arXiv:2003.10555