Building a deep learning-based QA system from a CQA dataset

被引:3
作者
Jin, Sol [1 ]
Lian, Xu [1 ]
Jung, Hanearl [1 ]
Park, Jinsoo [1 ]
Suh, Jihae [2 ]
机构
[1] Seoul Natl Univ, Coll Business Adm, Seoul, South Korea
[2] Seoul Natl Univ Sci & Technol, Coll Business Adm, Seoul, South Korea
关键词
Question answering (QA) system; Community question answering (CQA); BERT; T5; DECISION-SUPPORT; QUESTION; ANSWERS;
D O I
10.1016/j.dss.2023.114038
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
A man-made machine-reading comprehension (MRC) dataset is necessary to train the answer extraction part of existing Question Answering (QA) systems. However, a high-quality and well-structured dataset with question-paragraph-answer pairs is not usually found in the real world. Furthermore, updating or building an MRC dataset is a challenging and costly affair. To address these shortcomings, we propose a QA system that uses a large-scale English Community Question Answering (CQA) dataset (i.e., Stack Exchange) composed of 3,081,834 question-answer pairs. The QA system adopts a classifier-retriever-summarizer structure design. The question classifier and the answer retriever part are based on a Bidirectional Encoder Representations from Transformers (BERT) Natural Language Processing (NLP) model by Google, and the summarizer part introduces a deep learning-based Text-to-Text Transfer Transformer (T5) model to summarize the long answers. We instantiated the proposed QA system with 140 topics from the CQA dataset (including topics such as biology, law, politics, etc.) and conducted human and automatic evaluations. Our system presented encouraging results, considering that it provides high-quality answers to the questions in the test set and satisfied the requirements to develop a QA system without MRC datasets. Our results show the potential of building automatic and high-performance QA systems without being limited by man-made datasets, a significant step forward in the research of open-domain or specific-domain QA systems.
引用
收藏
页数:12
相关论文
共 32 条
  • [31] Methodology Based on BERT (Bidirectional Encoder Representations from Transformers) to Improve Solar Irradiance Prediction of Deep Learning Models Trained with Time Series of Spatiotemporal Meteorological Information
    Benavides-Cesar, Llinet
    Manso-Callejo, Miguel-angel
    Cira, Calimanut-Ionut
    FORECASTING, 2025, 7 (01):
  • [32] Use of BERT (Bidirectional Encoder Representations from Transformers)-Based Deep Learning Method for Extracting Evidences in Chinese Radiology Reports: Development of a Computer-Aided Liver Cancer Diagnosis Framework
    Liu, Honglei
    Zhang, Zhiqiang
    Xu, Yan
    Wang, Ni
    Huang, Yanqun
    Yang, Zhenghan
    Jiang, Rui
    Chen, Hui
    JOURNAL OF MEDICAL INTERNET RESEARCH, 2021, 23 (01)