Grapharizer: A Graph-Based Technique for Extractive Multi-Document Summarization

被引:6
作者
Jalil, Zakia [1 ]
Nasir, Muhammad [2 ]
Alazab, Moutaz [3 ]
Nasir, Jamal [4 ]
Amjad, Tehmina [1 ]
Alqammaz, Abdullah [5 ]
机构
[1] Int Islamic Univ, Dept Comp Sci, Islamabad 44000, Pakistan
[2] Int Islamic Univ, Dept Software Engn, Islamabad 44000, Pakistan
[3] Al Balqa Appl Univ, Fac Artificial Intelligence, Dept Intelligent Syst, Salt 19117, Jordan
[4] Univ Galway, Sch Comp Sci, Galway H91TK33, Ireland
[5] Zarqa Univ, Coll Informat Technol, Dept Cyber Secur, Zarqa 13110, Jordan
关键词
big data; automatic text summarization; extractive multi-document summarization; graph theory; machine learning; anaphora; cataphora; pronoun resolution; grammaticality; topic modeling; ChatGPT; TEXT; SEARCH;
D O I
10.3390/electronics12081895
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Featured Application A graph-based technique tested on a benchmark dataset and augmented by machine learning techniques to provide a concise, informative, and grammatically correct summary. In the age of big data, there is increasing growth of data on the Internet. It becomes frustrating for users to locate the desired data. Therefore, text summarization emerges as a solution to this problem. It summarizes and presents the users with the gist of the provided documents. However, summarizer systems face challenges, such as poor grammaticality, missing important information, and redundancy, particularly in multi-document summarization. This study involves the development of a graph-based extractive generic MDS technique, named Grapharizer (GRAPH-based summARIZER), focusing on resolving these challenges. Grapharizer addresses the grammaticality problems of the summary using lemmatization during pre-processing. Furthermore, synonym mapping, multi-word expression mapping, and anaphora and cataphora resolution, contribute positively to improving the grammaticality of the generated summary. Challenges, such as redundancy and proper coverage of all topics, are dealt with to achieve informativity and representativeness. Grapharizer is a novel approach which can also be used in combination with different machine learning models. The system was tested on DUC 2004 and Recent News Article datasets against various state-of-the-art techniques. Use of Grapharizer with machine learning increased accuracy by up to 23.05% compared with different baseline techniques on ROUGE scores. Expert evaluation of the proposed system indicated the accuracy to be more than 55%.
引用
收藏
页数:26
相关论文
共 50 条
  • [31] CENTRANK: A GRAPH CENTROID BASED RANKING FOR EXTRACTIVE TEXT SUMMARIZATION
    Hemamalini, S.
    Swaminathan, V.
    TWMS JOURNAL OF APPLIED AND ENGINEERING MATHEMATICS, 2023, 13 : 355 - 364
  • [32] EdgeSumm: Graph-based framework for automatic text summarization
    El-Kassas, Wafaa S.
    Salama, Cherif R.
    Rafea, Ahmed A.
    Mohamed, Hoda K.
    INFORMATION PROCESSING & MANAGEMENT, 2020, 57 (06)
  • [33] A developed framework for multi-document summarization using softmax regression and spider monkey optimization methods
    Wilson, Praveen K.
    Jeba, J. R.
    SOFT COMPUTING, 2022, 26 (07) : 3313 - 3328
  • [34] Heterogeneous-Length Text Topic Modeling for Reader-Aware Multi-Document Summarization
    Qiang, Jipeng
    Chen, Ping
    Ding, Wei
    Wang, Tong
    Xie, Fei
    Wu, Xindong
    ACM TRANSACTIONS ON KNOWLEDGE DISCOVERY FROM DATA, 2019, 13 (04)
  • [35] Improving Graph-Based Summarization with HTML']HTML Tag and Metadata Features
    Wardani, Dewi
    Susanti, Yuni
    ENGINEERING LETTERS, 2020, 28 (02) : 522 - 528
  • [36] A Probabilistic Approach for Extractive Summarization Based on Clustering Cum Graph Ranking Method
    Ahmad, Amreen
    Ahmad, Tanvir
    Masood, Sarfaraz
    Siddiqui, Mohd Khizir
    Abd El-Rahiem, Basma
    Plawiak, Pawel
    Alblehai, Fahad
    IEEE ACCESS, 2024, 12 : 70464 - 70479
  • [37] An Arabic Multi-source News Corpus: Experimenting on Single-document Extractive Summarization
    Amina Chouigui
    Oussama Ben Khiroun
    Bilel Elayeb
    Arabian Journal for Science and Engineering, 2021, 46 : 3925 - 3938
  • [38] An Arabic Multi-source News Corpus: Experimenting on Single-document Extractive Summarization
    Chouigui, Amina
    Ben Khiroun, Oussama
    Elayeb, Bilel
    ARABIAN JOURNAL FOR SCIENCE AND ENGINEERING, 2021, 46 (04) : 3925 - 3938
  • [39] Fuzzy logic based multi document summarization with improved sentence scoring and redundancy removal technique
    Patel, Darshna
    Shah, Saurabh
    Chhinkaniwala, Hitesh
    EXPERT SYSTEMS WITH APPLICATIONS, 2019, 134 : 167 - 177
  • [40] Impact of Similarity Measures in Graph-based Automatic Text Summarization of Konkani Texts
    D'Silva, Jovi
    Sharma, Uzzal
    ACM TRANSACTIONS ON ASIAN AND LOW-RESOURCE LANGUAGE INFORMATION PROCESSING, 2023, 22 (02)