Grapharizer: A Graph-Based Technique for Extractive Multi-Document Summarization

被引:6
作者
Jalil, Zakia [1 ]
Nasir, Muhammad [2 ]
Alazab, Moutaz [3 ]
Nasir, Jamal [4 ]
Amjad, Tehmina [1 ]
Alqammaz, Abdullah [5 ]
机构
[1] Int Islamic Univ, Dept Comp Sci, Islamabad 44000, Pakistan
[2] Int Islamic Univ, Dept Software Engn, Islamabad 44000, Pakistan
[3] Al Balqa Appl Univ, Fac Artificial Intelligence, Dept Intelligent Syst, Salt 19117, Jordan
[4] Univ Galway, Sch Comp Sci, Galway H91TK33, Ireland
[5] Zarqa Univ, Coll Informat Technol, Dept Cyber Secur, Zarqa 13110, Jordan
关键词
big data; automatic text summarization; extractive multi-document summarization; graph theory; machine learning; anaphora; cataphora; pronoun resolution; grammaticality; topic modeling; ChatGPT; TEXT; SEARCH;
D O I
10.3390/electronics12081895
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Featured Application A graph-based technique tested on a benchmark dataset and augmented by machine learning techniques to provide a concise, informative, and grammatically correct summary. In the age of big data, there is increasing growth of data on the Internet. It becomes frustrating for users to locate the desired data. Therefore, text summarization emerges as a solution to this problem. It summarizes and presents the users with the gist of the provided documents. However, summarizer systems face challenges, such as poor grammaticality, missing important information, and redundancy, particularly in multi-document summarization. This study involves the development of a graph-based extractive generic MDS technique, named Grapharizer (GRAPH-based summARIZER), focusing on resolving these challenges. Grapharizer addresses the grammaticality problems of the summary using lemmatization during pre-processing. Furthermore, synonym mapping, multi-word expression mapping, and anaphora and cataphora resolution, contribute positively to improving the grammaticality of the generated summary. Challenges, such as redundancy and proper coverage of all topics, are dealt with to achieve informativity and representativeness. Grapharizer is a novel approach which can also be used in combination with different machine learning models. The system was tested on DUC 2004 and Recent News Article datasets against various state-of-the-art techniques. Use of Grapharizer with machine learning increased accuracy by up to 23.05% compared with different baseline techniques on ROUGE scores. Expert evaluation of the proposed system indicated the accuracy to be more than 55%.
引用
收藏
页数:26
相关论文
共 50 条
  • [21] A New Automatic Multi-document Text Summarization using Topic Modeling
    Roul, Rajendra Kumar
    Mehrotra, Samarth
    Pungaliya, Yash
    Sahoo, Jajati Keshari
    DISTRIBUTED COMPUTING AND INTERNET TECHNOLOGY, ICDCIT 2019, 2019, 11319 : 212 - 221
  • [22] A Novel Contextual Topic Model for Query-focused Multi-document Summarization
    Yang, Guangbing
    2014 IEEE 26TH INTERNATIONAL CONFERENCE ON TOOLS WITH ARTIFICIAL INTELLIGENCE (ICTAI), 2014, : 576 - 583
  • [23] Mining Both Commonality and Specificity From Multiple Documents for Multi-Document Summarization
    Ma, Bing
    IEEE ACCESS, 2024, 12 : 54371 - 54381
  • [24] Assessing shallow sentence scoring techniques and combinations for single and multi-document summarization
    Oliveira, Hilario
    Ferreira, Rafael
    Lima, Rinaldo
    Lins, Rafael Dueire
    Freitas, Fred
    Riss, Marcelo
    Simske, Steven J.
    EXPERT SYSTEMS WITH APPLICATIONS, 2016, 65 : 68 - 86
  • [25] Extracting main content of a topic on online social network by multi-document summarization
    Liu, Chunyan
    Zhu, Conghui
    Zhao, Tiejun
    Zheng, Dequan
    PROCEEDINGS OF THE 2012 EIGHTH INTERNATIONAL CONFERENCE ON COMPUTATIONAL INTELLIGENCE AND SECURITY (CIS 2012), 2012, : 52 - 55
  • [26] Multi Document Summarization Based On Cross-Document Relation Using Voting Technique
    Kumar, Yogan Jaya
    Salim, Naomie
    Abuobieda, Albaraa
    Tawfik, Ameer
    2013 INTERNATIONAL CONFERENCE ON COMPUTING, ELECTRICAL AND ELECTRONICS ENGINEERING (ICCEEE), 2013, : 609 - 614
  • [27] CLUSTERING TECHNIQUES AND DISCRETE PARTICLE SWARM OPTIMIZATION ALGORITHM FOR MULTI-DOCUMENT SUMMARIZATION
    Aliguliyev, Ramiz M.
    COMPUTATIONAL INTELLIGENCE, 2010, 26 (04) : 420 - 448
  • [28] Extractive Document Summarization Based on Dynamic Feature Space Mapping
    Ghodratnama, Samira
    Beheshti, Amin
    Zakershahrak, Mehrdad
    Sobhanmanesh, Fariborz
    IEEE ACCESS, 2020, 8 : 139084 - 139095
  • [29] A hybrid deep learning architecture for opinion-oriented multi-document summarization based on multi-feature fusion
    Abdi, Asad
    Hasan, Shafaatunnur
    Shamsuddin, Siti Mariyam
    Idris, Norisma
    Piran, Jalil
    KNOWLEDGE-BASED SYSTEMS, 2021, 213
  • [30] What we achieve on text extractive summarization based on graph?
    Chen, Shuang
    Ren, Tao
    Qv, Ying
    Shi, Yang
    JOURNAL OF INTELLIGENT & FUZZY SYSTEMS, 2022, 43 (06) : 7057 - 7065