BD2TSumm: A Benchmark Dataset for Abstractive Disaster Tweet Summarization

被引:2
作者
Garg, Piyush Kumar [1 ]
Chakraborty, Roshni [2 ]
Dandapat, Sourav Kumar [1 ]
机构
[1] Indian Inst Technol Patna, Dept Comp Sci & Engn, Bihar, India
[2] ABV IIITM Gwalior, Dept Informat Technol, Gwalior, India
来源
ONLINE SOCIAL NETWORKS AND MEDIA | 2025年 / 45卷
关键词
Benchmark dataset; Abstractive summarization; Disaster; Tweet summarization; Social media;
D O I
10.1016/j.osnem.2024.100299
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Online social media platforms, such as Twitter, are mediums for valuable updates during disasters. However, the large scale of available information makes it difficult for humans to identify relevant information from the available information. An automatic summary of these tweets provides identification of relevant information easy and ensures a holistic overview of a disaster event to process the aid for disaster response. In literature, there are two types of abstractive disaster tweet summarization approaches based on the format of output summary: key-phrased-based (where summary is a set of key-phrases) and sentence-based (where summary is a paragraph consisting of sentences). Existing sentence-based abstractive approaches are either unsupervised or supervised. However, both types of approaches require a sizable amount of ground-truth summaries for training and/or evaluation such that they work on disaster events irrespective of type and location. The lack of abstractive disaster ground-truth summaries and guidelines for annotation motivates us to come up with a systematic procedure to create abstractive sentence ground-truth summaries of disaster events. Therefore, this paper presents a two-step systematic annotation procedure for sentence-based abstractive summary creation. Additionally, we release BD2TSumm, i.e., a benchmark ground-truth dataset for evaluating the sentence-based abstractive summarization approaches for disaster events. BD2TSumm consists of 15 ground-truth summaries belonging to 5 different continents and both natural and man-made disaster types. Furthermore, to ensure the high quality of the generated ground-truth summaries, we evaluate them qualitatively (using five metrics) and quantitatively (using two metrics). Finally, we compare 12 existing State-Of-The-Art (SOTA) abstractive summarization approaches on these ground-truth summaries using ROUGE-N F1-score.
引用
收藏
页数:10
相关论文
共 62 条
[11]   Tweet Summarization of News Articles: An Objective Ordering-Based Perspective [J].
Chakraborty, Roshni ;
Bhavsar, Maitry ;
Dandapat, Sourav Kumar ;
Chandra, Joydeep .
IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, 2019, 6 (04) :761-777
[12]   A Network Based Stratification Approach for Summarizing Relevant Comment Tweets of News Articles [J].
Chakraborty, Roshni ;
Bhavsar, Maitry ;
Dandapat, Sourav ;
Chandra, Joydeep .
WEB INFORMATION SYSTEMS ENGINEERING, WISE 2017, PT I, 2017, 10569 :33-48
[13]   An online multi-source summarization algorithm for text readability in topic-based search [J].
Curiel, Arturo ;
Gutierrez-Soto, Claudio ;
Rojano-Caceres, Jose-Rafael .
COMPUTER SPEECH AND LANGUAGE, 2021, 66 (66)
[14]  
Dongwook Lee, 2019, arXiv
[15]  
Duchi J, 2011, J MACH LEARN RES, V12, P2121
[16]   Ensemble Algorithms for Microblog Summarization [J].
Dutta, Soumi ;
Chandra, Vibhash ;
Mehra, Kanav ;
Das, Asit Kumar ;
Chakraborty, Tanmoy ;
Ghosh, Saptarshi .
IEEE INTELLIGENT SYSTEMS, 2018, 33 (03) :4-14
[17]   KEST: A graph-based keyphrase extraction technique for tweets summarization using Markov Decision Process [J].
Garg, Muskan ;
Kumar, Mukesh .
EXPERT SYSTEMS WITH APPLICATIONS, 2022, 209
[18]  
Garg P.K., 2022, CEUR WORKSHOP P, V3117, P91
[19]   ADSumm: annotated ground-truth summary datasets for disaster tweet summarization [J].
Garg, Piyush Kumar ;
Chakraborty, Roshni ;
Dandapat, Sourav Kumar .
SOCIAL NETWORK ANALYSIS AND MINING, 2024, 14 (01)
[20]  
Garg PK, 2024, Arxiv, DOI [arXiv:2405.06541, DOI 10.48550/ARXIV.2405.06541]