WASM: A Dataset for Hashtag Recommendation for Arabic Tweets

被引:1
|
作者
Al-Shaibani, Maged S. [1 ]
Luqman, Hamzah [1 ,2 ]
Al-Ghofaily, Abdulaziz S. [1 ]
Al-Najim, Abdullatif A. [1 ]
机构
[1] King Fahd Univ Petr & Minerals, Informat & Comp Sci Dept, Dhahran, Saudi Arabia
[2] SDAIA KFUPM Joint Res Ctr Artificial Intelligence, Dhahran 31261, Saudi Arabia
关键词
Hashtag Recommendation; Hashtag Generation; Tweets Classification; Arabic Tweets; Twitter; Hashtags;
D O I
10.1007/s13369-023-08567-1
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
As one of the largest microblogging websites in the world, Twitter generates a huge amount of information daily. The massive size of the generated data increases the difficulty for humans to follow and receive information relevant to their interests. Therefore, Twitter allows users to annotate and categorize their tweets using appropriate hashtags. However, finding an appropriate hashtag for a tweet is not always straightforward. Furthermore, many users violate the hashtag flow by posting irrelevant content to the hashtag topic. These problems increase the need for a hashtag recommendation and classification system. This topic has received considerable attention from researchers in some languages, such as English and Chinese. However, this problem has not yet been explored for the Arabic language owing to the lack of datasets. In this study, we bridge this gap by proposing WASM, an Arabic Twitter hashtag recommendation dataset consisting of more than 100,000 tweets annotated with 87 hashtags. The proposed dataset is subjected to several rounds of automatic and manual filtrations to ensure that it is suitable for tasks related to tweets and hashtags. Further, we propose three systems for hashtag recommendation and classification. Each of these systems approaches the task differently by considering it as classification, generation, and named entity recognition problems. The results obtained using these systems are promising and can be used to benchmark the WASM dataset. The data and code are available at https://github.com/Hamzah-Luqman/wasm.
引用
收藏
页码:12131 / 12145
页数:15
相关论文
共 50 条
  • [21] Research topics and trends of the hashtag recommendation domain
    Babak Amiri
    Ramin Karimianghadim
    Navid Yazdanjue
    Liaquat Hossain
    Scientometrics, 2021, 126 : 2689 - 2735
  • [22] Personalized Hashtag Recommendation for Micro-videos
    Wei, Yinwei
    Cheng, Zhiyong
    Yu, Xuzheng
    Zhao, Zhou
    Zhu, Lei
    Nie, Liqiang
    PROCEEDINGS OF THE 27TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA (MM'19), 2019, : 1446 - 1454
  • [23] Hashtag Recommendation Using Word Sequences' Embeddings
    Ben-Lhachemi, Nada
    Nfaoui, El Habib
    BIG DATA, CLOUD AND APPLICATIONS, BDCA 2018, 2018, 872 : 131 - 143
  • [24] Using Topic Models for Twitter Hashtag Recommendation
    Godin, Frederic
    Slavkovikj, Viktor
    De Neve, Wesley
    Schrauwen, Benjamin
    Van de Walle, Rik
    PROCEEDINGS OF THE 22ND INTERNATIONAL CONFERENCE ON WORLD WIDE WEB (WWW'13 COMPANION), 2013, : 593 - 596
  • [25] Research topics and trends of the hashtag recommendation domain
    Amiri, Babak
    Karimianghadim, Ramin
    Yazdanjue, Navid
    Hossain, Liaquat
    SCIENTOMETRICS, 2021, 126 (04) : 2689 - 2735
  • [26] The pragmatic functions of emojis in Arabic tweets
    Alharbi, Amjad
    Mahzari, Mohammad
    FRONTIERS IN PSYCHOLOGY, 2023, 13
  • [27] Empirical Analysis of Factors Influencing Twitter Hashtag Recommendation on Detected Communities
    Alsini, Areej
    Datta, Amitava
    Li, Jianxin
    Huynh, Du
    ADVANCED DATA MINING AND APPLICATIONS, ADMA 2017, 2017, 10604 : 119 - 131
  • [28] Emotional Tone Detection in Arabic Tweets
    Al-Khatib, Amr
    El-Beltagy, Samhaa R.
    COMPUTATIONAL LINGUISTICS AND INTELLIGENT TEXT PROCESSING, CICLING 2017, PT II, 2018, 10762 : 105 - 114
  • [29] Detecting Epidemic Diseases Using Sentiment Analysis of Arabic Tweets
    Baker, Qanita Bani
    Shatnawi, Farah
    Rawashdeh, Saif
    Al-Smadi, Mohammad
    Jararweh, Yaser
    JOURNAL OF UNIVERSAL COMPUTER SCIENCE, 2020, 26 (01) : 50 - 70
  • [30] Big Data Contextual Analytics Study on Arabic Tweets Summarization
    Al-Ibrahim, Fatimah
    Alzamil, Zakarya A.
    INTERNATIONAL JOURNAL OF KNOWLEDGE AND SYSTEMS SCIENCE, 2019, 10 (04) : 18 - 34