Unsupervised Domain-Agnostic Fake News Detection Using Multi-Modal Weak Signals

被引:3
作者
Silva, Amila [1 ]
Luo, Ling [1 ]
Karunasekera, Shanika [1 ]
Leckie, Christopher [1 ]
机构
[1] Univ Melbourne, Sch Comp & Informat Syst, Parkville, Vic 3010, Australia
关键词
unsupervised learning; Fake news detection; weak signals; SOCIAL MEDIA; INFORMATION;
D O I
10.1109/TKDE.2024.3392788
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
The emergence of social media as one of the main platforms for people to access news has enabled the wide dissemination of fake news, having serious impacts on society. Thus, it is really important to identify fake news with high confidence in a timely manner, which is not feasible using manual analysis. This has motivated numerous studies on automating fake news detection. Most of these approaches are supervised, which requires extensive time and labour to build a labelled dataset. Although there have been limited attempts at unsupervised fake news detection, their performance suffers due to not exploiting the knowledge from various modalities related to news records and due to the presence of various latent biases in the existing news datasets (e.g., unrealistic real and fake news distributions). To address these limitations, this work proposes an effective framework for unsupervised fake news detection, which first embeds the knowledge available in four modalities (i.e., source credibility, textual content, propagation speed, and user credibility) in news records and then proposes $(UMD)<^>{2}$(UMD)2, a novel noise-robust self-supervised learning technique, to identify the veracity of news records from the multi-modal embeddings. Also, we propose a novel technique to construct news datasets minimizing the latent biases in existing news datasets. Following the proposed approach for dataset construction, we produce a Large-scale Unlabelled News Dataset consisting 419,351 news articles related to COVID-19, acronymed as LUND-COVID. We trained the proposed unsupervised framework using LUND-COVID to exploit the potential of large datasets, and evaluate it using a set of existing labelled datasets. Our results show that the proposed unsupervised framework largely outperforms existing unsupervised baselines for different tasks such as multi-modal fake news detection, fake news early detection and few-shot fake news detection, while yielding notable improvements for unseen domains during training.
引用
收藏
页码:7283 / 7295
页数:13
相关论文
共 55 条
[1]  
Arevalo J, 2017, Arxiv, DOI arXiv:1702.01992
[2]  
Baevski A, 2022, Arxiv, DOI [arXiv:2202.03555, DOI 10.48550/ARXIV.2202.03555]
[3]  
Baly R, 2020, 58TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2020), P3364
[4]  
Baly R, 2018, 2018 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP 2018), P3528
[5]  
Berthon A, 2021, PR MACH LEARN RES, V139
[6]  
Bruff D., 2005, Notes Math., V20, P5
[7]  
Chakraborty A, 2016, PROCEEDINGS OF THE 2016 IEEE/ACM INTERNATIONAL CONFERENCE ON ADVANCES IN SOCIAL NETWORKS ANALYSIS AND MINING ASONAM 2016, P9, DOI 10.1109/ASONAM.2016.7752207
[8]  
Chen T, 2020, PR MACH LEARN RES, V119
[9]  
Chuang CY, 2022, Arxiv, DOI arXiv:2201.04309
[10]  
Cui LM, 2020, Arxiv, DOI arXiv:2006.00885