Detecting Tweets Containing Cannabidiol-Related COVID-19 Misinformation Using Transformer Language Models and Warning Letters From Food and Drug Administration: Content Analysis and Identification

被引:1
作者
Turner, Jason [1 ,4 ]
Kantardzic, Mehmed [1 ]
Vickers-Smith, Rachel [2 ]
Brown, Andrew G. [3 ]
机构
[1] Univ Louisville, J B Speed Sch Engn, Dept Comp Sci & Engn, Data Min Lab, Louisville, KY USA
[2] Univ Kentucky, Coll Publ Hlth, Dept Epidemiol & Environm Hlth, Lexington, KY USA
[3] No Arizona Univ, Dept Criminol & Criminal Justice, Flagstaff, AZ USA
[4] Univ Louisville, Engn Speed Sch Engn, Data Min Lab Dept Comp Sci, Louisville, KY 40292 USA
来源
JMIR INFODEMIOLOGY | 2023年 / 3卷 / 01期
关键词
transformer; misinformation; deep learning; COVID-19; infodemic; pandemic; language model; health information; social media; Twitter; content analysis; cannabidiol; sentence vector; infodemiology;
D O I
10.2196/38390
中图分类号
R19 [保健组织与事业(卫生事业管理)];
学科分类号
摘要
Background: COVID-19 has introduced yet another opportunity to web-based sellers of loosely regulated substances, such as cannabidiol (CBD), to promote sales under false pretenses of curing the disease. Therefore, it has become necessary to innovate ways to identify such instances of misinformation. Objective: We sought to identify COVID-19 misinformation as it relates to the sales or promotion of CBD and used transformer-based language models to identify tweets semantically similar to quotes taken from known instances of misinformation. In this case, the known misinformation was the publicly available Warning Letters from Food and Drug Administration (FDA). Methods: We collected tweets using CBD- and COVID-19-related terms. Using a previously trained model, we extracted the tweets indicating commercialization and sales of CBD and annotated those containing COVID-19 misinformation according to the FDA definitions. We encoded the collection of tweets and misinformation quotes into sentence vectors and then calculated the cosine similarity between each quote and each tweet. This allowed us to establish a threshold to identify tweets that were making false claims regarding CBD and COVID-19 while minimizing the instances of false positives. Results: We demonstrated that by using quotes taken from Warning Letters issued by FDA to perpetrators of similar misinformation, we can identify semantically similar tweets that also contain misinformation. This was accomplished by identifying a cosine distance threshold between the sentence vectors of the Warning Letters and tweets. Conclusions: This research shows that commercial CBD or COVID-19 misinformation can potentially be identified and curbed using transformer-based language models and known prior instances of misinformation. Our approach functions without the need for labeled data, potentially reducing the time at which misinformation can be identified. Our approach shows promise in that it is easily adapted to identify other forms of misinformation related to loosely regulated substances.
引用
收藏
页数:16
相关论文
共 57 条
[1]   Lies Kill, Facts Save: Detecting COVID-19 Misinformation in Twitter [J].
Al-Rakhami, Mabrook S. ;
Al-Amri, Atif M. .
IEEE ACCESS, 2020, 8 :155961-155970
[2]  
[Anonymous], Warning letter CBD online store
[3]  
[Anonymous], 2020, Fake news in the time of C-19
[4]  
[Anonymous], Warning Letters
[5]  
[Anonymous], 2010, REGULATORY PROCEDURE
[6]  
[Anonymous], 2020, FDA cautions against use of hydroxychloroquine or chloroquine for COVID. -19 outside of the hospital setting or a clinical trial due to risk of heart rhythm problems
[7]  
[Anonymous], Warning letter Nova Botanix LTD DBA CanaBD
[8]  
[Anonymous], SOCIAL NETWORKING SE
[9]  
[Anonymous], Warning letter Avazo-healthcare, llc
[10]  
[Anonymous], 2022, Fraudulent Coronavirus Disease 2019 (COVID-19) products.