Multilingual Hate Speech Detection: A Semi-Supervised Generative Adversarial Approach

被引:2
作者
Mnassri, Khouloud [1 ]
Farahbakhsh, Reza [1 ]
Crespi, Noel [1 ]
机构
[1] Inst Polytech Paris, Samovar Telecom SudParis, F-91120 Palaiseau, France
关键词
social media; hate speech; semisupervised; GAN; multilingual; PLMs; DATA AUGMENTATION; NETWORKS;
D O I
10.3390/e26040344
中图分类号
O4 [物理学];
学科分类号
0702 ;
摘要
Social media platforms have surpassed cultural and linguistic boundaries, thus enabling online communication worldwide. However, the expanded use of various languages has intensified the challenge of online detection of hate speech content. Despite the release of multiple Natural Language Processing (NLP) solutions implementing cutting-edge machine learning techniques, the scarcity of data, especially labeled data, remains a considerable obstacle, which further requires the use of semisupervised approaches along with Generative Artificial Intelligence (Generative AI) techniques. This paper introduces an innovative approach, a multilingual semisupervised model combining Generative Adversarial Networks (GANs) and Pretrained Language Models (PLMs), more precisely mBERT and XLM-RoBERTa. Our approach proves its effectiveness in the detection of hate speech and offensive language in Indo-European languages (in English, German, and Hindi) when employing only 20% annotated data from the HASOC2019 dataset, thereby presenting significantly high performances in each of multilingual, zero-shot crosslingual, and monolingual training scenarios. Our study provides a robust mBERT-based semisupervised GAN model (SS-GAN-mBERT) that outperformed the XLM-RoBERTa-based model (SS-GAN-XLM) and reached an average F1 score boost of 9.23% and an accuracy increase of 5.75% over the baseline semisupervised mBERT model.
引用
收藏
页数:19
相关论文
共 50 条
  • [1] Semi-Supervised Self-Learning for Arabic Hate Speech Detection
    Alsafari, Safa
    Sadaoui, Samira
    [J]. 2021 IEEE INTERNATIONAL CONFERENCE ON SYSTEMS, MAN, AND CYBERNETICS (SMC), 2021, : 863 - 868
  • [2] Internet, social media and online hate speech. Systematic review
    Andres Castano-Pulgarin, Sergio
    Suarez-Betancur, Natalia
    Tilano Vega, Luz Magnolia
    Herrera Lopez, Harvey Mauricio
    [J]. AGGRESSION AND VIOLENT BEHAVIOR, 2021, 58
  • [3] Hate Speech: A Systematized Review
    Antonia Paz, Maria
    Montero-Diaz, Julio
    Moreno-Delgado, Alicia
    [J]. SAGE OPEN, 2020, 10 (04):
  • [4] Auti Tapan, 2022, P 1 COMP SOC RESP WO, P52
  • [5] Brown TB, 2020, ADV NEUR IN, V33
  • [6] Cao R., 2020, P 28 INT C COMP LING, P6327
  • [7] Freepalestine on TikTok: from performative activism to (meaningful) playful activism
    Cervi, Laura
    Marin-Llado, Carles
    [J]. JOURNAL OF INTERNATIONAL AND INTERCULTURAL COMMUNICATION, 2022, 15 (04) : 414 - 434
  • [8] An Empirical Survey of Data Augmentation for Limited Data Learning in NLP
    Chen, Jiaao
    Tam, Derek
    Raffel, Colin
    Bansal, Mohit
    Yang, Diyi
    [J]. TRANSACTIONS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, 2023, 11 : 191 - 211
  • [9] LMGAN: Linguistically Informed Semi-Supervised GAN with Multiple Generators
    Cho, Whanhee
    Choi, Yongsuk
    [J]. SENSORS, 2022, 22 (22)
  • [10] Conneau A., 2020, P 58 ANN M ASS COMPU, P8440