Leveraging Emoji to Improve Sentiment Classification of Tweets

被引:5
作者
de Barros, Tiago Martinho [1 ]
Pedrini, Helio [1 ]
Dias, Zanoni [1 ]
机构
[1] Univ Estadual Campinas, Inst Comp, Campinas, SP, Brazil
来源
36TH ANNUAL ACM SYMPOSIUM ON APPLIED COMPUTING, SAC 2021 | 2021年
基金
巴西圣保罗研究基金会;
关键词
Natural Language Processing; Sentiment Analysis; Emoji; Social Media;
D O I
10.1145/3412841.3441960
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Recent advances in the Natural Language Processing field have brought good results to a number of interesting tasks, for instance, Linguistic Acceptability, Question Answering, Reading Comprehension, Natural Language Inference, and Sentiment Analysis. Methods, such as ULMFiT, ELMo, BERT, and their derivatives, have achieved increasing success with these tasks, but often requiring substantial amounts of pre-training data and computational resources. We propose a novel methodology to classify the sentiment of tweets, based on BERT but focusing on emojis, treating them as an important source of sentiment as opposed to considering them simple input tokens. Additionally, it is possible to use a previously pre-trained BERT model to warm start ours, greatly reducing the training time required. Experiments on two Brazilian Portuguese datasets - TweetSentBR and 2000-tweets-BR - show that our methodology produces better results than BERT and outperforms the previously published results for both datasets, thus establishing new state-of-the-art results on TweetSentBR with accuracy of 0.7577 (4.8 percentage points absolute improvement) and F-1 score of 0.7395 (8.4 percentage points absolute improvement); and on 2000-tweets-BR with accuracy of 0.8316 (15.2 percentage points absolute improvement) and F-1 score of 0.8151 (24.5 percentage points absolute improvement).
引用
收藏
页码:845 / 852
页数:8
相关论文
共 31 条
[21]  
Sakiyama KM, 2019, IEEE IJCNN
[22]  
Souza Fabio, 2019, Computing Research Repository abs/1909, V10649, P1
[23]   Cloze Procedure: A New Tool For Measuring Readability [J].
Taylor, Wilson L. .
JOURNALISM QUARTERLY, 1953, 30 (04) :415-433
[24]   Survey on mining subjective data on the web [J].
Tsytsarau, Mikalai ;
Palpanas, Themis .
DATA MINING AND KNOWLEDGE DISCOVERY, 2012, 24 (03) :478-514
[25]  
Vaswani A, 2017, ADV NEUR IN, V30
[26]  
Vitorio Douglas, 2017, P 11 BRAZ S INF HUM, P43
[27]  
Wagner JA, 2018, PROCEEDINGS OF THE ELEVENTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION (LREC 2018), P4339
[28]  
Wang H., 2012, P ACL 2012 SYST DEM, P115
[29]  
Wu Yonghui., 2016, Computing Research Repository (CoRR)
[30]  
Yang ZL, 2019, ADV NEUR IN, V32