Improved POS Tagging Model for Malay Twitter Data based on Machine Learning Algorithm

被引:0
|
作者
Ariffin, Siti Noor Allia Noor [1 ]
Tiun, Sabrina [1 ]
机构
[1] Univ Kebangsaan Malaysia, Fac Informat Sci & Technol, Bangi, Selangor, Malaysia
关键词
Informal Malay; Malay Twitter corpus; Malay POS tagging; Malay POS tagger model; Malay social media texts; Malay POS machine learning; SVM;
D O I
10.14569/IJACSA.2022.0130730
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Twitter is a popular social media platform in Malaysia that allows for 280-character microblogging. Almost everything that happens in a single day is tweeted by users. Because of the popularity of Twitter, most Malaysians use it daily, providing researchers and developers with a wealth of data on Malaysian users. This paper explains why and how this study chose to create a new Malay Twitter corpus, Malay Part-of-Speech (POS) tags, and a Malay POS tagger model. The goal of this paper is to improve existing Malay POS tags so that they are more compatible with the newly created Malay Twitter corpus, as well as to build a POS tagging model specifically tailored for Malay Twitter data using various machine learning algorithms. For instance, Support Vector Machine (SVM), Naive Bayes (NB), Decision Tree (DT), and K-Nearest Neighbor (KNN) classifiers. This study's data was gathered by using Twitter's Advanced Search function and relevant and related keywords associated with informal Malay. The data was fed into machine learning algorithms after several stages of processing to serve as the training and testing corpus. The evaluation and analysis of the developed Malay POS tagger model show that the SVM classifier, as well as the newly proposed Malay POS tags, is the best machine learning algorithm for Malay Twitter data. Furthermore, the prediction accuracy and POS tagging results show that this research outperformed a comparable previous study, indicating that the Malay POS tagger model and its POS were successfully improved.
引用
收藏
页码:229 / 234
页数:6
相关论文
共 50 条
  • [31] An efficient sentimental analysis using hybrid deep learning and optimization technique for Twitter using parts of speech (POS) tagging
    M. Divyapushpalakshmi
    R. Ramalakshmi
    International Journal of Speech Technology, 2021, 24 : 329 - 339
  • [32] Evolutionary extreme learning machine based on an improved MOPSO algorithm
    Qinghua Ling
    Kaimin Tan
    Yuyan Wang
    Zexu Li
    Wenkai Liu
    Neural Computing and Applications, 2025, 37 (12) : 7733 - 7750
  • [33] An efficient sentimental analysis using hybrid deep learning and optimization technique for Twitter using parts of speech (POS) tagging
    Divyapushpalakshmi, M.
    Ramalakshmi, R.
    INTERNATIONAL JOURNAL OF SPEECH TECHNOLOGY, 2021, 24 (02) : 329 - 339
  • [34] Geospatial data extraction algorithm based on machine learning
    Zhu, Xiao-Long
    Xie, Zhong
    Jilin Daxue Xuebao (Gongxueban)/Journal of Jilin University (Engineering and Technology Edition), 2021, 51 (03): : 1011 - 1016
  • [35] Improved Character-Based Neural Network for POS Tagging on Morphologically Rich Languages
    Ali, Samat
    Murat, Alim
    JOURNAL OF INFORMATION PROCESSING SYSTEMS, 2023, 19 (03): : 355 - 369
  • [36] On predictive modeling of the twitter-based sales data using anew probabilistic model and machine learning methods
    Wan, Min
    Alshahrani, Mohammed A.
    Aloraini, Najla M.
    Alkhathami, Alia A.
    Alqahtani, Haifa
    ALEXANDRIA ENGINEERING JOURNAL, 2025, 113 : 661 - 671
  • [37] Intelligent Recognition English Translation Model Based on Embedded Machine Learning and Improved GLR Algorithm
    Lei, Lei
    MOBILE INFORMATION SYSTEMS, 2022, 2022
  • [38] Construction of Primary and Secondary School Teachers' Competency Model Based on Improved Machine Learning Algorithm
    Wu, JunNa
    MATHEMATICAL PROBLEMS IN ENGINEERING, 2022, 2022
  • [39] A CFD Data-based Cavity Flow Surrogate Model with Machine Learning Algorithm
    Chung, Lee Shing
    Boon, Lua Kim
    JOURNAL OF AERONAUTICS ASTRONAUTICS AND AVIATION, 2023, 55 (02): : 215 - 230
  • [40] Predicting Personality with Twitter Data and Machine Learning Models
    Ergu, Izel
    Isik, Zerrin
    Yankayis, Ismail
    2019 INNOVATIONS IN INTELLIGENT SYSTEMS AND APPLICATIONS CONFERENCE (ASYU), 2019, : 386 - 390