Personality Classification from Online Text using Machine Learning Approach

被引:1
作者
Khan, Alam Sher [1 ]
Ahmad, Hussain [1 ]
Asghar, Muhammad Zubair [1 ]
Saddozai, Furcian Khan [1 ]
Arir, Areeba [1 ]
Khalid, Hassan Ali [1 ]
机构
[1] Gomal Univ, Inst Comp & Informat Technol, Dera Ismail Khan, Pakistan
关键词
Personality recognition; re-sampling; machine learning; XGBoost; class imbalanced; MBTI; social networks; SOCIAL MEDIA;
D O I
暂无
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Personality refer to the distinctive set of characteristics of a person that effect their habits, behaviour's, attitude and pattern of thoughts. Text available on Social Networking sites provide an opportunity to recognize individual's personality traits automatically. In this proposed work, Machine Learning Technique, XGBoost classifier is used to predict four personality traits based on Myers- Briggs Type Indicator (MBTI) model, namely Introversion-Extroversion(I-E), iNtuition-Sensing(N-S), Feeling-Thinking(F-T) and Judging-Perceiving(J-P) from input text. Publically available benchmark dataset from Kaggle is used in experiments. The skewness of the dataset is the main issue associated with the prior work, which is minimized by applying Re-sampling technique namely random over-sampling, resulting in better performance. For more exploration of the personality from text, pre-processing techniques including tokenization, word stemming, stop words elimination and feature selection using TF IDF are also exploited. This work provides the basis for developing a personality identification system which could assist organization for recruiting and selecting appropriate personnel and to improve their business by knowing the personality and preferences of their customers. The results obtained by all classifiers across all personality traits is good enough, however, the performance of XGBoost classifier is outstanding by achieving more than 99% precision and accuracy for different traits.
引用
收藏
页码:460 / 476
页数:17
相关论文
共 41 条
  • [1] Estimating Personality from Social Media Posts
    Alsadhan, N.
    Skillicorn, D. B.
    [J]. 2017 17TH IEEE INTERNATIONAL CONFERENCE ON DATA MINING WORKSHOPS (ICDMW 2017), 2017, : 350 - 356
  • [2] [Anonymous], 2018, International Research Journal of Engineering and Technology
  • [3] [Anonymous], 2012, 6 INT C DIG SOC NETW
  • [4] [Anonymous], 2013, WCPR ICWSM 13
  • [5] Arnoux Pierre-Hadrien, 2017, P INT AAAI C WEB SOC
  • [6] Arroju M., 2015, 6 C LABS EVALUATION, P23
  • [7] RIFT: A Rule Induction Framework for Twitter Sentiment Analysis
    Asghar, Muhammad Zubair
    Khan, Aurangzeb
    Khan, Furqan
    Kundi, Fazal Masud
    [J]. ARABIAN JOURNAL FOR SCIENCE AND ENGINEERING, 2018, 43 (02) : 857 - 877
  • [8] Bharadwaj S, 2018, 2018 INTERNATIONAL CONFERENCE ON ADVANCES IN COMPUTING, COMMUNICATIONS AND INFORMATICS (ICACCI), P1076, DOI 10.1109/ICACCI.2018.8554828
  • [9] Buraya K, 2017, AAAI CONF ARTIF INTE, P4909
  • [10] Cantador I., 2013, CEUR WORKSHOP P