Identifying Chinese Microblog Users With High Suicide Probability Using Internet-Based Profile and Linguistic Features: Classification Model

被引:56
作者
Guan, Li [1 ,2 ]
Hao, Bibo [2 ]
Cheng, Qijin [3 ]
Yip, Paul S. F. [3 ]
Zhu, Tingshao [1 ,4 ]
机构
[1] Chinese Acad Sci, Inst Psychol, Key Lab Behav Sci, Room 821,Bldg He Xie,16th Lincui Rd, Beijing 100101, Peoples R China
[2] Univ Chinese Acad Sci, Beijing, Peoples R China
[3] Univ Hong Kong, HKJC Ctr Suicide Res & Prevent, Hong Kong, Hong Kong, Peoples R China
[4] Chinese Acad Sci, Inst Comp Technol, Key Lab Intelligent Informat Proc, Beijing, Peoples R China
关键词
suicide probability; microblog; Chinese; classification model;
D O I
10.2196/mental.4227
中图分类号
R749 [精神病学];
学科分类号
100205 ;
摘要
Background: Traditional offline assessment of suicide probability is time consuming and difficult in convincing at-risk individuals to participate. Identifying individuals with high suicide probability through online social media has an advantage in its efficiency and potential to reach out to hidden individuals, yet little research has been focused on this specific field. Objective: The objective of this study was to apply two classification models, Simple Logistic Regression (SLR) and Random Forest (RF), to examine the feasibility and effectiveness of identifying high suicide possibility microblog users in China through profile and linguistic features extracted from Internet-based data. Methods: There were nine hundred and nine Chinese microblog users that completed an Internet survey, and those scoring one SD above the mean of the total Suicide Probability Scale (SPS) score, as well as one SD above the mean in each of the four subscale scores in the participant sample were labeled as high-risk individuals, respectively. Profile and linguistic features were fed into two machine learning algorithms (SLR and RF) to train the model that aims to identify high-risk individuals in general suicide probability and in its four dimensions. Models were trained and then tested by 5-fold cross validation; in which both training set and test set were generated under the stratified random sampling rule from the whole sample. There were three classic performance metrics (Precision, Recall, F1 measure) and a specifically defined metric "Screening Efficiency" that were adopted to evaluate model effectiveness. Results: Classification performance was generally matched between SLR and RF. Given the best performance of the classification models, we were able to retrieve over 70% of the labeled high-risk individuals in overall suicide probability as well as in the four dimensions. Screening Efficiency of most models varied from 1/4 to 1/2. Precision of the models was generally below 30%. Conclusions: Individuals in China with high suicide probability are recognizable by profile and text-based information from microblogs. Although there is still much space to improve the performance of classification models in the future, this study may shed light on preliminary screening of risky individuals via machine learning algorithms, which can work side-by-side with expert scrutiny to increase efficiency in large-scale-surveillance of suicide probability from online social media.
引用
收藏
页数:10
相关论文
共 47 条
[1]  
[Anonymous], 2013, ICWSM, DOI 10.1109/isads.2017.41
[2]  
[Anonymous], 2014, 2014 SINA WEIBO USER
[3]   Factors associated with suicidal ideation in an elderly urban Japanese population: A community-based, cross-sectional study [J].
Awata, S ;
Seki, T ;
Koizumi, Y ;
Sato, S ;
Hozawa, A ;
Omori, K ;
Kuriyama, S ;
Arai, H ;
Nagatomi, R ;
Matsuoka, H ;
Tsuji, I .
PSYCHIATRY AND CLINICAL NEUROSCIENCES, 2005, 59 (03) :327-336
[4]   Maternal Age at Child Birth, Birth Order, and Suicide at a Young Age: A Sibling Comparison [J].
Bjorngaard, Johan Hakon ;
Bjerkeset, Ottar ;
Vatten, Lars ;
Janszky, Imre ;
Gunnell, David ;
Romundstad, Pal .
AMERICAN JOURNAL OF EPIDEMIOLOGY, 2013, 177 (07) :638-644
[5]   Risk factors for the incidence and persistence of suicide-related outcomes: A 10-year follow-up study using the National Comorbidity Surveys [J].
Borges, Guilhertne ;
Angst, Jules ;
Nock, Matthew K. ;
Meron Ruscio, Ayelet ;
Kessler, Ronald C. .
JOURNAL OF AFFECTIVE DISORDERS, 2008, 105 (1-3) :25-33
[6]  
Chen P, 2013, INT C ADV INF ENG ED
[7]   Opportunities and challenges of online data collection for suicide prevention [J].
Cheng, Qijin ;
Chang, Shu-Sen ;
Yip, Paul S. F. .
LANCET, 2012, 379 (9830) :E53-E54
[8]   The association of irritability and impulsivity with suicidal ideation among 15-to 20-year-old males [J].
Conner, KR ;
Meldrum, S ;
Wieczorek, WF ;
Duberstein, PR ;
Welte, JW .
SUICIDE AND LIFE-THREATENING BEHAVIOR, 2004, 34 (04) :363-373
[9]   Applying Computer Adaptive Testing to Optimize Online Assessment of Suicidal Behavior: A Simulation Study [J].
De Beurs, Derek Paul ;
de Vries, Anton L. M. ;
de Groot, Marieke H. ;
de Keijser, Jos ;
Kerkhof, Ad J. F. M. .
JOURNAL OF MEDICAL INTERNET RESEARCH, 2014, 16 (09)
[10]  
De Chaudhury M, 2013, P SIGCHI C HUM FACT, DOI [DOI 10.1145/2470654.2466447, 10.1145/2470654.2466447]