Disentangling User Samples: A Supervised Machine Learning Approach to Proxy-population Mismatch in Twitter Research

被引:14
作者
Kwon, K. Hazel [1 ]
Priniski, J. Hunter [2 ]
Chadha, Monica [1 ]
机构
[1] Arizona State Univ, Walter Cronkite Sch Journalism & Mass Commun, Phoenix, AZ USA
[2] Arizona State Univ, Dept Math & Stat Sci, Phoenix, AZ USA
关键词
SOCIAL MEDIA; NETWORK; MESSAGES; TWEET; NEWS;
D O I
10.1080/19312458.2018.1430755
中图分类号
G2 [信息与知识传播];
学科分类号
05 ; 0503 ;
摘要
This study addresses the issue of sampling biases in social media data-driven communication research. The authors demonstrate how supervised machine learning could reduce Twitter sampling bias induced from "proxy-population mismatch". Particularly, this study used the Random Forest (RF) classifier to disentangle tweet samples representative of general publics' activities from non-general-or institutional-activities. By applying RF classifier models to Twitter data sets relevant to four news events and a randomly pooled dataset, the study finds systematic differences between general user samples and institutional user samples in their messaging patterns. This article calls for disentangling Twitter user samples when ordinary user behaviors are the focus of research. It also builds on the development of machine learning modeling in the context of communication research.
引用
收藏
页码:216 / 237
页数:22
相关论文
共 39 条
  • [1] Dissecting a Social Botnet: Growth, Content and Influence in Twitter
    Abokhodair, Norah
    Yoo, Daisy
    McDonald, David W.
    [J]. PROCEEDINGS OF THE 2015 ACM INTERNATIONAL CONFERENCE ON COMPUTER-SUPPORTED COOPERATIVE WORK AND SOCIAL COMPUTING (CSCW'15), 2015, : 839 - 851
  • [2] [Anonymous], 2015, LINGUISTIC INQUIRY A
  • [3] [Anonymous], 2017, ZOOL J LINN SOC-LOND, DOI [10.1093/gbe/evz266, DOI 10.1093/ZOOLINNEAN/ZLX067]
  • [4] [Anonymous], ARXIV13065204PHYSICS
  • [5] [Anonymous], 2011, Fifth International AAAI Conference on Weblogs and Social Media, DOI 10.1609/icwsm.v5i1.14127
  • [6] [Anonymous], 2017, BIT BIT SOCIAL RES D
  • [7] [Anonymous], 2011, 11114503 ARXIV
  • [8] CRITICAL QUESTIONS FOR BIG DATA Provocations for a cultural, technological, and scholarly phenomenon
    Boyd, Danah
    Crawford, Kate
    [J]. INFORMATION COMMUNICATION & SOCIETY, 2012, 15 (05) : 662 - 679
  • [9] Breiman L., 1983, OLSHEN STONE CLASSIF
  • [10] Random forests
    Breiman, L
    [J]. MACHINE LEARNING, 2001, 45 (01) : 5 - 32