Evaluating the Representativeness of Socio-Demographic Variables over Time for Geo-Social Media Data

被引:4
作者
Petutschnig, Andreas [1 ]
Resch, Bernd [1 ]
Lang, Stefan [1 ]
Havas, Clemens [1 ]
机构
[1] Univ Salzburg, Dept Geoinformat Z GIS, A-5020 Salzburg, Austria
基金
奥地利科学基金会;
关键词
geo-social media; Twitter; representativeness; spatial analysis; statistical correlations; temporal snapshots; SPATIAL BIAS; TWITTER; INFORMATION; SELECTION; PATTERNS; MODELS;
D O I
10.3390/ijgi10050323
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Geo-social media data are widely used as a data source to model populations and processes in a variety of contexts. However, if the data do not adequately represent the population they are drawn from, analysis results will be biased. Unaddressed, these biases may lead to false interpretations and conclusions. In this paper, we propose a generic methodology for investigating the representativeness of geo-social media data for population groups of similar statistical predictive power based on reference data. The groups are designed to be spatially coherent regions with similar prediction errors. Based on these units, we investigate the influence of different socio-demographic covariates on the representativeness. We perform experiments based on over 1.6 billion tweets and 90 socio-demographic covariates. We demonstrate that Twitter data representativeness varies strongly over time and space. Our results show that densely populated areas tend to be underrepresented consistently in non-spatial models. Over time, some covariates like the number of people aged 20 years exhibit highly different effects on the prediction models, whereas others are much more stable. The spatial effects can most frequently be explained using spatial error models, indicating spatially related errors that indicate the necessity of additional covariates. Finally, we provide hints for interpreting the results of our approach for researchers using the concepts presented in this paper.
引用
收藏
页数:18
相关论文
共 54 条