Collecting High-Dimensional and Correlation-Constrained Data with Local Differential Privacy

被引:21
作者
Du, Rong [1 ]
Ye, Qingqing [1 ]
Fu, Yue [1 ]
Hu, Haibo [1 ]
机构
[1] Hong Kong Polytech Univ, Dept Elect & Informat Engn, Hong Kong, Peoples R China
来源
2021 18TH ANNUAL IEEE INTERNATIONAL CONFERENCE ON SENSING, COMMUNICATION, AND NETWORKING (SECON) | 2021年
基金
中国国家自然科学基金;
关键词
ANSWERING RANGE QUERIES; DATA PUBLICATION;
D O I
10.1109/SECON52354.2021.9491591
中图分类号
TP3 [计算技术、计算机技术];
学科分类号
0812 ;
摘要
Local differential privacy (LDP) is a promising privacy model for distributed data collection. It has been widely deployed in real-world systems (e.g. Chrome, iOS, macOS). In LDP-based mechanisms, an aggregator collects private values perturbed by each user and then analyses these values to estimate their statistics, such as frequency and mean. Most existing works focus on simple scalar value types, such as hoolean and categorical values. However, with the emergence of smart sensors and internet of things, high-dimensional data are gaining increasing popularity. In many cases, correlations exist between various attributes of such data, e.g. temperature and luminance. To ensure LDP for high-dimensional data, existing solutions either partition the privacy budget c among these correlated attributes or adopt sampling, both of which dilute the density of useful information and thus result in poor data utility. In this paper, we propose a relaxed LDP model, namely, univariate dominance local differential privacy (UDLDP), for high-dimensional data. We quantify the correlations between attributtes and present a correlation-bounded perturbation (CBP) mechanism that optimizes the partitioning of privacy budget on each correlated attribute. Furthermore, we extend CBP to support sampling, which is a common bandwidth reduction technique in sensor networks and Internet of Things. We derive the best allocation strategy of sampling probabilities among attributes in terms of data utility, which leads to the correlation-bounded perturbation mechanism with sampling (CBPS). The performance of both mechanisms is evaluated and compared with state-of-the-art LDP mechanisms on real-world and synthetic datasets.
引用
收藏
页数:9
相关论文
共 36 条
[1]  
[Anonymous], 2022, TKDE, DOI DOI 10.1109/TKDE.2020.3047124
[2]   Dimensionality reduction of medical big data using neural-fuzzy classifier [J].
Azar, Ahmad Taher ;
Hassanien, Aboul Ella .
SOFT COMPUTING, 2015, 19 (04) :1115-1127
[3]  
Bassily R., 2017, P ADV NEUR INF PROC, P2285
[4]   Local, Private, Efficient Protocols for Succinct Histograms [J].
Bassily, Raef ;
Smith, Adam .
STOC'15: PROCEEDINGS OF THE 2015 ACM SYMPOSIUM ON THEORY OF COMPUTING, 2015, :127-135
[5]  
Chen R, 2016, PROC INT CONF DATA, P289, DOI 10.1109/ICDE.2016.7498248
[6]   Correlated network data publication via differential privacy [J].
Chen, Rui ;
Fung, Benjamin C. M. ;
Yu, Philip S. ;
Desai, Bipin C. .
VLDB JOURNAL, 2014, 23 (04) :653-676
[7]   Answering Range Queries Under Local Differential Privacy [J].
Cormode, Graham ;
Kulkarni, Tejas ;
Srivastava, Divesh .
PROCEEDINGS OF THE VLDB ENDOWMENT, 2019, 12 (10) :1126-1138
[8]   ] Marginal Release Under Local Differential Privacy [J].
Cormode, Graham ;
Kulkarni, Tejas ;
Srivastava, Divesh .
SIGMOD'18: PROCEEDINGS OF THE 2018 INTERNATIONAL CONFERENCE ON MANAGEMENT OF DATA, 2018, :131-146
[9]   Privacy Aware Learning [J].
Duchi, John C. ;
Jordan, Michael I. ;
Wainwright, Martin J. .
JOURNAL OF THE ACM, 2014, 61 (06) :1-57
[10]   Local Privacy and Statistical Minimax Rates [J].
Duchi, John C. ;
Jordan, Michael I. ;
Wainwright, Martin J. .
2013 IEEE 54TH ANNUAL SYMPOSIUM ON FOUNDATIONS OF COMPUTER SCIENCE (FOCS), 2013, :429-438