CampER: An Effective Framework for Privacy-Aware Deep Entity Resolution

被引:4
作者
Guo, Yuxiang [1 ]
Chen, Lu [1 ]
Zhou, Zhengjie [2 ]
Zheng, Baihua [3 ]
Fang, Ziquan [1 ]
Zhang, Zhikun [4 ]
Mao, Yuren [2 ]
Gao, Yunjun [1 ]
机构
[1] Zhejiang Univ, Hangzhou, Peoples R China
[2] Zhejiang Univ, Ningbo, Peoples R China
[3] Singapore Management Univ, Singapore, Singapore
[4] Stanford Univ, Palo Alto, CA 94304 USA
来源
PROCEEDINGS OF THE 29TH ACM SIGKDD CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING, KDD 2023 | 2023年
关键词
entity resolution; representation learning; similarity measurement; LINKAGE;
D O I
10.1145/3580305.3599266
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Entity Resolution (ER) is a fundamental problem in data preparation. Standard deep ER methods have achieved state-of-the-art effectiveness, assuming that relations from different organizations are centrally stored. However, due to privacy concerns, it can be difficult to centralize data in practice, rendering standard deep ER solutions inapplicable. Despite efforts to develop rule-based privacy-preserving ER methods, they often neglect subtle matching mechanisms and have poor effectiveness as a result. To bridge effectiveness and privacy, in this paper, we propose CampER, an effective framework for privacy-aware deep entity resolution. Specifically, we first design a training pair self-generation strategy to overcome the absence of manually labeled data in privacy-aware scenarios. Based on the self-constructed training pairs, we present a collaborative fine-tuning approach to learn the match-aware and uni-space individual tuple embeddings for accurate matching decisions. During the matching decision-making process, we first introduce a cryptographically secure approach to determine matches. Furthermore, we propose an order-preserving perturbation strategy to significantly accelerate the matching computation while guaranteeing the consistency of ER results. Extensive experiments on eight widely-used benchmark datasets demonstrate that CampER not only is comparable with the state-of-the-art standard deep ER solutions in effectiveness, but also preserves privacy.
引用
收藏
页码:626 / 637
页数:12
相关论文
共 39 条
[1]   Hybrid framework of differential privacy and secure multi-party computation for privacy-preserving entity resolution [J].
Dorgbefu, Maxwell ;
Missah, Yaw Marfo ;
Ussiph, Najim ;
Abdul-Salaam, Gaddafi ;
Kornyo, Oliver ;
Mensah, Joseph Mawulorm .
COMPUTERS & SECURITY, 2025, 157
[2]   Synthesizing Privacy Preserving Entity Resolution Datasets [J].
Qinl, Xuedi ;
Chai, Chengliang ;
Tang, Nan ;
Li, Jian ;
Luo, Yuyu ;
Li, Guoliang ;
Zhu, Yaoyu .
2022 IEEE 38TH INTERNATIONAL CONFERENCE ON DATA ENGINEERING (ICDE 2022), 2022, :2359-2371
[3]   Semantic-Aware Blocking for Entity Resolution [J].
Wang, Qing ;
Cui, Mingyuan ;
Liang, Huizhi .
IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2016, 28 (01) :166-180
[4]   Effective Explanations for Entity Resolution Models [J].
Teofili, Tommaso ;
Firmani, Donatella ;
Koudas, Nick ;
Martello, Vincenzo ;
Merialdo, Paolo ;
Srivastava, Divesh .
2022 IEEE 38TH INTERNATIONAL CONFERENCE ON DATA ENGINEERING (ICDE 2022), 2022, :2709-2721
[5]   Deep and Collective Entity Resolution in Parallel [J].
Deng, Ting ;
Fan, Wenfei ;
Lu, Ping ;
Luo, Xiaomeng ;
Zhu, Xiaoke ;
An, Wanhe .
2022 IEEE 38TH INTERNATIONAL CONFERENCE ON DATA ENGINEERING (ICDE 2022), 2022, :2060-2072
[6]   MFIBlocks: An effective blocking algorithm for entity resolution [J].
Kenig, Batya ;
Gal, Avigdor .
INFORMATION SYSTEMS, 2013, 38 (06) :908-926
[7]   Effective entity resolution in product review domain [J].
Liu, Jing-Jing ;
Cao, Yun-Bo ;
Huang, Ya-Lou .
PROCEEDINGS OF 2007 INTERNATIONAL CONFERENCE ON MACHINE LEARNING AND CYBERNETICS, VOLS 1-7, 2007, :106-+
[8]   Deep Sequence-to-Sequence Entity Matching for Heterogeneous Entity Resolution [J].
Nie, Hao ;
Han, Xianpei ;
He, Ben ;
Sun, Le ;
Chen, Bo ;
Zhang, Wei ;
Wu, Suhui ;
Kong, Hao .
PROCEEDINGS OF THE 28TH ACM INTERNATIONAL CONFERENCE ON INFORMATION & KNOWLEDGE MANAGEMENT (CIKM '19), 2019, :629-638
[9]   DeepBlock: A Novel Blocking Approach for Entity Resolution using Deep Learning [J].
Javdani, Delaram ;
Rahmani, Hossein ;
Allahgholi, Milad ;
Karimkhani, Fatemeh .
2019 5TH INTERNATIONAL CONFERENCE ON WEB RESEARCH (ICWR), 2019, :41-44
[10]   Scaling entity resolution: A loosely schema-aware approach [J].
Simonini, Giovanni ;
Gagliardelli, Luca ;
Bergamaschi, Sonia ;
Jagadish, H., V .
INFORMATION SYSTEMS, 2019, 83 :145-165