Efficient Model-Free Subsampling Method for Massive Data

被引:2
|
作者
Zhou, Zheng [1 ]
Yang, Zebin [2 ]
Zhang, Aijun [2 ]
Zhou, Yongdao [1 ,3 ]
机构
[1] Nankai Univ, Sch Stat & Data Sci, NITFID, Tianjin, Peoples R China
[2] Univ Hong Kong, Dept Stat & Actuarial Sci, Hong Kong, Peoples R China
[3] Nankai Univ, Sch Stat & Data Sci, NITFID, Tianjin 300071, Peoples R China
基金
中国国家自然科学基金;
关键词
Big data subsampling; Model robustness; Parallel computing; Uniform designs; VARIANCE TEST; DISCREPANCY;
D O I
10.1080/00401706.2023.2271091
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
Subsampling plays a crucial role in tackling problems associated with the storage and statistical learning of massive datasets. However, most existing subsampling methods are model-based, which means their performances can drop significantly when the underlying model is misspecified. Such an issue calls for model-free subsampling methods that are robust under diverse model specifications. Recently, several model-free subsampling methods have been developed. However, the computing time of these methods grows explosively with the sample size, making them impractical for handling massive data. In this article, an efficient model-free subsampling method is proposed, which segments the original data into some regular data blocks and obtains subsamples from each data block by the data-driven subsampling method. Compared with existing model-free subsampling methods, the proposed method has a significant speed advantage and performs more robustly for datasets with complex underlying distributions. As demonstrated in simulation experiments, the proposed method is an order of magnitude faster than other commonly used model-free subsampling methods when the sample size of the original dataset reaches the order of 107. Moreover, simulation experiments and case studies show that the proposed method is more robust than other model-free subsampling methods under diverse model specifications and subsample sizes.
引用
收藏
页码:240 / 252
页数:13
相关论文
共 50 条
  • [1] Model-free global likelihood subsampling for massive data
    Si-Yu Yi
    Yong-Dao Zhou
    Statistics and Computing, 2023, 33
  • [2] Model-free global likelihood subsampling for massive data
    Yi, Si-Yu
    Zhou, Yong-Dao
    STATISTICS AND COMPUTING, 2023, 33 (01)
  • [3] Model-Free Subsampling Method Based on Uniform Designs
    Zhang, Mei
    Zhou, Yongdao
    Zhou, Zheng
    Zhang, Aijun
    IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2024, 36 (03) : 1210 - 1220
  • [4] Efficient Model-free Anthropometry from Depth Data
    Probst, Thomas
    Fossati, Andrea
    Salzmann, Mathieu
    Van Gool, Luc
    PROCEEDINGS 2017 INTERNATIONAL CONFERENCE ON 3D VISION (3DV), 2017, : 486 - 495
  • [5] Robust and efficient subsampling algorithms for massive data logistic regression
    Jin, Jun
    Liu, Shuangzhe
    Ma, Tiefeng
    JOURNAL OF APPLIED STATISTICS, 2024, 51 (08) : 1427 - 1445
  • [6] Estimation and testing of expectile regression with efficient subsampling for massive data
    Chen, Baolin
    Song, Shanshan
    Zhou, Yong
    STATISTICAL PAPERS, 2024, 65 (09) : 5593 - 5613
  • [7] Efficient data structures for model-free data-driven computational mechanics
    Eggersmann, Robert
    Stainier, Laurent
    Ortiz, Michael
    Reese, Stefanie
    COMPUTER METHODS IN APPLIED MECHANICS AND ENGINEERING, 2021, 382
  • [8] MODEL-FREE DATA CONDENSATION
    FINCH, PD
    PHILOSOPHICAL TRANSACTIONS OF THE ROYAL SOCIETY OF LONDON SERIES A-MATHEMATICAL PHYSICAL AND ENGINEERING SCIENCES, 1991, 335 (1636): : 1 - 50
  • [9] An efficient model-free approach to interaction screening for high dimensional data
    Xiong, Wei
    Pan, Han
    Wang, Jianrong
    Tian, Maozai
    STATISTICS IN MEDICINE, 2023, 42 (10) : 1583 - 1605
  • [10] A novel residual subsampling method for skew-normal mode regression model with massive data
    Jiang, Zhe
    Wu, Yan
    Wang, Min
    Wu, Liucang
    COMMUNICATIONS IN STATISTICS-THEORY AND METHODS, 2024, 53 (16) : 5972 - 5988