Quantile-Composited Feature Screening for Ultrahigh-Dimensional Data

被引:0
作者
Chen, Shuaishuai [1 ]
Lu, Jun [2 ]
机构
[1] Shandong Univ, Sch Math, Jinan 250100, Peoples R China
[2] Natl Univ Def & Technol, Sch Sci, Changsha 410000, Peoples R China
基金
中国国家自然科学基金;
关键词
feature screening; discriminative analysis; quantile-composited; CLASSIFICATION;
D O I
10.3390/math11102398
中图分类号
O1 [数学];
学科分类号
0701 ; 070101 ;
摘要
Ultrahigh-dimensional grouped data are frequently encountered by biostatisticians working on multi-class categorical problems. To rapidly screen out the null predictors, this paper proposes a quantile-composited feature screening procedure. The new method first transforms the continuous predictor to a Bernoulli variable, by thresholding the predictor at a certain quantile. Consequently, the independence between the response and each predictor is easy to judge, by employing the Pearson chi-square statistic. The newly proposed method has the following salient features: (1) it is robust against high-dimensional heterogeneous data; (2) it is model-free, without specifying any regression structure between the covariate and outcome variable; (3) it enjoys a low computational cost, with the computational complexity controlled at the sample size level. Under some mild conditions, the new method was shown to achieve the sure screening property without imposing any moment condition on the predictors. Numerical studies and real data analyses further confirmed the effectiveness of the new screening procedure.
引用
收藏
页数:21
相关论文
共 26 条
  • [11] Feature Screening for Ultrahigh Dimensional Categorical Data With Applications
    Huang, Danyang
    Li, Runze
    Wang, Hansheng
    [J]. JOURNAL OF BUSINESS & ECONOMIC STATISTICS, 2014, 32 (02) : 237 - 244
  • [12] Feature Screening via Distance Correlation Learning
    Li, Runze
    Zhong, Wei
    Zhu, Liping
    [J]. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2012, 107 (499) : 1129 - 1139
  • [13] Nonparametric feature screening
    Lin, Lu
    Sun, Jing
    Zhu, Lixing
    [J]. COMPUTATIONAL STATISTICS & DATA ANALYSIS, 2013, 67 : 162 - 174
  • [14] Feature Selection for Varying Coefficient Models With Ultrahigh-Dimensional Covariates
    Liu, Jingyuan
    Li, Runze
    Wu, Rongling
    [J]. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2014, 109 (505) : 266 - 274
  • [15] Model-free conditional screening via conditional distance correlation
    Lu, Jun
    Lin, Lu
    [J]. STATISTICAL PAPERS, 2020, 61 (01) : 225 - 244
  • [16] THE FUSED KOLMOGOROV FILTER: A NONPARAMETRIC MODEL-FREE SCREENING METHOD
    Mai, Qing
    Zou, Hui
    [J]. ANNALS OF STATISTICS, 2015, 43 (04) : 1471 - 1497
  • [17] HIGH-DIMENSIONAL ADDITIVE MODELING
    Meier, Lukas
    van de Geer, Sara
    Buehlmann, Peter
    [J]. ANNALS OF STATISTICS, 2009, 37 (6B) : 3779 - 3821
  • [18] Ultrahigh-Dimensional Multiclass Linear Discriminant Analysis by Pairwise Sure Independence Screening
    Pan, Rui
    Wang, Hansheng
    Li, Runze
    [J]. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2016, 111 (513) : 169 - 179
  • [19] Shao J., 2003, Mathematical Statistics
  • [20] Model-Free Conditional Feature Screening with FDR Control
    Tong, Zhaoxue
    Cai, Zhanrui
    Yang, Songshan
    Li, Runze
    [J]. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2023, 118 (544) : 2575 - 2587