Model-free feature screening for high-dimensional survival data

被引:10
|
作者
Lin, Yuanyuan [1 ]
Liu, Xianhui [2 ,3 ]
Hao, Meiling [4 ]
机构
[1] Chinese Univ Hong Kong, Dept Stat, Hong Kong 999077, Hong Kong, Peoples R China
[2] Jiangxi Univ Finance & Econ, Sch Stat, Nanchang 330013, Jiangxi, Peoples R China
[3] Jiangxi Univ Finance & Econ, Res Ctr Appl Stat, Nanchang 330013, Jiangxi, Peoples R China
[4] Univ Hlth Network, Princess Margaret Canc Ctr, Toronto, ON M5G 2M9, Canada
基金
加拿大健康研究院; 中国国家自然科学基金;
关键词
feature screening; random censoring; robustness; sure independence screening; ultra-high dimension; PROPORTIONAL HAZARDS MODEL; VARIABLE SELECTION; HETEROGENEOUS DATA; NP-DIMENSIONALITY; ORACLE PROPERTIES; ADAPTIVE LASSO; COX MODEL; INEQUALITIES; REGRESSION;
D O I
10.1007/s11425-016-9116-6
中图分类号
O29 [应用数学];
学科分类号
070104 ;
摘要
With the rapid-growth-in-size scientific data in various disciplines, feature screening plays an important role to reduce the high-dimensionality to a moderate scale in many scientific fields. In this paper, we introduce a unified and robust model-free feature screening approach for high-dimensional survival data with censoring, which has several advantages: it is a model-free approach under a general model framework, and hence avoids the complication to specify an actual model form with huge number of candidate variables; under mild conditions without requiring the existence of any moment of the response, it enjoys the ranking consistency and sure screening properties in ultra-high dimension. In particular, we impose a conditional independence assumption of the response and the censoring variable given each covariate, instead of assuming the censoring variable is independent of the response and the covariates. Moreover, we also propose a more robust variant to the new procedure, which possesses desirable theoretical properties without any finite moment condition of the predictors and the response. The computation of the newly proposed methods does not require any complicated numerical optimization and it is fast and easy to implement. Extensive numerical studies demonstrate that the proposed methods perform competitively for various configurations. Application is illustrated with an analysis of a genetic data set.
引用
收藏
页码:1617 / 1636
页数:20
相关论文
共 50 条