Improving Lasso for model selection and prediction

被引：3

作者：

Pokarowski, Piotr ^{[1
]}

Rejchel, Wojciech ^{[2
]}

Soltys, Agnieszka ^{[1
]}

Frej, Michal ^{[1
]}

Mielniczuk, Jan ^{[3
,4
]}

机构：

[1] Univ Warsaw, Inst Appl Math & Mech, Warsaw, Poland

[2] Nicolaus Copernicus Univ, Fac Math & Comp Sci, Chopina 12-18, PL-87100 Torun, Poland

[3] Polish Acad Sci, Inst Comp Sci, Warsaw, Poland

[4] Warsaw Univ Technol, Fac Math & Informat Sci, Warsaw, Poland

来源：

SCANDINAVIAN JOURNAL OF STATISTICS | 2022年 / 49卷 / 02期

关键词：

convex loss function; empirical process; generalized information criterion; high-dimensional regression; penalized estimation; selection consistency; NONCONVEX PENALIZED REGRESSION; VARIABLE SELECTION; CRITERIA; REGULARIZATION; LIKELIHOOD;

D O I：

10.1111/sjos.12546

中图分类号：

O21 [概率论与数理统计]; C8 [统计学];

学科分类号：

020208 ; 070103 ; 0714 ;

摘要：

It is known that the Thresholded Lasso (TL), SCAD or MCP correct intrinsic estimation bias of the Lasso. In this paper we propose an alternative method of improving the Lasso for predictive models with general convex loss functions which encompass normal linear models, logistic regression, quantile regression, or support vector machines. For a given penalty we order the absolute values of the Lasso nonzero coefficients and then select the final model from a small nested family by the Generalized Information Criterion. We derive exponential upper bounds on the selection error of the method. These results confirm that, at least for normal linear models, our algorithm seems to be the benchmark for the theory of model selection as it is constructive, computationally efficient and leads to consistent model selection under weak assumptions. Constructivity of the algorithm means that, in contrast to the TL, SCAD or MCP, consistent selection does not rely on the unknown parameters as the cone invertibility factor. Instead, our algorithm only needs the sample size, the number of predictors and an upper bound on the noise parameter. We show in numerical experiments on synthetic and real-world datasets that an implementation of our algorithm is more accurate than implementations of studied concave regularizations. Our procedure is included in the R package DMRnet and available in the CRAN repository.

引用

页码：831 / 863

页数：33

共 40 条

[1] [Anonymous], 2012, Electronic Communications in Probability, DOI DOI 10.1214/ECP.V17-2079
[2] Convexity, classification, and risk bounds
Bartlett, PL
Jordan, MI
McAuliffe, JD
[J]. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2006, 101 (473) : 138 - 156
[3] SIMULTANEOUS ANALYSIS OF LASSO AND DANTZIG SELECTOR
Bickel, Peter J.
Ritov, Ya'acov
Tsybakov, Alexandre B.
[J]. ANNALS OF STATISTICS, 2009, 37 (04) : 1705 - 1732
[4] COORDINATE DESCENT ALGORITHMS FOR NONCONVEX PENALIZED REGRESSION, WITH APPLICATIONS TO BIOLOGICAL FEATURE SELECTION
Breheny, Patrick
Huang, Jian
[J]. ANNALS OF APPLIED STATISTICS, 2011, 5 (01) : 232 - 253
[5] Bühlmann P, 2011, SPRINGER SER STAT, P1, DOI 10.1007/978-3-642-20192-9
[6] Least angle regression - Rejoinder
Efron, B
Hastie, T
Johnstone, I
Tibshirani, R
[J]. ANNALS OF STATISTICS, 2004, 32 (02) : 494 - 499
[7] STRONG ORACLE OPTIMALITY OF FOLDED CONCAVE PENALIZED ESTIMATION
Fan, Jianqing
Xue, Lingzhou
Zou, Hui
[J]. ANNALS OF STATISTICS, 2014, 42 (03) : 819 - 849
[8] Variable selection via nonconcave penalized likelihood and its oracle properties
Fan, JQ
Li, RZ
[J]. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2001, 96 (456) : 1348 - 1360
[9] Tuning parameter selection in high dimensional penalized likelihood
Fan, Yingying
Tang, Cheng Yong
[J]. JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY, 2013, 75 (03) : 531 - 552
[10] THE RISK INFLATION CRITERION FOR MULTIPLE-REGRESSION
FOSTER, DP
GEORGE, EI
[J]. ANNALS OF STATISTICS, 1994, 22 (04) : 1947 - 1975

← 1 2 3 4 →