Cross-validation on extreme regions

被引:0
|
作者
Aghbalou, Anass [1 ]
Bertail, Patrice [2 ]
Portier, Francois [3 ]
Sabourin, Anne [4 ]
机构
[1] Inst Polytech Paris, Telecom Paris, LTCI, Palaiseau, France
[2] Univ Paris Nanterre, MODALX, Nanterre, France
[3] CREST, Ensai, Rennes, France
[4] Univ Paris, MAP5, CNRS, F-75006 Paris, France
关键词
Extreme value analysis; Cross-validation; Concentration inequalities; REGULAR VARIATION; TAIL DEPENDENCE; M-ESTIMATOR; CONSISTENCY; BOUNDS; CLASSIFICATION; STABILITY;
D O I
10.1007/s10687-024-00495-z
中图分类号
O1 [数学];
学科分类号
0701 ; 070101 ;
摘要
We conduct a non-asymptotic study of the Cross-Validation (CV) estimate of the generalization risk for learning algorithms dedicated to extreme regions of the covariates space. In this context which has recently been analysed from an Extreme Value Analysis perspective, the risk function measures the algorithm's error given that the norm of the input exceeds a high quantile. The main challenge within this framework is the negligible size of the extreme training sample with respect to the full sample size and the necessity to re-scale the risk function by a probability tending to zero. We open the road to a finite sample understanding of CV for extreme values by establishing two new results: an exponential probability bound on the K-fold CV error and a polynomial probability bound on the leave-p-out CV. Our bounds are sharp in the sense that they match state-of-the-art guarantees for standard CV estimates while extending them to encompass a conditioning event of small probability. We illustrate the significance of our results regarding high dimensional classification in extreme regions via a Lasso-type logistic regression algorithm. The tightness of our bounds is investigated in numerical experiments.
引用
收藏
页码:505 / 555
页数:51
相关论文
共 50 条
  • [21] A THEORY OF CROSS-VALIDATION ERROR
    TURNEY, P
    JOURNAL OF EXPERIMENTAL & THEORETICAL ARTIFICIAL INTELLIGENCE, 1994, 6 (04) : 361 - 391
  • [22] Cross-validation and median criterion
    Zheng, ZG
    Yang, Y
    STATISTICA SINICA, 1998, 8 (03) : 907 - 921
  • [23] Asymptotic properties of adaptive likelihood weights by cross-validation
    Wang, Xiaogang
    COMMUNICATIONS IN STATISTICS-THEORY AND METHODS, 2006, 35 (07) : 1257 - 1270
  • [24] Dynamic weighting ensemble classifiers based on cross-validation
    Zhu Yu-Quan
    Ou Ji-Shun
    Chen Geng
    Yu Hai-Ping
    NEURAL COMPUTING & APPLICATIONS, 2011, 20 (03): : 309 - 317
  • [25] Dynamic weighting ensemble classifiers based on cross-validation
    Zhu Yu-Quan
    Ou Ji-Shun
    Chen Geng
    Yu Hai-Ping
    Neural Computing and Applications, 2011, 20 : 309 - 317
  • [26] ON THE CONSISTENCY OF CROSS-VALIDATION IN NONLINEAR WAVELET REGRESSION ESTIMATION
    张双林
    郑忠国
    Acta Mathematica Scientia, 2000, (01) : 1 - 11
  • [27] Accelerating Cross-Validation in Multinomial Logistic Regression with l1-Regularization
    Obuchi, Tomoyuki
    Kabashima, Yoshiyuki
    JOURNAL OF MACHINE LEARNING RESEARCH, 2018, 19
  • [28] Fast Cross-Validation for Kernel-Based Algorithms
    Liu, Yong
    Liao, Shizhong
    Jiang, Shali
    Ding, Lizhong
    Lin, Hailun
    Wang, Weiping
    IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2020, 42 (05) : 1083 - 1096
  • [29] On the consistency of cross-validation in nonlinear wavelet regression estimation
    Zhang, SL
    Zheng, ZG
    ACTA MATHEMATICA SCIENTIA, 2000, 20 (01) : 1 - 11
  • [30] Concentration inequalities for cross-validation in scattered data approximation
    Bartel, Felix
    Hielscher, Ralf
    JOURNAL OF APPROXIMATION THEORY, 2022, 277