Cross-validation on extreme regions

被引:0
|
作者
Aghbalou, Anass [1 ]
Bertail, Patrice [2 ]
Portier, Francois [3 ]
Sabourin, Anne [4 ]
机构
[1] Inst Polytech Paris, Telecom Paris, LTCI, Palaiseau, France
[2] Univ Paris Nanterre, MODALX, Nanterre, France
[3] CREST, Ensai, Rennes, France
[4] Univ Paris, MAP5, CNRS, F-75006 Paris, France
关键词
Extreme value analysis; Cross-validation; Concentration inequalities; REGULAR VARIATION; TAIL DEPENDENCE; M-ESTIMATOR; CONSISTENCY; BOUNDS; CLASSIFICATION; STABILITY;
D O I
10.1007/s10687-024-00495-z
中图分类号
O1 [数学];
学科分类号
0701 ; 070101 ;
摘要
We conduct a non-asymptotic study of the Cross-Validation (CV) estimate of the generalization risk for learning algorithms dedicated to extreme regions of the covariates space. In this context which has recently been analysed from an Extreme Value Analysis perspective, the risk function measures the algorithm's error given that the norm of the input exceeds a high quantile. The main challenge within this framework is the negligible size of the extreme training sample with respect to the full sample size and the necessity to re-scale the risk function by a probability tending to zero. We open the road to a finite sample understanding of CV for extreme values by establishing two new results: an exponential probability bound on the K-fold CV error and a polynomial probability bound on the leave-p-out CV. Our bounds are sharp in the sense that they match state-of-the-art guarantees for standard CV estimates while extending them to encompass a conditioning event of small probability. We illustrate the significance of our results regarding high dimensional classification in extreme regions via a Lasso-type logistic regression algorithm. The tightness of our bounds is investigated in numerical experiments.
引用
收藏
页码:505 / 555
页数:51
相关论文
共 50 条
  • [41] Cross-validation rules for order estimation
    Stoica, P
    Selén, Y
    DIGITAL SIGNAL PROCESSING, 2004, 14 (04) : 355 - 371
  • [42] Granularity selection for cross-validation of SVM
    Liu, Yong
    Liao, Shizhong
    INFORMATION SCIENCES, 2017, 378 : 475 - 483
  • [43] Estimation Stability With Cross-Validation (ESCV)
    Lim, Chinghway
    Yu, Bin
    JOURNAL OF COMPUTATIONAL AND GRAPHICAL STATISTICS, 2016, 25 (02) : 464 - 492
  • [44] Cross-validation in nonparametric regression with outliers
    Leung, DHY
    ANNALS OF STATISTICS, 2005, 33 (05) : 2291 - 2310
  • [45] LOCAL MINIMA IN CROSS-VALIDATION FUNCTIONS
    HALL, P
    MARRON, JS
    JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-METHODOLOGICAL, 1991, 53 (01): : 245 - 252
  • [46] Cross-validation of multiway component models
    Louwerse, DJ
    Smilde, AK
    Kiers, HAL
    JOURNAL OF CHEMOMETRICS, 1999, 13 (05) : 491 - 510
  • [47] Theoretical Analysis of Cross-Validation for Estimating the Risk of the k-Nearest Neighbor Classifier
    Celisse, Alain
    Mary-Huard, Tristan
    JOURNAL OF MACHINE LEARNING RESEARCH, 2018, 19
  • [48] Cross-validation for parameter estimation in the BEM
    Golberg, MA
    ENGINEERING ANALYSIS WITH BOUNDARY ELEMENTS, 1997, 19 (02) : 157 - 166
  • [49] Procrustes Cross-Validation-A Bridge between Cross-Validation and Independent Validation Sets
    Kucheryavskiy, Sergey
    Zhilin, Sergei
    Rodionova, Oxana
    Pomerantsev, Alexey
    ANALYTICAL CHEMISTRY, 2020, 92 (17) : 11842 - 11850
  • [50] Nonparametric Goodness of Fit via Cross-Validation Bayes Factors
    Hart, Jeffrey D.
    Choi, Taeryon
    BAYESIAN ANALYSIS, 2017, 12 (03): : 653 - 677