What can be learnt with wide convolutional neural networks?

被引:0
作者
Cagnetta, Francesco [1 ]
Favero, Alessandro [1 ,2 ]
Wyart, Matthieu [1 ]
机构
[1] Ecole Polytech Fed Lausanne EPFL, Inst Phys, Lausanne, Switzerland
[2] Ecole Polytech Fed Lausanne EPFL, Inst Elect Engn, Lausanne, Switzerland
来源
JOURNAL OF STATISTICAL MECHANICS-THEORY AND EXPERIMENT | 2024年 / 2024卷 / 10期
关键词
machine learning; ICML; deep learning; NTK; convolutional neural networks; CNNs; Kernel methods; generalisation; curse of dimensionality; locality;
D O I
10.1088/1742-5468/ad65df
中图分类号
O3 [力学];
学科分类号
08 ; 0801 ;
摘要
Understanding how convolutional neural networks (CNNs) can efficiently learn high-dimensional functions remains a fundamental challenge. A popular belief is that these models harness the local and hierarchical structure of natural data such as images. Yet, we lack a quantitative understanding of how such structure affects performance, for example the rate of decay of the generalisation error with the number of training samples. In this paper, we study infinitely wide deep CNNs in the kernel regime. First, we show that the spectrum of the corresponding kernel inherits the hierarchical structure of the network, and we characterise its asymptotics. Then, we use this result together with generalisation bounds to prove that deep CNNs adapt to the spatial scale of the target function. In particular, we find that if the target function depends on low-dimensional subsets of adjacent input variables then the decay of the error is controlled by the effective dimensionality of these subsets. Conversely, if the target function depends on the full set of input variables then the error decay is controlled by the input dimension. We conclude by computing the generalisation error of a deep CNN trained on the output of another deep CNN with randomly initialised parameters. Interestingly, we find that, despite their hierarchical structure, the functions generated by infinitely wide deep CNNs are too rich to be efficiently learnable in high dimensions.
引用
收藏
页数:47
相关论文
共 62 条
[1]  
Abbe E, 2022, PR MACH LEARN RES, V178
[2]  
Arora S., 2019, Advances in Neural Information Processing Systems, P32
[3]   Spherical Harmonics and Approximations on the Unit Sphere: An Introduction Preface [J].
Atkinson, Kendall ;
Han, Weimin .
SPHERICAL HARMONICS AND APPROXIMATIONS ON THE UNIT SPHERE: AN INTRODUCTION, 2012, 2044 :V-+
[4]  
Azevedo D., 2015, Proc. Ser. Braz. Soc. Comput. Appl. Math, V3, P39, DOI [10.5540/03.2015.003.01.0039, DOI 10.5540/03.2015.003.01.0039]
[5]  
Bach F., 2024, The MIT Press Learning theory from first principles
[6]  
Bach F, 2017, J MACH LEARN RES, V18
[7]   RECOGNITION-BY-COMPONENTS - A THEORY OF HUMAN IMAGE UNDERSTANDING [J].
BIEDERMAN, I .
PSYCHOLOGICAL REVIEW, 1987, 94 (02) :115-147
[8]  
Bietti A., 2021, Advances in Neural Information Processing Systems, Vvol 34, ppp 18673
[9]  
Bietti A., 2021, ICLR 2021 INT C LEAR, ppp 1
[10]  
Bietti A., 2019, ADV NEUR IN, pp 32