Deep neural networks are central to recent breakthroughs in artificial intelligence, yet we still lack a full understanding of their theoretical foundations. For instance, in contrast with classical statistical learning theory, highly overparameterized networks can perfectly interpolate training data achieving zero training error while implicitly selecting solutions that generalize exceptionally well. This phenomenon, often conceptualized through the paradigms of ``double descent'' and ``benign overfitting'', has been shown to emerge in a broader class of analytically tractable learning tasks, such as Kernel Ridge Regression (KRR). Consequently, studying these proxy models in depth provides crucial theoretical insights into the fundamental mechanisms of modern machine learning. In this work, we investigate the impact of realistic data distributions on learning dynamics within a KRR framework in the polynomial scaling regime, i.e., when both the number of samples and the data dimension diverge to infinity following a polynomial relationship. Specifically, we analyze how the input data distribution shapes the spectral properties of the kernel operator, which directly govern the model's generalization capabilities through the deterministic equivalents for the bias and variance. Initially, we examine the kernel eigenvalue distribution under the assumption that training data are drawn from an isotropic multivariate Gaussian. We then extend our analysis to power-law distributed covariate structures, investigating how anisotropy in the training set alters the learning dynamics. Within this regime, we derive exact analytical learning curves for both the variance and bias terms, precisely capturing key phenomena such as interpolation peaks, multiple descent behavior, and scaling laws.
Sharp description of Kernel Ridge Regression under power-law data
RIZZI, LORENZO
2025/2026
Abstract
Deep neural networks are central to recent breakthroughs in artificial intelligence, yet we still lack a full understanding of their theoretical foundations. For instance, in contrast with classical statistical learning theory, highly overparameterized networks can perfectly interpolate training data achieving zero training error while implicitly selecting solutions that generalize exceptionally well. This phenomenon, often conceptualized through the paradigms of ``double descent'' and ``benign overfitting'', has been shown to emerge in a broader class of analytically tractable learning tasks, such as Kernel Ridge Regression (KRR). Consequently, studying these proxy models in depth provides crucial theoretical insights into the fundamental mechanisms of modern machine learning. In this work, we investigate the impact of realistic data distributions on learning dynamics within a KRR framework in the polynomial scaling regime, i.e., when both the number of samples and the data dimension diverge to infinity following a polynomial relationship. Specifically, we analyze how the input data distribution shapes the spectral properties of the kernel operator, which directly govern the model's generalization capabilities through the deterministic equivalents for the bias and variance. Initially, we examine the kernel eigenvalue distribution under the assumption that training data are drawn from an isotropic multivariate Gaussian. We then extend our analysis to power-law distributed covariate structures, investigating how anisotropy in the training set alters the learning dynamics. Within this regime, we derive exact analytical learning curves for both the variance and bias terms, precisely capturing key phenomena such as interpolation peaks, multiple descent behavior, and scaling laws.| File | Dimensione | Formato | |
|---|---|---|---|
|
Rizzi_Lorenzo.pdf
accesso aperto
Dimensione
1.8 MB
Formato
Adobe PDF
|
1.8 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/113159