Deep neural networks are central to recent breakthroughs in artificial intelligence, yet we still lack a full understanding of their theoretical foundations. For instance, in contrast with classical statistical learning theory, highly overparameterized networks can perfectly interpolate training data achieving zero training error while implicitly selecting solutions that generalize exceptionally well. This phenomenon, often conceptualized through the paradigms of ``double descent'' and ``benign overfitting'', has been shown to emerge in a broader class of analytically tractable learning tasks, such as Kernel Ridge Regression (KRR). Consequently, studying these proxy models in depth provides crucial theoretical insights into the fundamental mechanisms of modern machine learning. In this work, we investigate the impact of realistic data distributions on learning dynamics within a KRR framework in the polynomial scaling regime, i.e., when both the number of samples and the data dimension diverge to infinity following a polynomial relationship. Specifically, we analyze how the input data distribution shapes the spectral properties of the kernel operator, which directly govern the model's generalization capabilities through the deterministic equivalents for the bias and variance. Initially, we examine the kernel eigenvalue distribution under the assumption that training data are drawn from an isotropic multivariate Gaussian. We then extend our analysis to power-law distributed covariate structures, investigating how anisotropy in the training set alters the learning dynamics. Within this regime, we derive exact analytical learning curves for both the variance and bias terms, precisely capturing key phenomena such as interpolation peaks, multiple descent behavior, and scaling laws.

Sharp description of Kernel Ridge Regression under power-law data

RIZZI, LORENZO
2025/2026

Abstract

Deep neural networks are central to recent breakthroughs in artificial intelligence, yet we still lack a full understanding of their theoretical foundations. For instance, in contrast with classical statistical learning theory, highly overparameterized networks can perfectly interpolate training data achieving zero training error while implicitly selecting solutions that generalize exceptionally well. This phenomenon, often conceptualized through the paradigms of ``double descent'' and ``benign overfitting'', has been shown to emerge in a broader class of analytically tractable learning tasks, such as Kernel Ridge Regression (KRR). Consequently, studying these proxy models in depth provides crucial theoretical insights into the fundamental mechanisms of modern machine learning. In this work, we investigate the impact of realistic data distributions on learning dynamics within a KRR framework in the polynomial scaling regime, i.e., when both the number of samples and the data dimension diverge to infinity following a polynomial relationship. Specifically, we analyze how the input data distribution shapes the spectral properties of the kernel operator, which directly govern the model's generalization capabilities through the deterministic equivalents for the bias and variance. Initially, we examine the kernel eigenvalue distribution under the assumption that training data are drawn from an isotropic multivariate Gaussian. We then extend our analysis to power-law distributed covariate structures, investigating how anisotropy in the training set alters the learning dynamics. Within this regime, we derive exact analytical learning curves for both the variance and bias terms, precisely capturing key phenomena such as interpolation peaks, multiple descent behavior, and scaling laws.
2025
Sharp description of Kernel Ridge Regression under power-law data
Kernel Methods
Anisotropic Data
Multiple Descent
Benign Overfitting
Learning Curves
File in questo prodotto:
File Dimensione Formato  
Rizzi_Lorenzo.pdf

accesso aperto

Dimensione 1.8 MB
Formato Adobe PDF
1.8 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/113159