Canonical Correlation Analysis (CCA) is a fundamental multivariate statistical technique for exploring the dependence structure between two sets of variables. However, statistical inference in this context, especially in the presence of nuisance covariates (Partial CCA) or during the sequential extraction of components (step-down procedure), involves complex methodological challenges. Since the classical parametric approach relies on asymptotic approximations that are often inadequate for finite samples, modern research has increasingly turned to conditional resampling methods, which are typically praised for their flexibility and for the promise of “distribution-free” inference. This thesis aims to deconstruct this belief, showing that distributional independence in this specific context is, in fact, an illusion. Through a geometric and algebraic analysis, the work illustrates the failure of discrete transformations, such as permutations and sign-flipping, in the original data space: the application of these methods does not simultaneously preserve the orthogonality constraints imposed by both exogenous and endogenous nuisance factors, the latter being the canonical variates already extracted during the sequential step-down procedure. The only rigorous solution capable of restoring the geometric exactness of the inference lies in the rotatability-based approach proposed by Huh and Jhun, which requires transferring the entire analysis into a reduced subspace. However, the thesis highlights that this recovery of the randomization assumption comes at an unavoidable theoretical cost: the multivariate normality of the observations. In conclusion, the thesis establishes that, through the Huh-Jhun transformation, rotation-based tests, including permutations and sign changes as specific cases, provide exact inference for CCA, but fail to fulfill the original promise of resampling methods as a genuinely non-parametric inferential framework. To support the theoretical framework, the work also illustrates the development and optimization of the computational core of an R software package, designed to implement the analyses and inferential architectures discussed. The framework is further applied to a real neurobiological dataset concerning the cytoarchitecture of the primary visual cortex across different mammalian species. In particular, the application aims to explore the association between the morphometric characteristics of cortical cells and a quantitative representation of the phylogenetic relationships among the species considered, including both terrestrial and marine mammals. This analysis makes it possible to demonstrate the potential of CCA as a tool for multivariate synthesis, while also highlighting the inferential challenges that arise when dealing with real data characterized by small sample sizes, hierarchical structure, and non-ideal distributional conditions.
L'Analisi di Correlazione Canonica (CCA) è una tecnica statistica multivariata fondamentale per esplorare la struttura di dipendenza tra due insiemi di variabili. Tuttavia, l'inferenza statistica in questo contesto, specialmente in presenza di covariate di disturbo (CCA Parziale) o durante l'estrazione sequenziale delle componenti (step-down), presenta insidie metodologiche complesse. Poiché l'approccio parametrico classico si fonda su approssimazioni asintotiche spesso inadeguate per campioni finiti, la ricerca moderna si è orientata verso i metodi di ricampionamento condizionato, tipicamente elogiati per la loro flessibilità e per la promessa di un'inferenza "distribution-free". Il presente lavoro di tesi si propone di decostruire tale convinzione, dimostrando come l'indipendenza distribuzionale in questo specifico contesto rappresenti a tutti gli effetti un'illusione. Attraverso un'analisi geometrica e algebrica, l'elaborato illustra il fallimento delle trasformazioni discrete (permutazioni e sign-flipping) nello spazio originario dei dati: l'applicazione di tali metodi non rispetta simultaneamente i vincoli di ortogonalità imposti dai fattori confondenti esogeni ed endogeni (le variabili canoniche già estratte durante la procedura sequenziale step-down). L'unica soluzione rigorosa in grado di ripristinare l'esattezza geometrica dell'inferenza risiede nell'approccio basato sulla ruotabilità, proposto da Huh e Jhun, il quale prevede di traslare l'intera analisi all'interno di un sottospazio ridotto. Tuttavia, l'elaborato evidenzia come questo recupero dell'ipotesi di randomizzazione esiga un costo teorico ineludibile: la normalità multivariata delle osservazioni. In conclusione, la tesi sancisce che, grazie alla trasformazione di Huh e Jhun, i test basati sulle rotazioni, comprese permutazioni ed inversioni di segno come casi specifici, garantiscono un'inferenza esatta per la CCA, ma disattendono la promessa originaria dei metodi di ricampionamento di un'inferenza non parametrica. A supporto dell’impianto teorico, il lavoro illustra lo sviluppo e l’ottimizzazione del nucleo computazionale di un pacchetto software in ambiente R, progettato per implementare le analisi e le architetture inferenziali trattate. Il framework viene inoltre applicato a un dataset neurobiologico reale relativo allo studio della citoarchitettura della corteccia visiva primaria in diverse specie di mammiferi. In particolare, l’applicazione mira a indagare, in chiave esplorativa, l’associazione tra le caratteristiche morfometriche delle cellule corticali e una rappresentazione quantitativa delle relazioni filogenetiche tra le specie considerate, includendo mammiferi terrestri e marini. Questa analisi consente di mostrare concretamente le potenzialità della CCA come strumento di sintesi multivariata, ma anche le criticità inferenziali che emergono in presenza di dati reali caratterizzati da numerosità ridotta, struttura gerarchica e condizioni distribuzionali non ideali.
L’illusione dell’inferenza distribution-free nell’Analisi di Correlazione Canonica: il ruolo implicito della normalità nei metodi di ricampionamento condizionato
MARTINI, EDOARDO
2025/2026
Abstract
Canonical Correlation Analysis (CCA) is a fundamental multivariate statistical technique for exploring the dependence structure between two sets of variables. However, statistical inference in this context, especially in the presence of nuisance covariates (Partial CCA) or during the sequential extraction of components (step-down procedure), involves complex methodological challenges. Since the classical parametric approach relies on asymptotic approximations that are often inadequate for finite samples, modern research has increasingly turned to conditional resampling methods, which are typically praised for their flexibility and for the promise of “distribution-free” inference. This thesis aims to deconstruct this belief, showing that distributional independence in this specific context is, in fact, an illusion. Through a geometric and algebraic analysis, the work illustrates the failure of discrete transformations, such as permutations and sign-flipping, in the original data space: the application of these methods does not simultaneously preserve the orthogonality constraints imposed by both exogenous and endogenous nuisance factors, the latter being the canonical variates already extracted during the sequential step-down procedure. The only rigorous solution capable of restoring the geometric exactness of the inference lies in the rotatability-based approach proposed by Huh and Jhun, which requires transferring the entire analysis into a reduced subspace. However, the thesis highlights that this recovery of the randomization assumption comes at an unavoidable theoretical cost: the multivariate normality of the observations. In conclusion, the thesis establishes that, through the Huh-Jhun transformation, rotation-based tests, including permutations and sign changes as specific cases, provide exact inference for CCA, but fail to fulfill the original promise of resampling methods as a genuinely non-parametric inferential framework. To support the theoretical framework, the work also illustrates the development and optimization of the computational core of an R software package, designed to implement the analyses and inferential architectures discussed. The framework is further applied to a real neurobiological dataset concerning the cytoarchitecture of the primary visual cortex across different mammalian species. In particular, the application aims to explore the association between the morphometric characteristics of cortical cells and a quantitative representation of the phylogenetic relationships among the species considered, including both terrestrial and marine mammals. This analysis makes it possible to demonstrate the potential of CCA as a tool for multivariate synthesis, while also highlighting the inferential challenges that arise when dealing with real data characterized by small sample sizes, hierarchical structure, and non-ideal distributional conditions.| File | Dimensione | Formato | |
|---|---|---|---|
|
Martini_Edoardo.pdf
accesso aperto
Dimensione
1.77 MB
Formato
Adobe PDF
|
1.77 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/112198