The thesis is part of the field of high-dimensional statistics and probability, which, in the era of "Big Data," has assumed a fundamental role in the analysis of vectors and high-cardinality data structures. The thesis has two chapters. In the first chapter, sub-gaussian random matrices and packing and covering numbers are introduced. These are the theoretical tools that will be used in the rest of the thesis. The key result proved in this chapter is that the operator norm of sub-gaussian matrices is, with great probability and up to multiplicative constants, bounded from above by the sum of the roots of the number of rows and columns. The second chapter studies three application problems in data science and information theory: community detection in the stochastic block model, error-correcting codes for transmitting information via noisy channels, and n-dimensional point clustering in Gaussian mixture models. In all three cases, in the observations the data structure is “hidden” by a sub-gaussian perturbation. The norm estimation in the first chapter is the tool that allows us to control this perturbation by recovering the hidden structure of the data.

La tesi si colloca nell'ambito della statistica e della probabilità ad alta dimensione che, nell'era dei "Big Data", ha assunto un ruolo fondamentale per l'analisi di vettori e strutture dati ad alta cardinalità. La tesi ha due capitoli. Nel primo capitolo vengono introdotte le matrici aleatorie sub-gaussiane e i numeri di impacchettamento e ricoprimento. Questi sono gli strumenti teorici che verranno usati nel resto della tesi. Il risultato chiave dimostrato in questo capitolo è che la norma operatoriale delle matrici sub-gaussiane è, con grande probabilità e a meno di costanti moltiplicative, limitata dall’alto dalla somma delle radici del numero di righe e colonne. Nel secondo capitolo vengono studiati tre problemi applicativi, nell'ambito della data science e della teoria dell'informazione: la community detection nel modello stocastico a blocchi, i codici a correzione d'errore per trasmettere informazioni tramite canali rumorosi e il clustering di punti n-dimensionali in modelli a miscela di gaussiane. In tutti e tre i casi, nelle osservazioni la struttura dei dati è “nascosta” da una perturbazione sub-gaussiana. La stima sulla norma del primo capitolo è lo strumento che permette di controllare tale perturbazione, recuperando la struttura nascosta dei dati.

Matrici aleatorie sub-gaussiane e loro applicazioni.

SALE, CLAUDIA
2025/2026

Abstract

The thesis is part of the field of high-dimensional statistics and probability, which, in the era of "Big Data," has assumed a fundamental role in the analysis of vectors and high-cardinality data structures. The thesis has two chapters. In the first chapter, sub-gaussian random matrices and packing and covering numbers are introduced. These are the theoretical tools that will be used in the rest of the thesis. The key result proved in this chapter is that the operator norm of sub-gaussian matrices is, with great probability and up to multiplicative constants, bounded from above by the sum of the roots of the number of rows and columns. The second chapter studies three application problems in data science and information theory: community detection in the stochastic block model, error-correcting codes for transmitting information via noisy channels, and n-dimensional point clustering in Gaussian mixture models. In all three cases, in the observations the data structure is “hidden” by a sub-gaussian perturbation. The norm estimation in the first chapter is the tool that allows us to control this perturbation by recovering the hidden structure of the data.
2025
Applications of sub-gaussian matrices.
La tesi si colloca nell'ambito della statistica e della probabilità ad alta dimensione che, nell'era dei "Big Data", ha assunto un ruolo fondamentale per l'analisi di vettori e strutture dati ad alta cardinalità. La tesi ha due capitoli. Nel primo capitolo vengono introdotte le matrici aleatorie sub-gaussiane e i numeri di impacchettamento e ricoprimento. Questi sono gli strumenti teorici che verranno usati nel resto della tesi. Il risultato chiave dimostrato in questo capitolo è che la norma operatoriale delle matrici sub-gaussiane è, con grande probabilità e a meno di costanti moltiplicative, limitata dall’alto dalla somma delle radici del numero di righe e colonne. Nel secondo capitolo vengono studiati tre problemi applicativi, nell'ambito della data science e della teoria dell'informazione: la community detection nel modello stocastico a blocchi, i codici a correzione d'errore per trasmettere informazioni tramite canali rumorosi e il clustering di punti n-dimensionali in modelli a miscela di gaussiane. In tutti e tre i casi, nelle osservazioni la struttura dei dati è “nascosta” da una perturbazione sub-gaussiana. La stima sulla norma del primo capitolo è lo strumento che permette di controllare tale perturbazione, recuperando la struttura nascosta dei dati.
Community detection
Modello stocastico
Grafi
File in questo prodotto:
File Dimensione Formato  
Tesi_matrici_subgaussiane_.pdf

accesso aperto

Dimensione 1.05 MB
Formato Adobe PDF
1.05 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/115717