In machine learning, the main aim is to identify a single best model that ideally provides both interpretable results and accurate predictions for a given problem. However, recent research has shown that there can be many models that per- form approximately equally well for the same dataset, each of them potentially yielding different interpretations or indications on variable importance [Rudin et al., 2024]. This phenomenon is known as the Rashomon Effect and was coined by Leo Breiman in his seminal paper “Statistical Modeling: The Two Cultures” [Breiman, 2001]. One of the implications of the Rashomon Effect is that, when many almost-equivalently predictive models exist, conclusions about the significance of variables or their importance will depend on which model is chosen [Rudin et al., 2024]. In particular, while the Rashomon Effect has been explained by the presence of noise in the data [Semenova et al., 2023], in this work we investigate how to address the Rashomon Effect when there is hetero- geneity in the data. More specifically, we assume that groups of observations (e.g. individuals) express their responses in different ways (i.e. through differ- ent combinations of variables). This is the setting, for example, when dealing with mental disorders (e.g. ADHD, Autism) where it is commonly known that, also due to unobservable latent factors, the diagnosis and treatment of these disorders varies highly among groups of individuals. Therefore, given the need to find multiple good models based on different combinations of variables, in this work we make use of the Sparse Wrapper Algorithm (SWAG) [Molinari et al., 2020] consisting in a heuristic forward-search algorithm that is designed to identify highly-predictive sets of sparse models with diverse variable com- binations. Once this set is defined, the goal of this work is twofold: (i) when using GLMs within the SWAG, we want to explore ways to combine p-values for each variable from all models to obtain overall variable significance metrics; (ii) we want to explore ways to deliver individual-level predictions by combining those of all models in the set. These investigations can provide new statistical inference tools and more calibrated predictions to address the Rashomon Effect.
Nel machine learning, l'obiettivo principale è identificare un singolo modello migliore che idealmente fornisca sia risultati interpretabili sia predizioni accurate per un dato problema. Tuttavia, ricerche recenti hanno dimostrato che possono esistere molti modelli che ottengono prestazioni approssimativamente ugualmente buone per lo stesso dataset, ciascuno dei quali potenzialmente in grado di fornire interpretazioni o indicazioni differenti sull'importanza delle variabili [Rudin et al., 2024]. Questo fenomeno è noto come Effetto Rashomon ed è stato coniato da Leo Breiman nel suo lavoro seminale “Statistical Modeling: The Two Cultures” [Breiman, 2001]. Una delle implicazioni dell'Effetto Rashomon è che, quando esistono molti modelli con capacità predittiva quasi equivalente, le conclusioni sulla significatività delle variabili o sulla loro importanza dipenderanno da quale modello viene scelto [Rudin et al., 2024]. In particolare, mentre l'Effetto Rashomon è stato spiegato dalla presenza di rumore nei dati [Semenova et al., 2023], in questo lavoro indaghiamo come affrontare l'Effetto Rashomon quando vi è eterogeneità nei dati. Più nello specifico, assumiamo che gruppi di osservazioni (ad esempio individui) esprimano le loro risposte in modi diversi (ovvero attraverso diverse combinazioni di variabili). Questo è il contesto, ad esempio, in cui ci si trova quando si ha a che fare con disturbi mentali (ad esempio ADHD, Autismo), dove è comunemente noto che, anche a causa di fattori latenti non osservabili, la diagnosi e il trattamento di questi disturbi variano fortemente tra gruppi di individui. Pertanto, data la necessità di trovare molteplici modelli validi basati su diverse combinazioni di variabili, in questo lavoro facciamo uso dello Sparse Wrapper Algorithm (SWAG) [Molinari et al., 2020], che consiste in un algoritmo euristico di ricerca in avanti (forward-search) progettato per identificare insiemi altamente predittivi di modelli sparsi con diverse combinazioni di variabili. Una volta definito questo insieme, l'obiettivo di questo lavoro è duplice: (i) quando si utilizzano i GLM all'interno dello SWAG, vogliamo esplorare modi per combinare i p-value di ciascuna variabile provenienti da tutti i modelli per ottenere metriche complessive di significatività delle variabili; (ii) vogliamo esplorare modi per fornire predizioni a livello individuale combinando quelle di tutti i modelli presenti nell'insieme. Queste indagini possono fornire nuovi strumenti di inferenza statistica e predizioni più calibrate per affrontare l'Effetto Rashomon.
Inferenza statistica e previsione per insiemi di modelli sparsi
MORELLO, GIORGIA
2025/2026
Abstract
In machine learning, the main aim is to identify a single best model that ideally provides both interpretable results and accurate predictions for a given problem. However, recent research has shown that there can be many models that per- form approximately equally well for the same dataset, each of them potentially yielding different interpretations or indications on variable importance [Rudin et al., 2024]. This phenomenon is known as the Rashomon Effect and was coined by Leo Breiman in his seminal paper “Statistical Modeling: The Two Cultures” [Breiman, 2001]. One of the implications of the Rashomon Effect is that, when many almost-equivalently predictive models exist, conclusions about the significance of variables or their importance will depend on which model is chosen [Rudin et al., 2024]. In particular, while the Rashomon Effect has been explained by the presence of noise in the data [Semenova et al., 2023], in this work we investigate how to address the Rashomon Effect when there is hetero- geneity in the data. More specifically, we assume that groups of observations (e.g. individuals) express their responses in different ways (i.e. through differ- ent combinations of variables). This is the setting, for example, when dealing with mental disorders (e.g. ADHD, Autism) where it is commonly known that, also due to unobservable latent factors, the diagnosis and treatment of these disorders varies highly among groups of individuals. Therefore, given the need to find multiple good models based on different combinations of variables, in this work we make use of the Sparse Wrapper Algorithm (SWAG) [Molinari et al., 2020] consisting in a heuristic forward-search algorithm that is designed to identify highly-predictive sets of sparse models with diverse variable com- binations. Once this set is defined, the goal of this work is twofold: (i) when using GLMs within the SWAG, we want to explore ways to combine p-values for each variable from all models to obtain overall variable significance metrics; (ii) we want to explore ways to deliver individual-level predictions by combining those of all models in the set. These investigations can provide new statistical inference tools and more calibrated predictions to address the Rashomon Effect.| File | Dimensione | Formato | |
|---|---|---|---|
|
Morello_Giorgia.pdf
accesso aperto
Dimensione
1.87 MB
Formato
Adobe PDF
|
1.87 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/112246