This thesis explores the application of DNABERT, a language model based on the Transformer architecture and pre-trained on the human genome, for decoding DNA regulatory sequences. Starting from the analysis of the self-supervised learning process based on k-mers and Masked Language Modeling (MLM), the work investigates how the model captures global genomic syntax to solve various biological tasks. The research focuses on evaluating the model's performance in critical tasks such as promoter identification, splice site recognition, and protein binding site prediction, demonstrating superior accuracy compared to traditional computational methods. Particular attention is given to the model’s ability to interpret sequence context to identify functional motifs through the analysis of attention weights. Finally, the thesis presents a case study on core promoters and the TATA-box signal, highlighting DNABERT’s effectiveness in recognizing complex regulatory signals and predicting the impact of genetic variants, establishing it as an advanced tool for genomic research.
Questa tesi analizza l’applicazione di DNABERT, un modello linguistico basato sull’architettura Transformer e pre-addestrato sul genoma umano, per la decodifica delle sequenze regolatrici del DNA. Partendo dall'analisi del processo di apprendimento auto-supervisionato basato su k-mer e Masked Language Modeling (MLM), l'elaborato esplora come il modello sia in grado di catturare la sintassi genomica globale per risolvere diversi compiti biologici. La ricerca si concentra sulla valutazione delle prestazioni del modello in task critici quali l'identificazione dei promotori, dei siti di splicing e dei siti di legame proteico, dimostrando una precisione superiore rispetto ai metodi computazionali tradizionali. Particolare attenzione viene posta sulla capacità del modello di interpretare il contesto di sequenza per identificare motivi funzionali attraverso l'analisi dei pesi di attenzione. Infine, la tesi presenta un caso studio dedicato ai promotori core e al segnale TATA-box, evidenziando l’efficacia di DNABERT nel riconoscere segnali regolatori complessi e nel prevedere l'impatto di varianti genetiche, consolidandolo come uno strumento avanzato per la ricerca in ambito genomico.
Applicazione del modello DNABERT nell’analisi del DNA: studio dei promotori TATA box.
VADORI, ANDREA
2025/2026
Abstract
This thesis explores the application of DNABERT, a language model based on the Transformer architecture and pre-trained on the human genome, for decoding DNA regulatory sequences. Starting from the analysis of the self-supervised learning process based on k-mers and Masked Language Modeling (MLM), the work investigates how the model captures global genomic syntax to solve various biological tasks. The research focuses on evaluating the model's performance in critical tasks such as promoter identification, splice site recognition, and protein binding site prediction, demonstrating superior accuracy compared to traditional computational methods. Particular attention is given to the model’s ability to interpret sequence context to identify functional motifs through the analysis of attention weights. Finally, the thesis presents a case study on core promoters and the TATA-box signal, highlighting DNABERT’s effectiveness in recognizing complex regulatory signals and predicting the impact of genetic variants, establishing it as an advanced tool for genomic research.| File | Dimensione | Formato | |
|---|---|---|---|
|
Vadori_Andrea.pdf
accesso aperto
Dimensione
374.09 kB
Formato
Adobe PDF
|
374.09 kB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/111176