This thesis investigates the application of biological foundation models to the task of ortholog identification. Ortholog datasets were constructed from National Center for Biotechnology Information and OMA by selecting eukaryotic species with uniform evolutionary distance from humans and organizing genes into ortholog families. An evolutionary-aware contrastive learning loss was developed to incorporate phylogenetic information into representation learning. Embeddings extracted from Nucleotide Transformer v2 and Evo 2 were used to train a lightweight contrastive projection head for orthology prediction. Different sequence configurations were evaluated, including only coding regions, regulatory regions, and full genomic sequences. The work also includes representation alignment analyses through contrastive learning and preliminary scaling law evaluations for Evo 2.
This thesis investigates the application of biological foundation models to the task of ortholog identification. Ortholog datasets were constructed from National Center for Biotechnology Information and OMA by selecting eukaryotic species with uniform evolutionary distance from humans and organizing genes into ortholog families. An evolutionary-aware contrastive learning loss was developed to incorporate phylogenetic information into representation learning. Embeddings extracted from Nucleotide Transformer v2 and Evo 2 were used to train a lightweight contrastive projection head for orthology prediction. Different sequence configurations were evaluated, including only coding regions, regulatory regions, and full genomic sequences. The work also includes representation alignment analyses through contrastive learning and preliminary scaling law evaluations for Evo 2.
Transformers vs. State Space Models for Genomic Representation Learning: Orthology Detection, Latent Space Alignment, and Trade-offs of Long-Context Modeling
RIGATO, MARGHERITA
2025/2026
Abstract
This thesis investigates the application of biological foundation models to the task of ortholog identification. Ortholog datasets were constructed from National Center for Biotechnology Information and OMA by selecting eukaryotic species with uniform evolutionary distance from humans and organizing genes into ortholog families. An evolutionary-aware contrastive learning loss was developed to incorporate phylogenetic information into representation learning. Embeddings extracted from Nucleotide Transformer v2 and Evo 2 were used to train a lightweight contrastive projection head for orthology prediction. Different sequence configurations were evaluated, including only coding regions, regulatory regions, and full genomic sequences. The work also includes representation alignment analyses through contrastive learning and preliminary scaling law evaluations for Evo 2.| File | Dimensione | Formato | |
|---|---|---|---|
|
margherita_rigato_FINALE (1).pdf
Accesso riservato
Dimensione
5.55 MB
Formato
Adobe PDF
|
5.55 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110933