Bacteriophages are viruses that infect bacteria and play a central role in microbial ecosystems, while also representing a promising alternative to antibiotics in the context of increasing antimicrobial resistance. A key challenge in phage-based applications is the identification of suitable candidates, which depends on determining their lifestyle, typically classified as virulent or temperate. However, rapid growth of sequencing data has outpaced the availability of accurate annotations, motivating the development of computational approaches for inferring phage properties directly from sequence data. In this work, we investigate whether representations derived from a biological language model encode signals related to bacteriophage lifestyle. Using the Evo 2 model, we generate embeddings from protein-coding regions and analyze them at protein, genome, and functional group levels. Protein-level classification achieved moderate performance, with F1 score equal to 0.77 and Matthews correlation coefficient (MCC) equal to 0.56. Aggregating protein embeddings at the genome level using a simple mean significantly improves results, achieving F1 score equal to 0.89 and MCC equal to 0.78, indicating that lifestyle-related signals are distributed across multiple proteins. Interpretability analyses reveal that only a subset of proteins contributes strongly to classification, causing the classifier confidence score to shift from 0 to 1 or vice versa, while the majority of proteins have negligible impact with classifier confidence score shifts close to zero. Moreover, when proteins are grouped using keywords or functional categories, the score shifts direction is biologically consistent, with mean confidence shifts ranging from -0.05 to +0.07. Overall, the results indicate that protein embeddings capture meaningful signals related to phage lifestyle, despite noisy annotations, highlighting their potential for biological analysis and interpretation.
Bacteriophages are viruses that infect bacteria and play a central role in microbial ecosystems, while also representing a promising alternative to antibiotics in the context of increasing antimicrobial resistance. A key challenge in phage-based applications is the identification of suitable candidates, which depends on determining their lifestyle, typically classified as virulent or temperate. However, rapid growth of sequencing data has outpaced the availability of accurate annotations, motivating the development of computational approaches for inferring phage properties directly from sequence data. In this work, we investigate whether representations derived from a biological language model encode signals related to bacteriophage lifestyle. Using the Evo 2 model, we generate embeddings from protein-coding regions and analyze them at protein, genome, and functional group levels. Protein-level classification achieved moderate performance, with F1 score equal to 0.77 and Matthews correlation coefficient (MCC) equal to 0.56. Aggregating protein embeddings at the genome level using a simple mean significantly improves results, achieving F1 score equal to 0.89 and MCC equal to 0.78, indicating that lifestyle-related signals are distributed across multiple proteins. Interpretability analyses reveal that only a subset of proteins contributes strongly to classification, causing the classifier confidence score to shift from 0 to 1 or vice versa, while the majority of proteins have negligible impact with classifier confidence score shifts close to zero. Moreover, when proteins are grouped using keywords or functional categories, the score shifts direction is biologically consistent, with mean confidence shifts ranging from -0.05 to +0.07. Overall, the results indicate that protein embeddings capture meaningful signals related to phage lifestyle, despite noisy annotations, highlighting their potential for biological analysis and interpretation.
Using Evo 2 embeddings to classify phage genomes and identify decision driving proteins
TOKAYEV, DMYTRO
2025/2026
Abstract
Bacteriophages are viruses that infect bacteria and play a central role in microbial ecosystems, while also representing a promising alternative to antibiotics in the context of increasing antimicrobial resistance. A key challenge in phage-based applications is the identification of suitable candidates, which depends on determining their lifestyle, typically classified as virulent or temperate. However, rapid growth of sequencing data has outpaced the availability of accurate annotations, motivating the development of computational approaches for inferring phage properties directly from sequence data. In this work, we investigate whether representations derived from a biological language model encode signals related to bacteriophage lifestyle. Using the Evo 2 model, we generate embeddings from protein-coding regions and analyze them at protein, genome, and functional group levels. Protein-level classification achieved moderate performance, with F1 score equal to 0.77 and Matthews correlation coefficient (MCC) equal to 0.56. Aggregating protein embeddings at the genome level using a simple mean significantly improves results, achieving F1 score equal to 0.89 and MCC equal to 0.78, indicating that lifestyle-related signals are distributed across multiple proteins. Interpretability analyses reveal that only a subset of proteins contributes strongly to classification, causing the classifier confidence score to shift from 0 to 1 or vice versa, while the majority of proteins have negligible impact with classifier confidence score shifts close to zero. Moreover, when proteins are grouped using keywords or functional categories, the score shifts direction is biologically consistent, with mean confidence shifts ranging from -0.05 to +0.07. Overall, the results indicate that protein embeddings capture meaningful signals related to phage lifestyle, despite noisy annotations, highlighting their potential for biological analysis and interpretation.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tokayev_Dmytro.pdf
embargo fino al 22/04/2029
Dimensione
3.91 MB
Formato
Adobe PDF
|
3.91 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/108021