Bacteriophages are viruses that infect bacteria and play a central role in microbial ecosystems, while also representing a promising alternative to antibiotics in the context of increasing antimicrobial resistance. A key challenge in phage-based applications is the identification of suitable candidates, which depends on determining their lifestyle, typically classified as virulent or temperate. However, rapid growth of sequencing data has outpaced the availability of accurate annotations, motivating the development of computational approaches for inferring phage properties directly from sequence data. In this work, we investigate whether representations derived from a biological language model encode signals related to bacteriophage lifestyle. Using the Evo 2 model, we generate embeddings from protein-coding regions and analyze them at protein, genome, and functional group levels. Protein-level classification achieved moderate performance, with F1 score equal to 0.77 and Matthews correlation coefficient (MCC) equal to 0.56. Aggregating protein embeddings at the genome level using a simple mean significantly improves results, achieving F1 score equal to 0.89 and MCC equal to 0.78, indicating that lifestyle-related signals are distributed across multiple proteins. Interpretability analyses reveal that only a subset of proteins contributes strongly to classification, causing the classifier confidence score to shift from 0 to 1 or vice versa, while the majority of proteins have negligible impact with classifier confidence score shifts close to zero. Moreover, when proteins are grouped using keywords or functional categories, the score shifts direction is biologically consistent, with mean confidence shifts ranging from -0.05 to +0.07. Overall, the results indicate that protein embeddings capture meaningful signals related to phage lifestyle, despite noisy annotations, highlighting their potential for biological analysis and interpretation.

Bacteriophages are viruses that infect bacteria and play a central role in microbial ecosystems, while also representing a promising alternative to antibiotics in the context of increasing antimicrobial resistance. A key challenge in phage-based applications is the identification of suitable candidates, which depends on determining their lifestyle, typically classified as virulent or temperate. However, rapid growth of sequencing data has outpaced the availability of accurate annotations, motivating the development of computational approaches for inferring phage properties directly from sequence data. In this work, we investigate whether representations derived from a biological language model encode signals related to bacteriophage lifestyle. Using the Evo 2 model, we generate embeddings from protein-coding regions and analyze them at protein, genome, and functional group levels. Protein-level classification achieved moderate performance, with F1 score equal to 0.77 and Matthews correlation coefficient (MCC) equal to 0.56. Aggregating protein embeddings at the genome level using a simple mean significantly improves results, achieving F1 score equal to 0.89 and MCC equal to 0.78, indicating that lifestyle-related signals are distributed across multiple proteins. Interpretability analyses reveal that only a subset of proteins contributes strongly to classification, causing the classifier confidence score to shift from 0 to 1 or vice versa, while the majority of proteins have negligible impact with classifier confidence score shifts close to zero. Moreover, when proteins are grouped using keywords or functional categories, the score shifts direction is biologically consistent, with mean confidence shifts ranging from -0.05 to +0.07. Overall, the results indicate that protein embeddings capture meaningful signals related to phage lifestyle, despite noisy annotations, highlighting their potential for biological analysis and interpretation.

Using Evo 2 embeddings to classify phage genomes and identify decision driving proteins

TOKAYEV, DMYTRO
2025/2026

Abstract

Bacteriophages are viruses that infect bacteria and play a central role in microbial ecosystems, while also representing a promising alternative to antibiotics in the context of increasing antimicrobial resistance. A key challenge in phage-based applications is the identification of suitable candidates, which depends on determining their lifestyle, typically classified as virulent or temperate. However, rapid growth of sequencing data has outpaced the availability of accurate annotations, motivating the development of computational approaches for inferring phage properties directly from sequence data. In this work, we investigate whether representations derived from a biological language model encode signals related to bacteriophage lifestyle. Using the Evo 2 model, we generate embeddings from protein-coding regions and analyze them at protein, genome, and functional group levels. Protein-level classification achieved moderate performance, with F1 score equal to 0.77 and Matthews correlation coefficient (MCC) equal to 0.56. Aggregating protein embeddings at the genome level using a simple mean significantly improves results, achieving F1 score equal to 0.89 and MCC equal to 0.78, indicating that lifestyle-related signals are distributed across multiple proteins. Interpretability analyses reveal that only a subset of proteins contributes strongly to classification, causing the classifier confidence score to shift from 0 to 1 or vice versa, while the majority of proteins have negligible impact with classifier confidence score shifts close to zero. Moreover, when proteins are grouped using keywords or functional categories, the score shifts direction is biologically consistent, with mean confidence shifts ranging from -0.05 to +0.07. Overall, the results indicate that protein embeddings capture meaningful signals related to phage lifestyle, despite noisy annotations, highlighting their potential for biological analysis and interpretation.
2025
Using Evo 2 embeddings to classify phage genomes and identify decision driving proteins
Bacteriophages are viruses that infect bacteria and play a central role in microbial ecosystems, while also representing a promising alternative to antibiotics in the context of increasing antimicrobial resistance. A key challenge in phage-based applications is the identification of suitable candidates, which depends on determining their lifestyle, typically classified as virulent or temperate. However, rapid growth of sequencing data has outpaced the availability of accurate annotations, motivating the development of computational approaches for inferring phage properties directly from sequence data. In this work, we investigate whether representations derived from a biological language model encode signals related to bacteriophage lifestyle. Using the Evo 2 model, we generate embeddings from protein-coding regions and analyze them at protein, genome, and functional group levels. Protein-level classification achieved moderate performance, with F1 score equal to 0.77 and Matthews correlation coefficient (MCC) equal to 0.56. Aggregating protein embeddings at the genome level using a simple mean significantly improves results, achieving F1 score equal to 0.89 and MCC equal to 0.78, indicating that lifestyle-related signals are distributed across multiple proteins. Interpretability analyses reveal that only a subset of proteins contributes strongly to classification, causing the classifier confidence score to shift from 0 to 1 or vice versa, while the majority of proteins have negligible impact with classifier confidence score shifts close to zero. Moreover, when proteins are grouped using keywords or functional categories, the score shifts direction is biologically consistent, with mean confidence shifts ranging from -0.05 to +0.07. Overall, the results indicate that protein embeddings capture meaningful signals related to phage lifestyle, despite noisy annotations, highlighting their potential for biological analysis and interpretation.
LLM
Bacteriophages
DNA embeddings
File in questo prodotto:
File Dimensione Formato  
Tokayev_Dmytro.pdf

embargo fino al 22/04/2029

Dimensione 3.91 MB
Formato Adobe PDF
3.91 MB Adobe PDF

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/108021