In generative models of viral mutations under antibody pressure, additive models are commonly used to describe the binding energies of mutated viral sequences to the host-cell receptor and to antibodies. The corresponding binding vectors are inferred from deep mutational scanning data, in which the affinity of the viral spike-protein to a target is measured for libraries of mutants close to a wild-type sequence. Even neglecting epistatic effects, the fact that the model is learned from scarce and noisy experimental data makes it crucial to understand the expected performance of the inferred model in predicting the binding energy for novel antigen sequences, and to determine the optimal regularization to employ in such applications. This thesis develops a theoretical framework for high-dimensional ridge regression on protein sequence data, with the aim of understanding the interplay between sample structure, regularization strength, and generalization error. The analytical results, validated on synthetic data, reveal how the number of samples, the mutation probability, and the specificity of the target to the wild-type sequence determine the optimal regularization strength and the achievable prediction error. The framework is then extended to out-of-sample prediction, where test sequences carry more mutations than the sequences used for training. This setting captures a central difficulty in immunological applications, in which inferred models are used to make predictions beyond the region directly explored by experiments. The analysis quantifies the resulting deterioration in performance and shows that the optimal regularization depends on the test distribution. Additionally, epistatic effects are effectively integrated in this setting by considering a misspecified model. The theoretical predictions are tested on experimental measurements of the binding of SARS-CoV-2 spike-protein mutants to the host-cell receptor ACE2 and to a panel of antibodies. The comparison between the predicted and empirical risk landscapes across training-set sizes and regularization strengths allows to assess the validity of the theoretical description and to infer properties of the underlying binding problem. Finally, the inferred binding models are incorporated into a generative model of SARS-CoV-2 evolution under antibody pressure. By combining antibody escape and ACE2-binding predictions with a sequence-based model of viral fitness, the model generates candidate mutational trajectories toward variants with increased fitness under immune selection. The resulting escape mutations and evolutionary paths are compared with those observed in in-vitro escape experiments.
In generative models of viral mutations under antibody pressure, additive models are commonly used to describe the binding energies of mutated viral sequences to the host-cell receptor and to antibodies. The corresponding binding vectors are inferred from deep mutational scanning data, in which the affinity of the viral spike-protein to a target is measured for libraries of mutants close to a wild-type sequence. Even neglecting epistatic effects, the fact that the model is learned from scarce and noisy experimental data makes it crucial to understand the expected performance of the inferred model in predicting the binding energy for novel antigen sequences, and to determine the optimal regularization to employ in such applications. This thesis develops a theoretical framework for high-dimensional ridge regression on protein sequence data, with the aim of understanding the interplay between sample structure, regularization strength, and generalization error. The analytical results, validated on synthetic data, reveal how the number of samples, the mutation probability, and the specificity of the target to the wild-type sequence determine the optimal regularization strength and the achievable prediction error. The framework is then extended to out-of-sample prediction, where test sequences carry more mutations than the sequences used for training. This setting captures a central difficulty in immunological applications, in which inferred models are used to make predictions beyond the region directly explored by experiments. The analysis quantifies the resulting deterioration in performance and shows that the optimal regularization depends on the test distribution. Additionally, epistatic effects are effectively integrated in this setting by considering a misspecified model. The theoretical predictions are tested on experimental measurements of the binding of SARS-CoV-2 spike-protein mutants to the host-cell receptor ACE2 and to a panel of antibodies. The comparison between the predicted and empirical risk landscapes across training-set sizes and regularization strengths allows to assess the validity of the theoretical description and to infer properties of the underlying binding problem. Finally, the inferred binding models are incorporated into a generative model of SARS-CoV-2 evolution under antibody pressure. By combining antibody escape and ACE2-binding predictions with a sequence-based model of viral fitness, the model generates candidate mutational trajectories toward variants with increased fitness under immune selection. The resulting escape mutations and evolutionary paths are compared with those observed in in-vitro escape experiments.
Inference and Generalization in Models of Viral Escape and Antibody Binding
CORSO, ALESSANDRO
2025/2026
Abstract
In generative models of viral mutations under antibody pressure, additive models are commonly used to describe the binding energies of mutated viral sequences to the host-cell receptor and to antibodies. The corresponding binding vectors are inferred from deep mutational scanning data, in which the affinity of the viral spike-protein to a target is measured for libraries of mutants close to a wild-type sequence. Even neglecting epistatic effects, the fact that the model is learned from scarce and noisy experimental data makes it crucial to understand the expected performance of the inferred model in predicting the binding energy for novel antigen sequences, and to determine the optimal regularization to employ in such applications. This thesis develops a theoretical framework for high-dimensional ridge regression on protein sequence data, with the aim of understanding the interplay between sample structure, regularization strength, and generalization error. The analytical results, validated on synthetic data, reveal how the number of samples, the mutation probability, and the specificity of the target to the wild-type sequence determine the optimal regularization strength and the achievable prediction error. The framework is then extended to out-of-sample prediction, where test sequences carry more mutations than the sequences used for training. This setting captures a central difficulty in immunological applications, in which inferred models are used to make predictions beyond the region directly explored by experiments. The analysis quantifies the resulting deterioration in performance and shows that the optimal regularization depends on the test distribution. Additionally, epistatic effects are effectively integrated in this setting by considering a misspecified model. The theoretical predictions are tested on experimental measurements of the binding of SARS-CoV-2 spike-protein mutants to the host-cell receptor ACE2 and to a panel of antibodies. The comparison between the predicted and empirical risk landscapes across training-set sizes and regularization strengths allows to assess the validity of the theoretical description and to infer properties of the underlying binding problem. Finally, the inferred binding models are incorporated into a generative model of SARS-CoV-2 evolution under antibody pressure. By combining antibody escape and ACE2-binding predictions with a sequence-based model of viral fitness, the model generates candidate mutational trajectories toward variants with increased fitness under immune selection. The resulting escape mutations and evolutionary paths are compared with those observed in in-vitro escape experiments.| File | Dimensione | Formato | |
|---|---|---|---|
|
Corso_Alessandro.pdf
Accesso riservato
Dimensione
2.3 MB
Formato
Adobe PDF
|
2.3 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110074