As machine learning systems are increasingly deployed in dynamic environments, the ability to continuously acquire new knowledge has become a fundamental requirement rather than a desirable property, and retraining a model from scratch whenever new data becomes available is often infeasible. In this context, continual learning is a paradigm where models encounter a sequence of tasks or data distributions and must accumulate knowledge as if all data are observed simultaneously. Training samples from different distributions arrive sequentially and the model must adapt to new tasks with limited or no access to previous data while maintaining high performance across all past domains. The primary challenge in this field is catastrophic forgetting, a phenomenon where the adaptation to a new distribution results in a significant degradation of the model's ability to recall previously learned information. A variety of strategies have been developed to counter it, ranging from the regularisation of the parameters that encode past knowledge to the rehearsal of a small memory of previously seen examples. Traditional offline continual learning assumes that each task can be revisited for multiple epochs before moving on to the next one, so that the model updates its weights on the same data several times. This assumption does not hold in real-world applications where data arrives as a continuous stream and cannot be stored and replayed in full, and it is precisely the limitation that the online approach removes. Online Continual Learning (OCL) tightens this setting further. Whereas in offline continual learning each task is delivered as a complete dataset on which training can proceed over as many epochs as required, in OCL the data of a task arrives as a stream of small mini-batches and the model must update its weights as each of them is received, since no later moment exists at which the whole task will be available. A single-pass constraint follows: each sample is observed once and cannot be revisited after its first encounter, so the multi-epoch training on which offline continual learning relies is no longer possible. Although a large number of methods have been proposed for both settings, the analysis of how they act on the model has lagged behind their development, and it remains unclear how the different solutions influence performance, especially with respect to forgetting. A valuable instrument in this direction is to determine where in the model forgetting originates and recent work in continual learning has begun to separate shallow forgetting, the degradation observed at the output layer, from deep forgetting, the degradation of the internal representations, characterising the two analytically in the offline regime. No comparable analysis exists for the online setting, where the single-pass constraint changes the conditions under which such measurements are taken. The objective of this thesis is to fill this gap by investigating deep and shallow forgetting in Online Continual Learning. This work analyses the impact of such strategies on forgetting, focusing primarily on Experience Replay, which stores a small memory of past samples and interleaves them with the incoming data. The analysis relies on instruments that are well established in the literature, such as linear probing and the signal-to-noise ratio of the feature space. Their common purpose is to make the hidden part of the network observable since the internal representation produces no output of its own, it cannot be assessed through accuracy alone and these tools quantify how much of the final forgetting originates there rather than in the output layer. By bringing this analysis to the online setting, the thesis aims to establish how much the stringent constraints of that paradigm affect the way forgetting arises and is measured and whether the conclusions drawn in the offline regime carry over to it.

As machine learning systems are increasingly deployed in dynamic environments, the ability to continuously acquire new knowledge has become a fundamental requirement rather than a desirable property, and retraining a model from scratch whenever new data becomes available is often infeasible. In this context, continual learning is a paradigm where models encounter a sequence of tasks or data distributions and must accumulate knowledge as if all data are observed simultaneously. Training samples from different distributions arrive sequentially and the model must adapt to new tasks with limited or no access to previous data while maintaining high performance across all past domains. The primary challenge in this field is catastrophic forgetting, a phenomenon where the adaptation to a new distribution results in a significant degradation of the model's ability to recall previously learned information. A variety of strategies have been developed to counter it, ranging from the regularisation of the parameters that encode past knowledge to the rehearsal of a small memory of previously seen examples. Traditional offline continual learning assumes that each task can be revisited for multiple epochs before moving on to the next one, so that the model updates its weights on the same data several times. This assumption does not hold in real-world applications where data arrives as a continuous stream and cannot be stored and replayed in full, and it is precisely the limitation that the online approach removes. Online Continual Learning (OCL) tightens this setting further. Whereas in offline continual learning each task is delivered as a complete dataset on which training can proceed over as many epochs as required, in OCL the data of a task arrives as a stream of small mini-batches and the model must update its weights as each of them is received, since no later moment exists at which the whole task will be available. A single-pass constraint follows: each sample is observed once and cannot be revisited after its first encounter, so the multi-epoch training on which offline continual learning relies is no longer possible. Although a large number of methods have been proposed for both settings, the analysis of how they act on the model has lagged behind their development, and it remains unclear how the different solutions influence performance, especially with respect to forgetting. A valuable instrument in this direction is to determine where in the model forgetting originates and recent work in continual learning has begun to separate shallow forgetting, the degradation observed at the output layer, from deep forgetting, the degradation of the internal representations, characterising the two analytically in the offline regime. No comparable analysis exists for the online setting, where the single-pass constraint changes the conditions under which such measurements are taken. The objective of this thesis is to fill this gap by investigating deep and shallow forgetting in Online Continual Learning. This work analyses the impact of such strategies on forgetting, focusing primarily on Experience Replay, which stores a small memory of past samples and interleaves them with the incoming data. The analysis relies on instruments that are well established in the literature, such as linear probing and the signal-to-noise ratio of the feature space. Their common purpose is to make the hidden part of the network observable since the internal representation produces no output of its own, it cannot be assessed through accuracy alone and these tools quantify how much of the final forgetting originates there rather than in the output layer. By bringing this analysis to the online setting, the thesis aims to establish how much the stringent constraints of that paradigm affect the way forgetting arises and is measured and whether the conclusions drawn in the offline regime carry over to it.

Investigating deep and shallow forgetting in Online Continual Learning

CALZAROTTO, ANDREA
2025/2026

Abstract

As machine learning systems are increasingly deployed in dynamic environments, the ability to continuously acquire new knowledge has become a fundamental requirement rather than a desirable property, and retraining a model from scratch whenever new data becomes available is often infeasible. In this context, continual learning is a paradigm where models encounter a sequence of tasks or data distributions and must accumulate knowledge as if all data are observed simultaneously. Training samples from different distributions arrive sequentially and the model must adapt to new tasks with limited or no access to previous data while maintaining high performance across all past domains. The primary challenge in this field is catastrophic forgetting, a phenomenon where the adaptation to a new distribution results in a significant degradation of the model's ability to recall previously learned information. A variety of strategies have been developed to counter it, ranging from the regularisation of the parameters that encode past knowledge to the rehearsal of a small memory of previously seen examples. Traditional offline continual learning assumes that each task can be revisited for multiple epochs before moving on to the next one, so that the model updates its weights on the same data several times. This assumption does not hold in real-world applications where data arrives as a continuous stream and cannot be stored and replayed in full, and it is precisely the limitation that the online approach removes. Online Continual Learning (OCL) tightens this setting further. Whereas in offline continual learning each task is delivered as a complete dataset on which training can proceed over as many epochs as required, in OCL the data of a task arrives as a stream of small mini-batches and the model must update its weights as each of them is received, since no later moment exists at which the whole task will be available. A single-pass constraint follows: each sample is observed once and cannot be revisited after its first encounter, so the multi-epoch training on which offline continual learning relies is no longer possible. Although a large number of methods have been proposed for both settings, the analysis of how they act on the model has lagged behind their development, and it remains unclear how the different solutions influence performance, especially with respect to forgetting. A valuable instrument in this direction is to determine where in the model forgetting originates and recent work in continual learning has begun to separate shallow forgetting, the degradation observed at the output layer, from deep forgetting, the degradation of the internal representations, characterising the two analytically in the offline regime. No comparable analysis exists for the online setting, where the single-pass constraint changes the conditions under which such measurements are taken. The objective of this thesis is to fill this gap by investigating deep and shallow forgetting in Online Continual Learning. This work analyses the impact of such strategies on forgetting, focusing primarily on Experience Replay, which stores a small memory of past samples and interleaves them with the incoming data. The analysis relies on instruments that are well established in the literature, such as linear probing and the signal-to-noise ratio of the feature space. Their common purpose is to make the hidden part of the network observable since the internal representation produces no output of its own, it cannot be assessed through accuracy alone and these tools quantify how much of the final forgetting originates there rather than in the output layer. By bringing this analysis to the online setting, the thesis aims to establish how much the stringent constraints of that paradigm affect the way forgetting arises and is measured and whether the conclusions drawn in the offline regime carry over to it.
2025
Investigating deep and shallow forgetting in Online Continual Learning
As machine learning systems are increasingly deployed in dynamic environments, the ability to continuously acquire new knowledge has become a fundamental requirement rather than a desirable property, and retraining a model from scratch whenever new data becomes available is often infeasible. In this context, continual learning is a paradigm where models encounter a sequence of tasks or data distributions and must accumulate knowledge as if all data are observed simultaneously. Training samples from different distributions arrive sequentially and the model must adapt to new tasks with limited or no access to previous data while maintaining high performance across all past domains. The primary challenge in this field is catastrophic forgetting, a phenomenon where the adaptation to a new distribution results in a significant degradation of the model's ability to recall previously learned information. A variety of strategies have been developed to counter it, ranging from the regularisation of the parameters that encode past knowledge to the rehearsal of a small memory of previously seen examples. Traditional offline continual learning assumes that each task can be revisited for multiple epochs before moving on to the next one, so that the model updates its weights on the same data several times. This assumption does not hold in real-world applications where data arrives as a continuous stream and cannot be stored and replayed in full, and it is precisely the limitation that the online approach removes. Online Continual Learning (OCL) tightens this setting further. Whereas in offline continual learning each task is delivered as a complete dataset on which training can proceed over as many epochs as required, in OCL the data of a task arrives as a stream of small mini-batches and the model must update its weights as each of them is received, since no later moment exists at which the whole task will be available. A single-pass constraint follows: each sample is observed once and cannot be revisited after its first encounter, so the multi-epoch training on which offline continual learning relies is no longer possible. Although a large number of methods have been proposed for both settings, the analysis of how they act on the model has lagged behind their development, and it remains unclear how the different solutions influence performance, especially with respect to forgetting. A valuable instrument in this direction is to determine where in the model forgetting originates and recent work in continual learning has begun to separate shallow forgetting, the degradation observed at the output layer, from deep forgetting, the degradation of the internal representations, characterising the two analytically in the offline regime. No comparable analysis exists for the online setting, where the single-pass constraint changes the conditions under which such measurements are taken. The objective of this thesis is to fill this gap by investigating deep and shallow forgetting in Online Continual Learning. This work analyses the impact of such strategies on forgetting, focusing primarily on Experience Replay, which stores a small memory of past samples and interleaves them with the incoming data. The analysis relies on instruments that are well established in the literature, such as linear probing and the signal-to-noise ratio of the feature space. Their common purpose is to make the hidden part of the network observable since the internal representation produces no output of its own, it cannot be assessed through accuracy alone and these tools quantify how much of the final forgetting originates there rather than in the output layer. By bringing this analysis to the online setting, the thesis aims to establish how much the stringent constraints of that paradigm affect the way forgetting arises and is measured and whether the conclusions drawn in the offline regime carry over to it.
Neural Networks
Continual Learning
Forgetting
Linear Probe
Feature Space
File in questo prodotto:
File Dimensione Formato  
Calzarotto_Andrea.pdf

Accesso riservato

Dimensione 13.22 MB
Formato Adobe PDF
13.22 MB Adobe PDF

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/115824