Multivariate time series data plays a critical role in domains such as energy markets, where data completeness and accuracy are essential for reliable analysis and decision-making. However, real-world datasets frequently contain missing values due to factors such as low market liquidity, reporting delays, and arbitrage inconsistencies, which can significantly degrade the performance of data-driven models. Despite advances in deep learning-based imputation, limited work has examined these methods under realistic conditions of missingness specific to energy pricing data. This thesis investigates the effectiveness of various deep learning imputation techniques for multivariate energy pricing time series, intending to develop a robust and reliable imputation pipeline. A range of methods is evaluated, including transformer-based architectures such as SAITS and Transformer, a token-based approach such as TOTEM, and other neural network models, including TSLANet, and an LSTM autoencoder serving as a baseline, with models implemented within the PyPOTS framework. The evaluation is conducted under realistic missingness scenarios, incorporating both random and structured block missing patterns. Experimental results demonstrate that model performance varies according to the type and extent of missingness: token-based approaches perform well under random missing conditions, while SAITS demonstrates the most robust and consistent performance across all evaluated models, achieving the lowest Mean Absolute Percentage Error (MAPE) at the highest missing rate of 90%, with values of 0.780 under random missingness and 0.750 under block missingness. Notably, SAITS outperforms all other evaluated models at the highest missingness rate of 90%, suggesting that its performance advantage becomes more pronounced as the degree of missing data increases, demonstrating strong robustness under severe data degradation. This work presents a systematic evaluation framework for deep learning-based imputation of multivariate energy pricing time series, offering practical guidance on model selection across varying levels of missingness.

Multivariate time series data plays a critical role in domains such as energy markets, where data completeness and accuracy are essential for reliable analysis and decision-making. However, real-world datasets frequently contain missing values due to factors such as low market liquidity, reporting delays, and arbitrage inconsistencies, which can significantly degrade the performance of data-driven models. Despite advances in deep learning-based imputation, limited work has examined these methods under realistic conditions of missingness specific to energy pricing data. This thesis investigates the effectiveness of various deep learning imputation techniques for multivariate energy pricing time series, intending to develop a robust and reliable imputation pipeline. A range of methods is evaluated, including transformer-based architectures such as SAITS and Transformer, a token-based approach such as TOTEM, and other neural network models, including TSLANet, and an LSTM autoencoder serving as a baseline, with models implemented within the PyPOTS framework. The evaluation is conducted under realistic missingness scenarios, incorporating both random and structured block missing patterns. Experimental results demonstrate that model performance varies according to the type and extent of missingness: token-based approaches perform well under random missing conditions, while SAITS demonstrates the most robust and consistent performance across all evaluated models, achieving the lowest Mean Absolute Percentage Error (MAPE) at the highest missing rate of 90%, with values of 0.780 under random missingness and 0.750 under block missingness. Notably, SAITS outperforms all other evaluated models at the highest missingness rate of 90%, suggesting that its performance advantage becomes more pronounced as the degree of missing data increases, demonstrating strong robustness under severe data degradation. This work presents a systematic evaluation framework for deep learning-based imputation of multivariate energy pricing time series, offering practical guidance on model selection across varying levels of missingness.

A Survey and Experimental Evaluation of Deep Learning Imputation Models for Real-World Energy Pricing Multivariate Time Series

KAUSHAL, APARNA
2025/2026

Abstract

Multivariate time series data plays a critical role in domains such as energy markets, where data completeness and accuracy are essential for reliable analysis and decision-making. However, real-world datasets frequently contain missing values due to factors such as low market liquidity, reporting delays, and arbitrage inconsistencies, which can significantly degrade the performance of data-driven models. Despite advances in deep learning-based imputation, limited work has examined these methods under realistic conditions of missingness specific to energy pricing data. This thesis investigates the effectiveness of various deep learning imputation techniques for multivariate energy pricing time series, intending to develop a robust and reliable imputation pipeline. A range of methods is evaluated, including transformer-based architectures such as SAITS and Transformer, a token-based approach such as TOTEM, and other neural network models, including TSLANet, and an LSTM autoencoder serving as a baseline, with models implemented within the PyPOTS framework. The evaluation is conducted under realistic missingness scenarios, incorporating both random and structured block missing patterns. Experimental results demonstrate that model performance varies according to the type and extent of missingness: token-based approaches perform well under random missing conditions, while SAITS demonstrates the most robust and consistent performance across all evaluated models, achieving the lowest Mean Absolute Percentage Error (MAPE) at the highest missing rate of 90%, with values of 0.780 under random missingness and 0.750 under block missingness. Notably, SAITS outperforms all other evaluated models at the highest missingness rate of 90%, suggesting that its performance advantage becomes more pronounced as the degree of missing data increases, demonstrating strong robustness under severe data degradation. This work presents a systematic evaluation framework for deep learning-based imputation of multivariate energy pricing time series, offering practical guidance on model selection across varying levels of missingness.
2025
A Survey and Experimental Evaluation of Deep Learning Imputation Models for Real-World Energy Pricing Multivariate Time Series
Multivariate time series data plays a critical role in domains such as energy markets, where data completeness and accuracy are essential for reliable analysis and decision-making. However, real-world datasets frequently contain missing values due to factors such as low market liquidity, reporting delays, and arbitrage inconsistencies, which can significantly degrade the performance of data-driven models. Despite advances in deep learning-based imputation, limited work has examined these methods under realistic conditions of missingness specific to energy pricing data. This thesis investigates the effectiveness of various deep learning imputation techniques for multivariate energy pricing time series, intending to develop a robust and reliable imputation pipeline. A range of methods is evaluated, including transformer-based architectures such as SAITS and Transformer, a token-based approach such as TOTEM, and other neural network models, including TSLANet, and an LSTM autoencoder serving as a baseline, with models implemented within the PyPOTS framework. The evaluation is conducted under realistic missingness scenarios, incorporating both random and structured block missing patterns. Experimental results demonstrate that model performance varies according to the type and extent of missingness: token-based approaches perform well under random missing conditions, while SAITS demonstrates the most robust and consistent performance across all evaluated models, achieving the lowest Mean Absolute Percentage Error (MAPE) at the highest missing rate of 90%, with values of 0.780 under random missingness and 0.750 under block missingness. Notably, SAITS outperforms all other evaluated models at the highest missingness rate of 90%, suggesting that its performance advantage becomes more pronounced as the degree of missing data increases, demonstrating strong robustness under severe data degradation. This work presents a systematic evaluation framework for deep learning-based imputation of multivariate energy pricing time series, offering practical guidance on model selection across varying levels of missingness.
Time Series
Missing data
Imputation models
Deep Learning
Multivariate data
File in questo prodotto:
File Dimensione Formato  
AK_Master_Thesis.pdf

Accesso riservato

Dimensione 2.41 MB
Formato Adobe PDF
2.41 MB Adobe PDF

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/110928