This thesis presents a comprehensive framework for Anomaly Detection and Root Cause Analysis (RCA) on multivariate time-series data from industrial hypoxic air generation plants, used for maintaining controlled atmospheres. The lack of exhaustively labeled historical datasets for every possible failure mode often renders classical supervised approaches inapplicable. The study systematically compares six distinct detection paradigms, spanning statistical methods, traditional machine learning, and deep learning techniques. Specifically, the implemented approaches include Z-Score, One-Class SVM, Isolation Forest, Autoencoders, Variational Autoencoders, and Long Short-Term Memory Autoencoders (LSTM-AE). To overcome the inherent validation challenges of unsupervised models, these approaches were first evaluated using synthetic anomaly injection to establish an objective ground truth. Subsequently, the models were tested and compared against real-world scenarios involving physical machine tampering. The primary contribution of this project is enhancing model explainability to overcome the typical "black-box" nature of artificial intelligence. This is achieved by directly integrating SHAP (SHapley Additive exPlanations) values to interpret the models' outputs. Through this interpretability layer, the proposed system extends beyond binary anomaly detection to assist in localizing the most probable triggering factors. Ultimately, this framework demonstrates how raw industrial data can be leveraged to extract interpretable insights, offering a structured decision-support mechanism for proactive maintenance strategies.
This thesis presents a comprehensive framework for Anomaly Detection and Root Cause Analysis (RCA) on multivariate time-series data from industrial hypoxic air generation plants, used for maintaining controlled atmospheres. The lack of exhaustively labeled historical datasets for every possible failure mode often renders classical supervised approaches inapplicable. The study systematically compares six distinct detection paradigms, spanning statistical methods, traditional machine learning, and deep learning techniques. Specifically, the implemented approaches include Z-Score, One-Class SVM, Isolation Forest, Autoencoders, Variational Autoencoders, and Long Short-Term Memory Autoencoders (LSTM-AE). To overcome the inherent validation challenges of unsupervised models, these approaches were first evaluated using synthetic anomaly injection to establish an objective ground truth. Subsequently, the models were tested and compared against real-world scenarios involving physical machine tampering. The primary contribution of this project is enhancing model explainability to overcome the typical "black-box" nature of artificial intelligence. This is achieved by directly integrating SHAP (SHapley Additive exPlanations) values to interpret the models' outputs. Through this interpretability layer, the proposed system extends beyond binary anomaly detection to assist in localizing the most probable triggering factors. Ultimately, this framework demonstrates how raw industrial data can be leveraged to extract interpretable insights, offering a structured decision-support mechanism for proactive maintenance strategies.
Anomaly Detection and Root Cause Analysis on Industrial Sensor Data: A Multi-Model Comparison with SHAP Interpretability
GIORGIO, MARCO
2025/2026
Abstract
This thesis presents a comprehensive framework for Anomaly Detection and Root Cause Analysis (RCA) on multivariate time-series data from industrial hypoxic air generation plants, used for maintaining controlled atmospheres. The lack of exhaustively labeled historical datasets for every possible failure mode often renders classical supervised approaches inapplicable. The study systematically compares six distinct detection paradigms, spanning statistical methods, traditional machine learning, and deep learning techniques. Specifically, the implemented approaches include Z-Score, One-Class SVM, Isolation Forest, Autoencoders, Variational Autoencoders, and Long Short-Term Memory Autoencoders (LSTM-AE). To overcome the inherent validation challenges of unsupervised models, these approaches were first evaluated using synthetic anomaly injection to establish an objective ground truth. Subsequently, the models were tested and compared against real-world scenarios involving physical machine tampering. The primary contribution of this project is enhancing model explainability to overcome the typical "black-box" nature of artificial intelligence. This is achieved by directly integrating SHAP (SHapley Additive exPlanations) values to interpret the models' outputs. Through this interpretability layer, the proposed system extends beyond binary anomaly detection to assist in localizing the most probable triggering factors. Ultimately, this framework demonstrates how raw industrial data can be leveraged to extract interpretable insights, offering a structured decision-support mechanism for proactive maintenance strategies.| File | Dimensione | Formato | |
|---|---|---|---|
|
Giorgio_Marco_2103675.pdf
accesso aperto
Dimensione
1.92 MB
Formato
Adobe PDF
|
1.92 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/115838