This thesis presents the implementation and evaluation of PILCO (Probabilistic Inference for Learning Control), a model-based policy search algorithm applied to precision thermal control on the TCLab (Temperature Control Lab) hardware system. The central problem addressed is the sample inefficiency typical of model-free reinforcement learning algorithms, which require thousands of interactions with the environment before convergence, making them difficult to apply to real physical systems subject to wear and time costs. PILCO addresses this limitation using a probabilistic dynamics model based on Gaussian Processes (GP), which allows model uncertainty to be quantified and propagated analytically during trajectory planning.

L'elaborato presenta l'implementazione e la valutazione di PILCO (Probabilistic Inference for Learning Control), un algoritmo di policy search model-based applicato al controllo termico di precisione sul sistema hardware TCLab (Temperature Control Lab). Il problema centrale affrontato è la sample inefficiency tipica degli algoritmi di Reinforcement Learning model-free, che richiedono migliaia di interazioni con l'ambiente prima di convergere, rendendone difficile l'applicazione su sistemi fisici reali soggetti a usura e costi temporali. PILCO risolve tale limitazione utilizzando un modello probabilistico della dinamica basato su Gaussian Processes (GP), che consente di quantificare l'incertezza del modello e di propagarla analiticamente durante la pianificazione delle traiettorie.

Controllo della temperatura in un dispositivo TCLab mediante un algoritmo di tipo PILCO-RL

BORTOLETTO, ALBERTO
2025/2026

Abstract

This thesis presents the implementation and evaluation of PILCO (Probabilistic Inference for Learning Control), a model-based policy search algorithm applied to precision thermal control on the TCLab (Temperature Control Lab) hardware system. The central problem addressed is the sample inefficiency typical of model-free reinforcement learning algorithms, which require thousands of interactions with the environment before convergence, making them difficult to apply to real physical systems subject to wear and time costs. PILCO addresses this limitation using a probabilistic dynamics model based on Gaussian Processes (GP), which allows model uncertainty to be quantified and propagated analytically during trajectory planning.
2025
Temperature control in a TCLab device using a PILCO-RL type algorithm
L'elaborato presenta l'implementazione e la valutazione di PILCO (Probabilistic Inference for Learning Control), un algoritmo di policy search model-based applicato al controllo termico di precisione sul sistema hardware TCLab (Temperature Control Lab). Il problema centrale affrontato è la sample inefficiency tipica degli algoritmi di Reinforcement Learning model-free, che richiedono migliaia di interazioni con l'ambiente prima di convergere, rendendone difficile l'applicazione su sistemi fisici reali soggetti a usura e costi temporali. PILCO risolve tale limitazione utilizzando un modello probabilistico della dinamica basato su Gaussian Processes (GP), che consente di quantificare l'incertezza del modello e di propagarla analiticamente durante la pianificazione delle traiettorie.
RL
PILCO
Temperature Control
Controllo automatico
MATLAB
File in questo prodotto:
File Dimensione Formato  
Bortoletto_Alberto.pdf

accesso aperto

Dimensione 1.75 MB
Formato Adobe PDF
1.75 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/111138