Deep reinforcement learning for continuous control is commonly stabilized through replay buffers, mini-batch updates, and target networks. While effective in replay-based off-policy algorithms, these mechanisms are incom- patible with strict streaming settings in which transitions are processed directly and not stored for later reuse. This thesis investigates the performance and stability gap between replay-based and streaming actor–critic learning un- der a shared continuous-control evaluation protocol. The study compares Twin Delayed Deep Deterministic Policy Gradient (TD3), Streaming Actor–Critic with eligibility traces (Streaming AC(λ)), and Action Value Gradient (AVG) on MuJoCo and DeepMind Control Suite tasks. TD3 serves as the replay-based reference method, while Streaming AC(λ) and AVG represent streaming methods that update directly from newly observed transitions. The results show that TD3 achieves the strongest final performance under the implemented protocol. Stream- ing AC(λ) remains functional without replay-based stabilizers, but reaches lower final returns than TD3. AVG shows larger task dependence and higher sensitivity to stabilization choices. Ablation experiments indicate that scaling and normalization can improve individual environments, but do not produce a uniformly robust config- uration. A preliminary pixel-based extension of Streaming AC(λ) shows that learning from visual representations is feasible, but less stable than state-based learning. Overall, the findings show that streaming deep actor–critic learning is possible, but remains less robust than replay-based learning under the evaluated protocol. The results identify replay availability, update scale, normal- ization, and learned representation dynamics as central factors shaping the stability gap between replay-based and streaming deep actor–critic learning.

Deep reinforcement learning for continuous control is commonly stabilized through replay buffers, mini-batch updates, and target networks. While effective in replay-based off-policy algorithms, these mechanisms are incom- patible with strict streaming settings in which transitions are processed directly and not stored for later reuse. This thesis investigates the performance and stability gap between replay-based and streaming actor–critic learning un- der a shared continuous-control evaluation protocol. The study compares Twin Delayed Deep Deterministic Policy Gradient (TD3), Streaming Actor–Critic with eligibility traces (Streaming AC(λ)), and Action Value Gradient (AVG) on MuJoCo and DeepMind Control Suite tasks. TD3 serves as the replay-based reference method, while Streaming AC(λ) and AVG represent streaming methods that update directly from newly observed transitions. The results show that TD3 achieves the strongest final performance under the implemented protocol. Stream- ing AC(λ) remains functional without replay-based stabilizers, but reaches lower final returns than TD3. AVG shows larger task dependence and higher sensitivity to stabilization choices. Ablation experiments indicate that scaling and normalization can improve individual environments, but do not produce a uniformly robust config- uration. A preliminary pixel-based extension of Streaming AC(λ) shows that learning from visual representations is feasible, but less stable than state-based learning. Overall, the findings show that streaming deep actor–critic learning is possible, but remains less robust than replay-based learning under the evaluated protocol. The results identify replay availability, update scale, normal- ization, and learned representation dynamics as central factors shaping the stability gap between replay-based and streaming deep actor–critic learning.

Stability in Streaming Deep Reinforcement Learning: A Study of Scaling, Normalization, and Visual Representation.

HORN, MAX HANS JURGEN HENRY
2025/2026

Abstract

Deep reinforcement learning for continuous control is commonly stabilized through replay buffers, mini-batch updates, and target networks. While effective in replay-based off-policy algorithms, these mechanisms are incom- patible with strict streaming settings in which transitions are processed directly and not stored for later reuse. This thesis investigates the performance and stability gap between replay-based and streaming actor–critic learning un- der a shared continuous-control evaluation protocol. The study compares Twin Delayed Deep Deterministic Policy Gradient (TD3), Streaming Actor–Critic with eligibility traces (Streaming AC(λ)), and Action Value Gradient (AVG) on MuJoCo and DeepMind Control Suite tasks. TD3 serves as the replay-based reference method, while Streaming AC(λ) and AVG represent streaming methods that update directly from newly observed transitions. The results show that TD3 achieves the strongest final performance under the implemented protocol. Stream- ing AC(λ) remains functional without replay-based stabilizers, but reaches lower final returns than TD3. AVG shows larger task dependence and higher sensitivity to stabilization choices. Ablation experiments indicate that scaling and normalization can improve individual environments, but do not produce a uniformly robust config- uration. A preliminary pixel-based extension of Streaming AC(λ) shows that learning from visual representations is feasible, but less stable than state-based learning. Overall, the findings show that streaming deep actor–critic learning is possible, but remains less robust than replay-based learning under the evaluated protocol. The results identify replay availability, update scale, normal- ization, and learned representation dynamics as central factors shaping the stability gap between replay-based and streaming deep actor–critic learning.
2025
Stability in Streaming Deep Reinforcement Learning: A Study of Scaling, Normalization, and Visual Representation.
Deep reinforcement learning for continuous control is commonly stabilized through replay buffers, mini-batch updates, and target networks. While effective in replay-based off-policy algorithms, these mechanisms are incom- patible with strict streaming settings in which transitions are processed directly and not stored for later reuse. This thesis investigates the performance and stability gap between replay-based and streaming actor–critic learning un- der a shared continuous-control evaluation protocol. The study compares Twin Delayed Deep Deterministic Policy Gradient (TD3), Streaming Actor–Critic with eligibility traces (Streaming AC(λ)), and Action Value Gradient (AVG) on MuJoCo and DeepMind Control Suite tasks. TD3 serves as the replay-based reference method, while Streaming AC(λ) and AVG represent streaming methods that update directly from newly observed transitions. The results show that TD3 achieves the strongest final performance under the implemented protocol. Stream- ing AC(λ) remains functional without replay-based stabilizers, but reaches lower final returns than TD3. AVG shows larger task dependence and higher sensitivity to stabilization choices. Ablation experiments indicate that scaling and normalization can improve individual environments, but do not produce a uniformly robust config- uration. A preliminary pixel-based extension of Streaming AC(λ) shows that learning from visual representations is feasible, but less stable than state-based learning. Overall, the findings show that streaming deep actor–critic learning is possible, but remains less robust than replay-based learning under the evaluated protocol. The results identify replay availability, update scale, normal- ization, and learned representation dynamics as central factors shaping the stability gap between replay-based and streaming deep actor–critic learning.
Neural Networks
Data Science
Actor Critic
AVG
Computer Vision
File in questo prodotto:
File Dimensione Formato  
Horn_Max_2143471_Master_Thesis.pdf

accesso aperto

Dimensione 2.32 MB
Formato Adobe PDF
2.32 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/110925