Visual end-to-end deep reinforcement learning (DRL) is a powerful method for autonomous vehicle control, but closing the sim-to-real "reality gap" remains a major engineering challenge. Formulated as a partially observable Markov decision process (POMDP), this thesis presents a detailed empirical ablation study that tests Soft Actor-Critic (SAC) and Twin Delayed Deep Deterministic Policy Gradient (TD3) on a miniature differential-drive robot. The models are evaluated over 1,500,000 training steps per configuration in simulation before being deployed directly onto physical hardware powered by an NVIDIA Jetson Nano. To find the exact components needed for robust zero-shot transfer, we systematically remove and test a variety of sensorimotor wrappers designed to match real-world hardware limits. We evaluate image horizon cropping, temporal frame stacking, highly downscaled spatial resolutions to improve feature alignment, kinematically constrained steering spaces, input latency simulations, a stateful recovery training mechanism, and multi-objective rewards that penalize lane deviation and high-frequency steering jerk. Furthermore, we study the transfer pipeline by comparing standard domain randomization against a progressive curriculum learning strategy that gradually introduces visual and physical noise over time. Our experimental results show a clear trade-off: design choices that maximize sample efficiency in ideal simulation settings consistently fail when deployed on the real robot. Conversely, we demonstrate that a constrained action space, managed latencies, a margin-recovery tool, and a performance-driven curriculum are essential for real-world robustness, ultimately achieving a 100% physical lap completion rate. This breakdown provides a clear, practical engineering guide for choosing optimal wrappers and data augmentation schedules in embedded visual robotic control.
From Simulation to Real-World Driving: Visual Reinforcement Learning for a Miniature Autonomous Vehicle
ESMAEILI NASAB LAHIJAN, ALI
2025/2026
Abstract
Visual end-to-end deep reinforcement learning (DRL) is a powerful method for autonomous vehicle control, but closing the sim-to-real "reality gap" remains a major engineering challenge. Formulated as a partially observable Markov decision process (POMDP), this thesis presents a detailed empirical ablation study that tests Soft Actor-Critic (SAC) and Twin Delayed Deep Deterministic Policy Gradient (TD3) on a miniature differential-drive robot. The models are evaluated over 1,500,000 training steps per configuration in simulation before being deployed directly onto physical hardware powered by an NVIDIA Jetson Nano. To find the exact components needed for robust zero-shot transfer, we systematically remove and test a variety of sensorimotor wrappers designed to match real-world hardware limits. We evaluate image horizon cropping, temporal frame stacking, highly downscaled spatial resolutions to improve feature alignment, kinematically constrained steering spaces, input latency simulations, a stateful recovery training mechanism, and multi-objective rewards that penalize lane deviation and high-frequency steering jerk. Furthermore, we study the transfer pipeline by comparing standard domain randomization against a progressive curriculum learning strategy that gradually introduces visual and physical noise over time. Our experimental results show a clear trade-off: design choices that maximize sample efficiency in ideal simulation settings consistently fail when deployed on the real robot. Conversely, we demonstrate that a constrained action space, managed latencies, a margin-recovery tool, and a performance-driven curriculum are essential for real-world robustness, ultimately achieving a 100% physical lap completion rate. This breakdown provides a clear, practical engineering guide for choosing optimal wrappers and data augmentation schedules in embedded visual robotic control.| File | Dimensione | Formato | |
|---|---|---|---|
|
EsmaeiliNasablahijan_Ali.pdf
accesso aperto
Dimensione
4.36 MB
Formato
Adobe PDF
|
4.36 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/109371