This thesis investigates the use of deep reinforcement learning for time-optimal autonomous driving of miniature vehicles using image-based observations. The driving task was formulated in the Duckietown environment, where the agent was trained in simulation and evaluated both in simulation and on a physical Duckiebot platform. The objective was to learn policies capable of completing laps efficiently while maintaining stable driving behavior and remaining within the drivable area. Soft Actor-Critic (SAC) and Twin Delayed Deep Deterministic Policy Gradient (TD3) were implemented for continuous control using stacked camera observations. Different observation representations, reward formulations and sim-to-real transfer techniques were evaluated. The reward design was developed to balance forward progress, lap completion, driving stability and time-optimal behavior. In addition, visual and dynamics randomization techniques were applied during training to improve robustness to real-world deployment conditions. The results show that TD3 was able to learn stable lap completion behavior in simulation using both RGB and grayscale observations. A time-optimal MPC baseline achieved faster lap times in simulation, as expected due to its access to vehicle state information, track geometry and a prediction model. However, the reinforcement learning policies relied only on camera observations and did not require explicit localization or predefined trajectories. Real-world experiments showed that transferring policies from simulation to the physical Duckiebot remains challenging, but policies trained with visual distortions achieved better transfer performance than those trained without distortions. Overall, the thesis shows that image based deep reinforcement learning can be used to learn autonomous driving policies for miniature vehicles. However, robust time-optimal real-world driving remains limited by the sim-to-real gap, particularly when the learned policy operates close to the track boundaries.

Deep Reinforcement Learning for Time-Optimal Autonomous Driving in a Duckietown Environment

MYRHAUG, BJOERN MAGNUS
2025/2026

Abstract

This thesis investigates the use of deep reinforcement learning for time-optimal autonomous driving of miniature vehicles using image-based observations. The driving task was formulated in the Duckietown environment, where the agent was trained in simulation and evaluated both in simulation and on a physical Duckiebot platform. The objective was to learn policies capable of completing laps efficiently while maintaining stable driving behavior and remaining within the drivable area. Soft Actor-Critic (SAC) and Twin Delayed Deep Deterministic Policy Gradient (TD3) were implemented for continuous control using stacked camera observations. Different observation representations, reward formulations and sim-to-real transfer techniques were evaluated. The reward design was developed to balance forward progress, lap completion, driving stability and time-optimal behavior. In addition, visual and dynamics randomization techniques were applied during training to improve robustness to real-world deployment conditions. The results show that TD3 was able to learn stable lap completion behavior in simulation using both RGB and grayscale observations. A time-optimal MPC baseline achieved faster lap times in simulation, as expected due to its access to vehicle state information, track geometry and a prediction model. However, the reinforcement learning policies relied only on camera observations and did not require explicit localization or predefined trajectories. Real-world experiments showed that transferring policies from simulation to the physical Duckiebot remains challenging, but policies trained with visual distortions achieved better transfer performance than those trained without distortions. Overall, the thesis shows that image based deep reinforcement learning can be used to learn autonomous driving policies for miniature vehicles. However, robust time-optimal real-world driving remains limited by the sim-to-real gap, particularly when the learned policy operates close to the track boundaries.
2025
Deep Reinforcement Learning for Time-Optimal Autonomous Driving in a Duckietown Environment
RL
Duckietown
Autonomous driving
Time-optimal
File in questo prodotto:
File Dimensione Formato  
Myrhuag_BjoernMagnus.pdf

accesso aperto

Dimensione 9.8 MB
Formato Adobe PDF
9.8 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/109373