Urban traffic signal control must adapt to changing demand while balancing the needs of multiple road users. Fixed-time and rule-based controllers are reliable, but they can perform poorly when traffic conditions vary across a network or when pedestrian crossings introduce additional constraints. This thesis studies adaptive traffic signal control as a reinforcement learning (RL) problem with multi-agent coordination in a simulated urban setting, with particular attention to the trade-off between vehicle efficiency and pedestrian fairness. The work develops a reproducible experimental pipeline based on a Simulation of Urban MObility (SUMO) 2$\times$2 grid of four signalized intersections. Vehicle and pedestrian demand scenarios are generated at low, medium, and high intensity levels using fixed seeds and held-out test routes, so that all controllers can be compared under common conditions. Four Proximal Policy Optimization (PPO)-based control formulations are implemented and evaluated against the conventional SUMO controller: centralized PPO, independent PPO, shared-policy PPO, and Multi-Agent PPO (MAPPO). The study first considers a vehicle-only objective and then extends the environment to a multimodal setting in which pedestrians are simulated natively in SUMO and the reward includes vehicle delay, pedestrian delay, maximum-wait, and starvation terms. The results show that all RL controllers substantially improve over the SUMO baseline in the vehicle-only experiments, reducing waiting time, queue length, and travel time across the tested demand levels. However, the pedestrian-aware experiments reveal that the relative ranking of the architectures changes when the objective becomes multimodal. Centralized PPO degrades in more demanding pedestrian scenarios, while shared-policy PPO and MAPPO provide the strongest overall trade-off between vehicle performance, pedestrian service, and seed-to-seed stability. These findings indicate that adaptive traffic signal controllers should not be evaluated only through vehicle-side metrics: architectural choices and reward design interact strongly with multimodal traffic objectives. The thesis therefore provides both a reusable benchmark pipeline and empirical evidence that pedestrian-aware evaluation is essential for realistic RL-based traffic signal control.
Urban traffic signal control must adapt to changing demand while balancing the needs of multiple road users. Fixed-time and rule-based controllers are reliable, but they can perform poorly when traffic conditions vary across a network or when pedestrian crossings introduce additional constraints. This thesis studies adaptive traffic signal control as a reinforcement learning (RL) problem with multi-agent coordination in a simulated urban setting, with particular attention to the trade-off between vehicle efficiency and pedestrian fairness. The work develops a reproducible experimental pipeline based on a Simulation of Urban MObility (SUMO) 2$\times$2 grid of four signalized intersections. Vehicle and pedestrian demand scenarios are generated at low, medium, and high intensity levels using fixed seeds and held-out test routes, so that all controllers can be compared under common conditions. Four Proximal Policy Optimization (PPO)-based control formulations are implemented and evaluated against the conventional SUMO controller: centralized PPO, independent PPO, shared-policy PPO, and Multi-Agent PPO (MAPPO). The study first considers a vehicle-only objective and then extends the environment to a multimodal setting in which pedestrians are simulated natively in SUMO and the reward includes vehicle delay, pedestrian delay, maximum-wait, and starvation terms. The results show that all RL controllers substantially improve over the SUMO baseline in the vehicle-only experiments, reducing waiting time, queue length, and travel time across the tested demand levels. However, the pedestrian-aware experiments reveal that the relative ranking of the architectures changes when the objective becomes multimodal. Centralized PPO degrades in more demanding pedestrian scenarios, while shared-policy PPO and MAPPO provide the strongest overall trade-off between vehicle performance, pedestrian service, and seed-to-seed stability. These findings indicate that adaptive traffic signal controllers should not be evaluated only through vehicle-side metrics: architectural choices and reward design interact strongly with multimodal traffic objectives. The thesis therefore provides both a reusable benchmark pipeline and empirical evidence that pedestrian-aware evaluation is essential for realistic RL-based traffic signal control.
Adaptive Traffic Signal Control via Multi-Agent Reinforcement Learning: Balancing Vehicle and Pedestrian Flows
FERRARI, DAVIDE
2025/2026
Abstract
Urban traffic signal control must adapt to changing demand while balancing the needs of multiple road users. Fixed-time and rule-based controllers are reliable, but they can perform poorly when traffic conditions vary across a network or when pedestrian crossings introduce additional constraints. This thesis studies adaptive traffic signal control as a reinforcement learning (RL) problem with multi-agent coordination in a simulated urban setting, with particular attention to the trade-off between vehicle efficiency and pedestrian fairness. The work develops a reproducible experimental pipeline based on a Simulation of Urban MObility (SUMO) 2$\times$2 grid of four signalized intersections. Vehicle and pedestrian demand scenarios are generated at low, medium, and high intensity levels using fixed seeds and held-out test routes, so that all controllers can be compared under common conditions. Four Proximal Policy Optimization (PPO)-based control formulations are implemented and evaluated against the conventional SUMO controller: centralized PPO, independent PPO, shared-policy PPO, and Multi-Agent PPO (MAPPO). The study first considers a vehicle-only objective and then extends the environment to a multimodal setting in which pedestrians are simulated natively in SUMO and the reward includes vehicle delay, pedestrian delay, maximum-wait, and starvation terms. The results show that all RL controllers substantially improve over the SUMO baseline in the vehicle-only experiments, reducing waiting time, queue length, and travel time across the tested demand levels. However, the pedestrian-aware experiments reveal that the relative ranking of the architectures changes when the objective becomes multimodal. Centralized PPO degrades in more demanding pedestrian scenarios, while shared-policy PPO and MAPPO provide the strongest overall trade-off between vehicle performance, pedestrian service, and seed-to-seed stability. These findings indicate that adaptive traffic signal controllers should not be evaluated only through vehicle-side metrics: architectural choices and reward design interact strongly with multimodal traffic objectives. The thesis therefore provides both a reusable benchmark pipeline and empirical evidence that pedestrian-aware evaluation is essential for realistic RL-based traffic signal control.| File | Dimensione | Formato | |
|---|---|---|---|
|
Ferrari_Davide.pdf
Accesso riservato
Dimensione
1.68 MB
Formato
Adobe PDF
|
1.68 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110014