Current traffic systems adapt poorly to dynamic demand. While Reinforcement Learn- ing provides a solution, multi-agent approaches struggle on real-world networks with di- verse intersection topologies. Independent policies scale linearly and isolate knowledge, whereas parameter-shared architectures demand rigid, fixed-dimensional spaces that in- herently exclude heterogeneous networks. This thesis presents AnyLight, a generalizable Multi-Agent Reinforcement Learning ar- chitecture for heterogeneous Traffic Signal Control, featuring three contributions. First, a movement-centric state representation ensures semantically aligned inputs across diverse junctions. Second, universal parameter-sharing enables a single Proximal Policy Opti- mization network to govern all intersections via dynamic padding and masking. Third, a cross-attention decoder and centralized critic exploit privileged neighbor-action informa- tion during Centralized Training, Decentralized Execution. AnyLight is evaluated on six synthetic and real-world benchmark scenarios from the RESCO and MA2C suites. Compared to classical heuristics and learning-based baselines, AnyLight consistently demonstrates superior performance by minimizing intersection de- lays during high-density flows. Ablation studies confirm that combining the movement- centric representation, universal parameter sharing, and collaborative reward is crucial for scaling to complex urban environments.
Current traffic systems adapt poorly to dynamic demand. While Reinforcement Learn- ing provides a solution, multi-agent approaches struggle on real-world networks with di- verse intersection topologies. Independent policies scale linearly and isolate knowledge, whereas parameter-shared architectures demand rigid, fixed-dimensional spaces that in- herently exclude heterogeneous networks. This thesis presents AnyLight, a generalizable Multi-Agent Reinforcement Learning ar- chitecture for heterogeneous Traffic Signal Control, featuring three contributions. First, a movement-centric state representation ensures semantically aligned inputs across diverse junctions. Second, universal parameter-sharing enables a single Proximal Policy Opti- mization network to govern all intersections via dynamic padding and masking. Third, a cross-attention decoder and centralized critic exploit privileged neighbor-action informa- tion during Centralized Training, Decentralized Execution. AnyLight is evaluated on six synthetic and real-world benchmark scenarios from the RESCO and MA2C suites. Compared to classical heuristics and learning-based baselines, AnyLight consistently demonstrates superior performance by minimizing intersection de- lays during high-density flows. Ablation studies confirm that combining the movement- centric representation, universal parameter sharing, and collaborative reward is crucial for scaling to complex urban environments.
AnyLight: A Generalizable Multi-Agent Reinforcement Learning Architecture for Heterogeneous Traffic Networks
BUSTAFFA, MARCO
2025/2026
Abstract
Current traffic systems adapt poorly to dynamic demand. While Reinforcement Learn- ing provides a solution, multi-agent approaches struggle on real-world networks with di- verse intersection topologies. Independent policies scale linearly and isolate knowledge, whereas parameter-shared architectures demand rigid, fixed-dimensional spaces that in- herently exclude heterogeneous networks. This thesis presents AnyLight, a generalizable Multi-Agent Reinforcement Learning ar- chitecture for heterogeneous Traffic Signal Control, featuring three contributions. First, a movement-centric state representation ensures semantically aligned inputs across diverse junctions. Second, universal parameter-sharing enables a single Proximal Policy Opti- mization network to govern all intersections via dynamic padding and masking. Third, a cross-attention decoder and centralized critic exploit privileged neighbor-action informa- tion during Centralized Training, Decentralized Execution. AnyLight is evaluated on six synthetic and real-world benchmark scenarios from the RESCO and MA2C suites. Compared to classical heuristics and learning-based baselines, AnyLight consistently demonstrates superior performance by minimizing intersection de- lays during high-density flows. Ablation studies confirm that combining the movement- centric representation, universal parameter sharing, and collaborative reward is crucial for scaling to complex urban environments.| File | Dimensione | Formato | |
|---|---|---|---|
|
AnyLight.pdf
accesso aperto
Dimensione
11.05 MB
Formato
Adobe PDF
|
11.05 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110920