Multi-agent path finding under partial observability requires agents to coordinate from limited local infor- mation. Recent learning-based planners address this by allowing agents to exchange learned messages, but they assume unconstrained communication: every agent transmits at every timestep, regardless of whether the message affects any recipient’s decision. This is wasteful under realistic bandwidth constraints, and it worsens the chatter problem as team size grows. This thesis studies goal-oriented communication for decentralized multi-agent path finding, where agents transmit only when transmission serves the goal. We build on SCRIMP, a decentralized policy that couples a small 3 × 3 field of view with a transformer-based communication block, trained through centralized- training/decentralized-execution using Proximal Policy Optimization and imitation learning. We extend it with a learned communication gate, realized as an additional output head that reads the same fused repre- sentation as the policy and value heads, so that each agent decides whether to transmit after having received the messages of its teammates. The decision is sampled from a categorical distribution and trained by a clipped surrogate objective sharing the advantage estimate of the navigation policy, with a penalty applied to the probability of transmitting rather than to the environment reward. A message cache held at the com- munication block supplies an agent’s most recent transmission whenever it stays silent, so that the attention computation receives one entry per agent and no bandwidth is consumed. Two training strategies are evaluated for the tendency of agents to suppress communication before their messages become informative: a curriculum that introduces the penalty only after navigation and coordina- tion have been learned, and an alternation between penalized and unpenalized episodes in which the policy observes which regime applies. All experiments are conducted in a Python simulation environment and compared against unmodified SCRIMP across fifteen scenarios spanning team sizes from 8 to 128 agents and obstacle densities up to 30%. Results show that communication can be reduced by more than 97% while the success rate is preserved across the vast majority of scenarios, indicating that the great majority of the messages exchanged under an unconditional scheme carry no bearing upon the decisions of their recipients.

Un Approccio alla Comunicazione Orientato agli Obiettivi per Robot Mobili Autonomi

ALY, MAHMOUD MOHAMED SHAABAN MOHAMED SHAABAN
2025/2026

Abstract

Multi-agent path finding under partial observability requires agents to coordinate from limited local infor- mation. Recent learning-based planners address this by allowing agents to exchange learned messages, but they assume unconstrained communication: every agent transmits at every timestep, regardless of whether the message affects any recipient’s decision. This is wasteful under realistic bandwidth constraints, and it worsens the chatter problem as team size grows. This thesis studies goal-oriented communication for decentralized multi-agent path finding, where agents transmit only when transmission serves the goal. We build on SCRIMP, a decentralized policy that couples a small 3 × 3 field of view with a transformer-based communication block, trained through centralized- training/decentralized-execution using Proximal Policy Optimization and imitation learning. We extend it with a learned communication gate, realized as an additional output head that reads the same fused repre- sentation as the policy and value heads, so that each agent decides whether to transmit after having received the messages of its teammates. The decision is sampled from a categorical distribution and trained by a clipped surrogate objective sharing the advantage estimate of the navigation policy, with a penalty applied to the probability of transmitting rather than to the environment reward. A message cache held at the com- munication block supplies an agent’s most recent transmission whenever it stays silent, so that the attention computation receives one entry per agent and no bandwidth is consumed. Two training strategies are evaluated for the tendency of agents to suppress communication before their messages become informative: a curriculum that introduces the penalty only after navigation and coordina- tion have been learned, and an alternation between penalized and unpenalized episodes in which the policy observes which regime applies. All experiments are conducted in a Python simulation environment and compared against unmodified SCRIMP across fifteen scenarios spanning team sizes from 8 to 128 agents and obstacle densities up to 30%. Results show that communication can be reduced by more than 97% while the success rate is preserved across the vast majority of scenarios, indicating that the great majority of the messages exchanged under an unconditional scheme carry no bearing upon the decisions of their recipients.
2025
A Goal-Oriented Communication approach for autonomous mobile robots
Goal-oriented Comm.
Semantic Comm.
Multi-Agent systems
safe navigation
Path Planning
File in questo prodotto:
File Dimensione Formato  
Aly_MahmoudMohamedShaabanMohamedShaaban.pdf

accesso aperto

Dimensione 831.73 kB
Formato Adobe PDF
831.73 kB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/112989