Recent advancements in Vision-Language Models (VLMs) have demonstrated significant potential for high-level robotic task planning by mapping multimodal observations and natural language instructions to structured control policies, such as Behavior Trees (BTs). While recent methodologies have successfully utilized synthetic domain-randomized data to fine-tune models for zero-shot sim-to-real transfer, their performance and reliability are primarily validated within the strict bounds of their training distribution. The limits of execution under Out-Of-Distribution (OOD) conditions remain critical and underexplored areas of investigation. This thesis aims to systematically evaluate the generalization properties and structural robustness of instruction-tuned VLMs deployed as symbolic policy synthesizers. By exposing the planning framework to a diverse range of OOD perturbations - such as novel object geometries, unseen semantic tasks, and new hardware primitives - this research investigates how effectively these models adapt to scenarios absent from their foundational training data. The objective is to identify and characterize potential failure modes, including logical breakdowns and functional grounding gaps, when the system faces unfamiliar real-world umpredictable complexity. In conclusion, this work seeks to provide a comprehensive understanding of the limitations inherent in synthetic neuro-symbolic supervision, contributing valuable insights toward the development of more robust and adaptable robotic control architectures.

Recent advancements in Vision-Language Models (VLMs) have demonstrated significant potential for high-level robotic task planning by mapping multimodal observations and natural language instructions to structured control policies, such as Behavior Trees (BTs). While recent methodologies have successfully utilized synthetic domain-randomized data to fine-tune models for zero-shot sim-to-real transfer, their performance and reliability are primarily validated within the strict bounds of their training distribution. The limits of execution under Out-Of-Distribution (OOD) conditions remain critical and underexplored areas of investigation. This thesis aims to systematically evaluate the generalization properties and structural robustness of instruction-tuned VLMs deployed as symbolic policy synthesizers. By exposing the planning framework to a diverse range of OOD perturbations - such as novel object geometries, unseen semantic tasks, and new hardware primitives - this research investigates how effectively these models adapt to scenarios absent from their foundational training data. The objective is to identify and characterize potential failure modes, including logical breakdowns and functional grounding gaps, when the system faces unfamiliar real-world umpredictable complexity. In conclusion, this work seeks to provide a comprehensive understanding of the limitations inherent in synthetic neuro-symbolic supervision, contributing valuable insights toward the development of more robust and adaptable robotic control architectures.

Evaluating Robustness in Instruction-Tuned VLMs for Robotic Planning

ADAMI, SIMONE
2025/2026

Abstract

Recent advancements in Vision-Language Models (VLMs) have demonstrated significant potential for high-level robotic task planning by mapping multimodal observations and natural language instructions to structured control policies, such as Behavior Trees (BTs). While recent methodologies have successfully utilized synthetic domain-randomized data to fine-tune models for zero-shot sim-to-real transfer, their performance and reliability are primarily validated within the strict bounds of their training distribution. The limits of execution under Out-Of-Distribution (OOD) conditions remain critical and underexplored areas of investigation. This thesis aims to systematically evaluate the generalization properties and structural robustness of instruction-tuned VLMs deployed as symbolic policy synthesizers. By exposing the planning framework to a diverse range of OOD perturbations - such as novel object geometries, unseen semantic tasks, and new hardware primitives - this research investigates how effectively these models adapt to scenarios absent from their foundational training data. The objective is to identify and characterize potential failure modes, including logical breakdowns and functional grounding gaps, when the system faces unfamiliar real-world umpredictable complexity. In conclusion, this work seeks to provide a comprehensive understanding of the limitations inherent in synthetic neuro-symbolic supervision, contributing valuable insights toward the development of more robust and adaptable robotic control architectures.
2025
Evaluating Robustness in Instruction-Tuned VLMs for Robotic Planning
Recent advancements in Vision-Language Models (VLMs) have demonstrated significant potential for high-level robotic task planning by mapping multimodal observations and natural language instructions to structured control policies, such as Behavior Trees (BTs). While recent methodologies have successfully utilized synthetic domain-randomized data to fine-tune models for zero-shot sim-to-real transfer, their performance and reliability are primarily validated within the strict bounds of their training distribution. The limits of execution under Out-Of-Distribution (OOD) conditions remain critical and underexplored areas of investigation. This thesis aims to systematically evaluate the generalization properties and structural robustness of instruction-tuned VLMs deployed as symbolic policy synthesizers. By exposing the planning framework to a diverse range of OOD perturbations - such as novel object geometries, unseen semantic tasks, and new hardware primitives - this research investigates how effectively these models adapt to scenarios absent from their foundational training data. The objective is to identify and characterize potential failure modes, including logical breakdowns and functional grounding gaps, when the system faces unfamiliar real-world umpredictable complexity. In conclusion, this work seeks to provide a comprehensive understanding of the limitations inherent in synthetic neuro-symbolic supervision, contributing valuable insights toward the development of more robust and adaptable robotic control architectures.
VLM
Robotic Planning
Behaviour trees
Fine tuning
File in questo prodotto:
File Dimensione Formato  
Adami_Simone.pdf

accesso aperto

Dimensione 1.28 MB
Formato Adobe PDF
1.28 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/114224