Vision-Language-Action (VLA) models have demonstrated strong generalisation capabilities in robotic manipulation, yet their action generation relies primarily on visual input, which is often degraded or entirely unavailable during contact-rich interactions such as grasping and in-hand manipulation. This thesis investigates whether a frozen VLA can maintain task performance under degraded or absent vision by learning a lightweight projection module, f_theta, that maps tactile embeddings into the policy's frozen visual latent space without modifying the pretrained policy weights. Rather than committing immediately to full-scale training, the work follows a feasibility-first methodology designed to establish whether such an approach is well-founded. First, a representational similarity analysis of tactile and visual embedding spaces is conducted on a small exploratory dataset, leading to the evidence-based selection of a generic ViT-Small tactile backbone over the official TVL encoder. Second, a fully reproducible benchmark is developed by integrating a custom DIGIT-equipped Panda gripper into the LIBERO/robosuite/MuJoCo simulation stack, extending TACTO's tactile rendering pipeline to a new simulator. This integration also reveals an independent mechanical confound: the gripper modification alone reduces the frozen policy's success rate on libero_object from 98.2% to 56.8%, before any tactile alignment is introduced. Third, the projection module is trained and evaluated under a leave-one-object-out protocol against a naive baseline. Although it achieves a mean cross-object cosine similarity of 0.852, it does not consistently outperform the baseline, providing inconclusive evidence of cross-object generalisation at the scale of the available data. Taken together, these findings support a qualified rather than unqualified answer to the central hypothesis of this thesis. Tactile and visual embeddings appear sufficiently compatible, in aggregate, to justify learning a shared representation, yet generalisation to unseen objects remains unproven, while an independent mechanical confound limits policy performance regardless of the alignment strategy. Overall, the principal contribution of this thesis is a validated experimental infrastructure and a rigorous feasibility study, rather than a demonstrated end-to-end capability.

Vision-Language-Action (VLA) models have demonstrated strong generalisation capabilities in robotic manipulation, yet their action generation relies primarily on visual input, which is often degraded or entirely unavailable during contact-rich interactions such as grasping and in-hand manipulation. This thesis investigates whether a frozen VLA can maintain task performance under degraded or absent vision by learning a lightweight projection module, f_theta, that maps tactile embeddings into the policy's frozen visual latent space without modifying the pretrained policy weights. Rather than committing immediately to full-scale training, the work follows a feasibility-first methodology designed to establish whether such an approach is well-founded. First, a representational similarity analysis of tactile and visual embedding spaces is conducted on a small exploratory dataset, leading to the evidence-based selection of a generic ViT-Small tactile backbone over the official TVL encoder. Second, a fully reproducible benchmark is developed by integrating a custom DIGIT-equipped Panda gripper into the LIBERO/robosuite/MuJoCo simulation stack, extending TACTO's tactile rendering pipeline to a new simulator. This integration also reveals an independent mechanical confound: the gripper modification alone reduces the frozen policy's success rate on libero_object from 98.2% to 56.8%, before any tactile alignment is introduced. Third, the projection module is trained and evaluated under a leave-one-object-out protocol against a naive baseline. Although it achieves a mean cross-object cosine similarity of 0.852, it does not consistently outperform the baseline, providing inconclusive evidence of cross-object generalisation at the scale of the available data. Taken together, these findings support a qualified rather than unqualified answer to the central hypothesis of this thesis. Tactile and visual embeddings appear sufficiently compatible, in aggregate, to justify learning a shared representation, yet generalisation to unseen objects remains unproven, while an independent mechanical confound limits policy performance regardless of the alignment strategy. Overall, the principal contribution of this thesis is a validated experimental infrastructure and a rigorous feasibility study, rather than a demonstrated end-to-end capability.

Tactile-to-Visual Representation Learning for Vision-Language-Action Models

GIRARDELLO, NICOLO'
2025/2026

Abstract

Vision-Language-Action (VLA) models have demonstrated strong generalisation capabilities in robotic manipulation, yet their action generation relies primarily on visual input, which is often degraded or entirely unavailable during contact-rich interactions such as grasping and in-hand manipulation. This thesis investigates whether a frozen VLA can maintain task performance under degraded or absent vision by learning a lightweight projection module, f_theta, that maps tactile embeddings into the policy's frozen visual latent space without modifying the pretrained policy weights. Rather than committing immediately to full-scale training, the work follows a feasibility-first methodology designed to establish whether such an approach is well-founded. First, a representational similarity analysis of tactile and visual embedding spaces is conducted on a small exploratory dataset, leading to the evidence-based selection of a generic ViT-Small tactile backbone over the official TVL encoder. Second, a fully reproducible benchmark is developed by integrating a custom DIGIT-equipped Panda gripper into the LIBERO/robosuite/MuJoCo simulation stack, extending TACTO's tactile rendering pipeline to a new simulator. This integration also reveals an independent mechanical confound: the gripper modification alone reduces the frozen policy's success rate on libero_object from 98.2% to 56.8%, before any tactile alignment is introduced. Third, the projection module is trained and evaluated under a leave-one-object-out protocol against a naive baseline. Although it achieves a mean cross-object cosine similarity of 0.852, it does not consistently outperform the baseline, providing inconclusive evidence of cross-object generalisation at the scale of the available data. Taken together, these findings support a qualified rather than unqualified answer to the central hypothesis of this thesis. Tactile and visual embeddings appear sufficiently compatible, in aggregate, to justify learning a shared representation, yet generalisation to unseen objects remains unproven, while an independent mechanical confound limits policy performance regardless of the alignment strategy. Overall, the principal contribution of this thesis is a validated experimental infrastructure and a rigorous feasibility study, rather than a demonstrated end-to-end capability.
2025
Tactile-to-Visual Representation Learning for Vision-Language-Action Models
Vision-Language-Action (VLA) models have demonstrated strong generalisation capabilities in robotic manipulation, yet their action generation relies primarily on visual input, which is often degraded or entirely unavailable during contact-rich interactions such as grasping and in-hand manipulation. This thesis investigates whether a frozen VLA can maintain task performance under degraded or absent vision by learning a lightweight projection module, f_theta, that maps tactile embeddings into the policy's frozen visual latent space without modifying the pretrained policy weights. Rather than committing immediately to full-scale training, the work follows a feasibility-first methodology designed to establish whether such an approach is well-founded. First, a representational similarity analysis of tactile and visual embedding spaces is conducted on a small exploratory dataset, leading to the evidence-based selection of a generic ViT-Small tactile backbone over the official TVL encoder. Second, a fully reproducible benchmark is developed by integrating a custom DIGIT-equipped Panda gripper into the LIBERO/robosuite/MuJoCo simulation stack, extending TACTO's tactile rendering pipeline to a new simulator. This integration also reveals an independent mechanical confound: the gripper modification alone reduces the frozen policy's success rate on libero_object from 98.2% to 56.8%, before any tactile alignment is introduced. Third, the projection module is trained and evaluated under a leave-one-object-out protocol against a naive baseline. Although it achieves a mean cross-object cosine similarity of 0.852, it does not consistently outperform the baseline, providing inconclusive evidence of cross-object generalisation at the scale of the available data. Taken together, these findings support a qualified rather than unqualified answer to the central hypothesis of this thesis. Tactile and visual embeddings appear sufficiently compatible, in aggregate, to justify learning a shared representation, yet generalisation to unseen objects remains unproven, while an independent mechanical confound limits policy performance regardless of the alignment strategy. Overall, the principal contribution of this thesis is a validated experimental infrastructure and a rigorous feasibility study, rather than a demonstrated end-to-end capability.
Tactile
Learning
Robotics
File in questo prodotto:
File Dimensione Formato  
presentation.pdf

accesso aperto

Dimensione 5.33 MB
Formato Adobe PDF
5.33 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/110886