Vision-Language-Action (VLA) models have demonstrated strong generalisation capabilities in robotic manipulation, yet their action generation relies primarily on visual input, which is often degraded or entirely unavailable during contact-rich interactions such as grasping and in-hand manipulation. This thesis investigates whether a frozen VLA can maintain task performance under degraded or absent vision by learning a lightweight projection module, f_theta, that maps tactile embeddings into the policy's frozen visual latent space without modifying the pretrained policy weights. Rather than committing immediately to full-scale training, the work follows a feasibility-first methodology designed to establish whether such an approach is well-founded. First, a representational similarity analysis of tactile and visual embedding spaces is conducted on a small exploratory dataset, leading to the evidence-based selection of a generic ViT-Small tactile backbone over the official TVL encoder. Second, a fully reproducible benchmark is developed by integrating a custom DIGIT-equipped Panda gripper into the LIBERO/robosuite/MuJoCo simulation stack, extending TACTO's tactile rendering pipeline to a new simulator. This integration also reveals an independent mechanical confound: the gripper modification alone reduces the frozen policy's success rate on libero_object from 98.2% to 56.8%, before any tactile alignment is introduced. Third, the projection module is trained and evaluated under a leave-one-object-out protocol against a naive baseline. Although it achieves a mean cross-object cosine similarity of 0.852, it does not consistently outperform the baseline, providing inconclusive evidence of cross-object generalisation at the scale of the available data. Taken together, these findings support a qualified rather than unqualified answer to the central hypothesis of this thesis. Tactile and visual embeddings appear sufficiently compatible, in aggregate, to justify learning a shared representation, yet generalisation to unseen objects remains unproven, while an independent mechanical confound limits policy performance regardless of the alignment strategy. Overall, the principal contribution of this thesis is a validated experimental infrastructure and a rigorous feasibility study, rather than a demonstrated end-to-end capability.
Vision-Language-Action (VLA) models have demonstrated strong generalisation capabilities in robotic manipulation, yet their action generation relies primarily on visual input, which is often degraded or entirely unavailable during contact-rich interactions such as grasping and in-hand manipulation. This thesis investigates whether a frozen VLA can maintain task performance under degraded or absent vision by learning a lightweight projection module, f_theta, that maps tactile embeddings into the policy's frozen visual latent space without modifying the pretrained policy weights. Rather than committing immediately to full-scale training, the work follows a feasibility-first methodology designed to establish whether such an approach is well-founded. First, a representational similarity analysis of tactile and visual embedding spaces is conducted on a small exploratory dataset, leading to the evidence-based selection of a generic ViT-Small tactile backbone over the official TVL encoder. Second, a fully reproducible benchmark is developed by integrating a custom DIGIT-equipped Panda gripper into the LIBERO/robosuite/MuJoCo simulation stack, extending TACTO's tactile rendering pipeline to a new simulator. This integration also reveals an independent mechanical confound: the gripper modification alone reduces the frozen policy's success rate on libero_object from 98.2% to 56.8%, before any tactile alignment is introduced. Third, the projection module is trained and evaluated under a leave-one-object-out protocol against a naive baseline. Although it achieves a mean cross-object cosine similarity of 0.852, it does not consistently outperform the baseline, providing inconclusive evidence of cross-object generalisation at the scale of the available data. Taken together, these findings support a qualified rather than unqualified answer to the central hypothesis of this thesis. Tactile and visual embeddings appear sufficiently compatible, in aggregate, to justify learning a shared representation, yet generalisation to unseen objects remains unproven, while an independent mechanical confound limits policy performance regardless of the alignment strategy. Overall, the principal contribution of this thesis is a validated experimental infrastructure and a rigorous feasibility study, rather than a demonstrated end-to-end capability.
Tactile-to-Visual Representation Learning for Vision-Language-Action Models
GIRARDELLO, NICOLO'
2025/2026
Abstract
Vision-Language-Action (VLA) models have demonstrated strong generalisation capabilities in robotic manipulation, yet their action generation relies primarily on visual input, which is often degraded or entirely unavailable during contact-rich interactions such as grasping and in-hand manipulation. This thesis investigates whether a frozen VLA can maintain task performance under degraded or absent vision by learning a lightweight projection module, f_theta, that maps tactile embeddings into the policy's frozen visual latent space without modifying the pretrained policy weights. Rather than committing immediately to full-scale training, the work follows a feasibility-first methodology designed to establish whether such an approach is well-founded. First, a representational similarity analysis of tactile and visual embedding spaces is conducted on a small exploratory dataset, leading to the evidence-based selection of a generic ViT-Small tactile backbone over the official TVL encoder. Second, a fully reproducible benchmark is developed by integrating a custom DIGIT-equipped Panda gripper into the LIBERO/robosuite/MuJoCo simulation stack, extending TACTO's tactile rendering pipeline to a new simulator. This integration also reveals an independent mechanical confound: the gripper modification alone reduces the frozen policy's success rate on libero_object from 98.2% to 56.8%, before any tactile alignment is introduced. Third, the projection module is trained and evaluated under a leave-one-object-out protocol against a naive baseline. Although it achieves a mean cross-object cosine similarity of 0.852, it does not consistently outperform the baseline, providing inconclusive evidence of cross-object generalisation at the scale of the available data. Taken together, these findings support a qualified rather than unqualified answer to the central hypothesis of this thesis. Tactile and visual embeddings appear sufficiently compatible, in aggregate, to justify learning a shared representation, yet generalisation to unseen objects remains unproven, while an independent mechanical confound limits policy performance regardless of the alignment strategy. Overall, the principal contribution of this thesis is a validated experimental infrastructure and a rigorous feasibility study, rather than a demonstrated end-to-end capability.| File | Dimensione | Formato | |
|---|---|---|---|
|
presentation.pdf
accesso aperto
Dimensione
5.33 MB
Formato
Adobe PDF
|
5.33 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110886