Robust 3D object recognition in real-world environments requires models that generalise beyond clean, canonical shapes to objects that are deformed, damaged, or geometrically ambiguous, cases that are systematically underrepresented in existing 3D datasets. This thesis addresses that gap through a training-free framework for generating synthetic 3D meshes that combine geometric properties from two distinct source objects, providing a principled data augmentation strategy for hard, out-of-distribution cases without requiring any retraining of the generative model. The investigation proceeds in two stages. Single-image transformations are first applied in both pixel and DINOv2 feature space, confirming that both are steerable for simple manipulations. When scaled to combining two images, however, pixel-level blending fails entirely while feature-space blending succeeds, motivating a principled framework in which a blending operator $\Phi$ acts on the encoded representations of two source images before injection into the frozen decoder, with no retraining required. Four strategies are developed and evaluated, linear interpolation, patch substitution, similarity-based adaptive blending, and CLS token manipulation, on diverse objects and all pairwise combinations, assessed via Chamfer Distance, Hausdorff Distance, mesh volume, surface area, and their relative differences. In addition, we used LLM(GPT-5) to carefully develop a more nuanced human-aligned ranking system. Results demonstrate that feature-space blending consistently produces geometrically coherent hybrid meshes, opening a path toward more robust 3D recognition in ambiguous and real-world settings.
Robust 3D object recognition in real-world environments requires models that generalise beyond clean, canonical shapes to objects that are deformed, damaged, or geometrically ambiguous, cases that are systematically underrepresented in existing 3D datasets. This thesis addresses that gap through a training-free framework for generating synthetic 3D meshes that combine geometric properties from two distinct source objects, providing a principled data augmentation strategy for hard, out-of-distribution cases without requiring any retraining of the generative model. The investigation proceeds in two stages. Single-image transformations are first applied in both pixel and DINOv2 feature space, confirming that both are steerable for simple manipulations. When scaled to combining two images, however, pixel-level blending fails entirely while feature-space blending succeeds, motivating a principled framework in which a blending operator $\Phi$ acts on the encoded representations of two source images before injection into the frozen decoder, with no retraining required. Four strategies are developed and evaluated, linear interpolation, patch substitution, similarity-based adaptive blending, and CLS token manipulation, on diverse objects and all pairwise combinations, assessed via Chamfer Distance, Hausdorff Distance, mesh volume, surface area, and their relative differences. In addition, we used LLM(GPT-5) to carefully develop a more nuanced human-aligned ranking system. Results demonstrate that feature-space blending consistently produces geometrically coherent hybrid meshes, opening a path toward more robust 3D recognition in ambiguous and real-world settings.
Generate Novel Objects: Combining Visual Concepts in the Latent Space of a Pre-Trained Image-to-3D Reconstruction Model
KALANTARY, HANNANEH
2025/2026
Abstract
Robust 3D object recognition in real-world environments requires models that generalise beyond clean, canonical shapes to objects that are deformed, damaged, or geometrically ambiguous, cases that are systematically underrepresented in existing 3D datasets. This thesis addresses that gap through a training-free framework for generating synthetic 3D meshes that combine geometric properties from two distinct source objects, providing a principled data augmentation strategy for hard, out-of-distribution cases without requiring any retraining of the generative model. The investigation proceeds in two stages. Single-image transformations are first applied in both pixel and DINOv2 feature space, confirming that both are steerable for simple manipulations. When scaled to combining two images, however, pixel-level blending fails entirely while feature-space blending succeeds, motivating a principled framework in which a blending operator $\Phi$ acts on the encoded representations of two source images before injection into the frozen decoder, with no retraining required. Four strategies are developed and evaluated, linear interpolation, patch substitution, similarity-based adaptive blending, and CLS token manipulation, on diverse objects and all pairwise combinations, assessed via Chamfer Distance, Hausdorff Distance, mesh volume, surface area, and their relative differences. In addition, we used LLM(GPT-5) to carefully develop a more nuanced human-aligned ranking system. Results demonstrate that feature-space blending consistently produces geometrically coherent hybrid meshes, opening a path toward more robust 3D recognition in ambiguous and real-world settings.| File | Dimensione | Formato | |
|---|---|---|---|
|
Final Thesis_KALANTARY.pdf
accesso aperto
Dimensione
40.71 MB
Formato
Adobe PDF
|
40.71 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110961