Product classification in retail remains a significant challenge due to dynamic catalogs, high inter-class similarity, and the domain gap between studio catalog images and in-store user shots. Traditional closed-set models are increasingly inefficient in these environments, as they require continuous fine-tuning to accommodate new product entries or packaging restyles. This thesis proposes a two-stage hybrid architecture operating in a zero-shot regime. The first stage employs CLIP-based dense retrieval to extract the top-K most likely candidates. The second stage leverages a Multimodal Large Language Model (MLLM) for precise reranking. To mitigate the high latency and computational costs typically associated with MLLMs, a token-efficient approach based on structured textual descriptions is introduced. Catalog images are processed offline to generate JSON dictionaries capturing key visual attributes. During inference, the MLLM performs semantic matching by comparing the query image exclusively against these structured strings. This strategy eliminates the need to process multiple high-resolution reference images in real-time, optimizing throughput without compromising classification accuracy. Empirical evaluations demonstrate that this approach significantly elevates the baseline performance, achieving an overall accuracy of 85.7% with top-performing models such as Gemini 3.5 Flash. The analysis confirms that structured JSON descriptions outperform freeform natural language, effectively reducing token consumption. Ultimately, this work validates zero-shot multimodal reasoning as a robust solution for retail applications.

Product classification in retail remains a significant challenge due to dynamic catalogs, high inter-class similarity, and the domain gap between studio catalog images and in-store user shots. Traditional closed-set models are increasingly inefficient in these environments, as they require continuous fine-tuning to accommodate new product entries or packaging restyles. This thesis proposes a two-stage hybrid architecture operating in a zero-shot regime. The first stage employs CLIP-based dense retrieval to extract the top-K most likely candidates. The second stage leverages a Multimodal Large Language Model (MLLM) for precise reranking. To mitigate the high latency and computational costs typically associated with MLLMs, a token-efficient approach based on structured textual descriptions is introduced. Catalog images are processed offline to generate JSON dictionaries capturing key visual attributes. During inference, the MLLM performs semantic matching by comparing the query image exclusively against these structured strings. This strategy eliminates the need to process multiple high-resolution reference images in real-time, optimizing throughput without compromising classification accuracy. Empirical evaluations demonstrate that this approach significantly elevates the baseline performance, achieving an overall accuracy of 85.7% with top-performing models such as Gemini 3.5 Flash. The analysis confirms that structured JSON descriptions outperform freeform natural language, effectively reducing token consumption. Ultimately, this work validates zero-shot multimodal reasoning as a robust solution for retail applications.

Zero-Shot Retail Product Classification: A Compact MLLM Reranking Approach

MARTINEZ, ZOREN
2025/2026

Abstract

Product classification in retail remains a significant challenge due to dynamic catalogs, high inter-class similarity, and the domain gap between studio catalog images and in-store user shots. Traditional closed-set models are increasingly inefficient in these environments, as they require continuous fine-tuning to accommodate new product entries or packaging restyles. This thesis proposes a two-stage hybrid architecture operating in a zero-shot regime. The first stage employs CLIP-based dense retrieval to extract the top-K most likely candidates. The second stage leverages a Multimodal Large Language Model (MLLM) for precise reranking. To mitigate the high latency and computational costs typically associated with MLLMs, a token-efficient approach based on structured textual descriptions is introduced. Catalog images are processed offline to generate JSON dictionaries capturing key visual attributes. During inference, the MLLM performs semantic matching by comparing the query image exclusively against these structured strings. This strategy eliminates the need to process multiple high-resolution reference images in real-time, optimizing throughput without compromising classification accuracy. Empirical evaluations demonstrate that this approach significantly elevates the baseline performance, achieving an overall accuracy of 85.7% with top-performing models such as Gemini 3.5 Flash. The analysis confirms that structured JSON descriptions outperform freeform natural language, effectively reducing token consumption. Ultimately, this work validates zero-shot multimodal reasoning as a robust solution for retail applications.
2025
Zero-Shot Retail Product Classification: A Compact MLLM Reranking Approach
Product classification in retail remains a significant challenge due to dynamic catalogs, high inter-class similarity, and the domain gap between studio catalog images and in-store user shots. Traditional closed-set models are increasingly inefficient in these environments, as they require continuous fine-tuning to accommodate new product entries or packaging restyles. This thesis proposes a two-stage hybrid architecture operating in a zero-shot regime. The first stage employs CLIP-based dense retrieval to extract the top-K most likely candidates. The second stage leverages a Multimodal Large Language Model (MLLM) for precise reranking. To mitigate the high latency and computational costs typically associated with MLLMs, a token-efficient approach based on structured textual descriptions is introduced. Catalog images are processed offline to generate JSON dictionaries capturing key visual attributes. During inference, the MLLM performs semantic matching by comparing the query image exclusively against these structured strings. This strategy eliminates the need to process multiple high-resolution reference images in real-time, optimizing throughput without compromising classification accuracy. Empirical evaluations demonstrate that this approach significantly elevates the baseline performance, achieving an overall accuracy of 85.7% with top-performing models such as Gemini 3.5 Flash. The analysis confirms that structured JSON descriptions outperform freeform natural language, effectively reducing token consumption. Ultimately, this work validates zero-shot multimodal reasoning as a robust solution for retail applications.
Retail products
LLM
Vision Transformer
Computer Vision
File in questo prodotto:
File Dimensione Formato  
Martinez_Zoren.pdf

Accesso riservato

Dimensione 14.82 MB
Formato Adobe PDF
14.82 MB Adobe PDF

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/111351