Product classification in retail remains a significant challenge due to dynamic catalogs, high inter-class similarity, and the domain gap between studio catalog images and in-store user shots. Traditional closed-set models are increasingly inefficient in these environments, as they require continuous fine-tuning to accommodate new product entries or packaging restyles. This thesis proposes a two-stage hybrid architecture operating in a zero-shot regime. The first stage employs CLIP-based dense retrieval to extract the top-K most likely candidates. The second stage leverages a Multimodal Large Language Model (MLLM) for precise reranking. To mitigate the high latency and computational costs typically associated with MLLMs, a token-efficient approach based on structured textual descriptions is introduced. Catalog images are processed offline to generate JSON dictionaries capturing key visual attributes. During inference, the MLLM performs semantic matching by comparing the query image exclusively against these structured strings. This strategy eliminates the need to process multiple high-resolution reference images in real-time, optimizing throughput without compromising classification accuracy. Empirical evaluations demonstrate that this approach significantly elevates the baseline performance, achieving an overall accuracy of 85.7% with top-performing models such as Gemini 3.5 Flash. The analysis confirms that structured JSON descriptions outperform freeform natural language, effectively reducing token consumption. Ultimately, this work validates zero-shot multimodal reasoning as a robust solution for retail applications.
Product classification in retail remains a significant challenge due to dynamic catalogs, high inter-class similarity, and the domain gap between studio catalog images and in-store user shots. Traditional closed-set models are increasingly inefficient in these environments, as they require continuous fine-tuning to accommodate new product entries or packaging restyles. This thesis proposes a two-stage hybrid architecture operating in a zero-shot regime. The first stage employs CLIP-based dense retrieval to extract the top-K most likely candidates. The second stage leverages a Multimodal Large Language Model (MLLM) for precise reranking. To mitigate the high latency and computational costs typically associated with MLLMs, a token-efficient approach based on structured textual descriptions is introduced. Catalog images are processed offline to generate JSON dictionaries capturing key visual attributes. During inference, the MLLM performs semantic matching by comparing the query image exclusively against these structured strings. This strategy eliminates the need to process multiple high-resolution reference images in real-time, optimizing throughput without compromising classification accuracy. Empirical evaluations demonstrate that this approach significantly elevates the baseline performance, achieving an overall accuracy of 85.7% with top-performing models such as Gemini 3.5 Flash. The analysis confirms that structured JSON descriptions outperform freeform natural language, effectively reducing token consumption. Ultimately, this work validates zero-shot multimodal reasoning as a robust solution for retail applications.
Zero-Shot Retail Product Classification: A Compact MLLM Reranking Approach
MARTINEZ, ZOREN
2025/2026
Abstract
Product classification in retail remains a significant challenge due to dynamic catalogs, high inter-class similarity, and the domain gap between studio catalog images and in-store user shots. Traditional closed-set models are increasingly inefficient in these environments, as they require continuous fine-tuning to accommodate new product entries or packaging restyles. This thesis proposes a two-stage hybrid architecture operating in a zero-shot regime. The first stage employs CLIP-based dense retrieval to extract the top-K most likely candidates. The second stage leverages a Multimodal Large Language Model (MLLM) for precise reranking. To mitigate the high latency and computational costs typically associated with MLLMs, a token-efficient approach based on structured textual descriptions is introduced. Catalog images are processed offline to generate JSON dictionaries capturing key visual attributes. During inference, the MLLM performs semantic matching by comparing the query image exclusively against these structured strings. This strategy eliminates the need to process multiple high-resolution reference images in real-time, optimizing throughput without compromising classification accuracy. Empirical evaluations demonstrate that this approach significantly elevates the baseline performance, achieving an overall accuracy of 85.7% with top-performing models such as Gemini 3.5 Flash. The analysis confirms that structured JSON descriptions outperform freeform natural language, effectively reducing token consumption. Ultimately, this work validates zero-shot multimodal reasoning as a robust solution for retail applications.| File | Dimensione | Formato | |
|---|---|---|---|
|
Martinez_Zoren.pdf
Accesso riservato
Dimensione
14.82 MB
Formato
Adobe PDF
|
14.82 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/111351