Image retrieval systems increasingly need to support precise user intent rather than only broad semantic similarity. In instance-based composed image retrieval, the query combines a refer- ence image and a textual modification. A correct target must depict the same literal instance, such as the same product, character, landmark, or object, and must also satisfy the requested vi- sual change. This makes the task stricter than ordinary composed retrieval: a candidate is wrong if it shows the correct modification on a different instance, and it is also wrong if it preserves the instance but ignores the modification. This thesis studies the task as a conjunction of two verification signals: identity preservation and modification satisfaction. The central idea is to keep these two forms of evidence visible during scoring and combine them multiplicatively, so that a candidate is rewarded only when both are supported. I first audit several frozen embedding backbones and training-free retrieval components under the intended query-local hard-negative protocol. The strongest first-stage system uses SigLIP2 SO400M Patch16 384 with leave-one-instance-out image centering and text contextualization, reaching 59.64 mean average precision at 100 on the full benchmark. The thesis then introduces two second-stage reranking routes that preserve the first-stage im- age and text scores while adding more instance-focused evidence. Route A uses ILIAS, an image instance-level retrieval dataset, as an auxiliary source of supervision. Lightweight linear probes trained on frozen embeddings show that strong instance-level structure is already present in the representation; when used as a dual-encoder reranking signal together with a contextual- ized text branch, this route raises the ICIR score to 66.05 mean average precision at 100. Route B uses RZEnEmbed-7B as an instruction-aware multimodal embedder. Candidate images are embedded with an instance descriptor generated offline by open-source multimodal language models, encouraging the reranker to focus on the literal target instance rather than on generic scene content. The best final system is the descriptor-guided Route B product, which multiplies the first- stage image score, the first-stage text score, and two RZEn reranking scores. It reaches 66.42 mean average precision at 100 and 65.65 macro mean average precision at 100 while preserving the same candidate-set recall as the first stage. The results show that modern frozen embed- dings already contain useful instance information, but that high-performing instance-based composed image retrieval requires careful separation and recombination of identity and modi- fication evidence.

Image retrieval systems increasingly need to support precise user intent rather than only broad semantic similarity. In instance-based composed image retrieval, the query combines a refer- ence image and a textual modification. A correct target must depict the same literal instance, such as the same product, character, landmark, or object, and must also satisfy the requested vi- sual change. This makes the task stricter than ordinary composed retrieval: a candidate is wrong if it shows the correct modification on a different instance, and it is also wrong if it preserves the instance but ignores the modification. This thesis studies the task as a conjunction of two verification signals: identity preservation and modification satisfaction. The central idea is to keep these two forms of evidence visible during scoring and combine them multiplicatively, so that a candidate is rewarded only when both are supported. I first audit several frozen embedding backbones and training-free retrieval components under the intended query-local hard-negative protocol. The strongest first-stage system uses SigLIP2 SO400M Patch16 384 with leave-one-instance-out image centering and text contextualization, reaching 59.64 mean average precision at 100 on the full benchmark. The thesis then introduces two second-stage reranking routes that preserve the first-stage im- age and text scores while adding more instance-focused evidence. Route A uses ILIAS, an image instance-level retrieval dataset, as an auxiliary source of supervision. Lightweight linear probes trained on frozen embeddings show that strong instance-level structure is already present in the representation; when used as a dual-encoder reranking signal together with a contextual- ized text branch, this route raises the ICIR score to 66.05 mean average precision at 100. Route B uses RZEnEmbed-7B as an instruction-aware multimodal embedder. Candidate images are embedded with an instance descriptor generated offline by open-source multimodal language models, encouraging the reranker to focus on the literal target instance rather than on generic scene content. The best final system is the descriptor-guided Route B product, which multiplies the first- stage image score, the first-stage text score, and two RZEn reranking scores. It reaches 66.42 mean average precision at 100 and 65.65 macro mean average precision at 100 while preserving the same candidate-set recall as the first stage. The results show that modern frozen embed- dings already contain useful instance information, but that high-performing instance-based composed image retrieval requires careful separation and recombination of identity and modification evidence.

Instance-Based Composed Image Retrieval with Vision-Language Models

ALMASI KOLOR, MOHAMMADAMIN
2025/2026

Abstract

Image retrieval systems increasingly need to support precise user intent rather than only broad semantic similarity. In instance-based composed image retrieval, the query combines a refer- ence image and a textual modification. A correct target must depict the same literal instance, such as the same product, character, landmark, or object, and must also satisfy the requested vi- sual change. This makes the task stricter than ordinary composed retrieval: a candidate is wrong if it shows the correct modification on a different instance, and it is also wrong if it preserves the instance but ignores the modification. This thesis studies the task as a conjunction of two verification signals: identity preservation and modification satisfaction. The central idea is to keep these two forms of evidence visible during scoring and combine them multiplicatively, so that a candidate is rewarded only when both are supported. I first audit several frozen embedding backbones and training-free retrieval components under the intended query-local hard-negative protocol. The strongest first-stage system uses SigLIP2 SO400M Patch16 384 with leave-one-instance-out image centering and text contextualization, reaching 59.64 mean average precision at 100 on the full benchmark. The thesis then introduces two second-stage reranking routes that preserve the first-stage im- age and text scores while adding more instance-focused evidence. Route A uses ILIAS, an image instance-level retrieval dataset, as an auxiliary source of supervision. Lightweight linear probes trained on frozen embeddings show that strong instance-level structure is already present in the representation; when used as a dual-encoder reranking signal together with a contextual- ized text branch, this route raises the ICIR score to 66.05 mean average precision at 100. Route B uses RZEnEmbed-7B as an instruction-aware multimodal embedder. Candidate images are embedded with an instance descriptor generated offline by open-source multimodal language models, encouraging the reranker to focus on the literal target instance rather than on generic scene content. The best final system is the descriptor-guided Route B product, which multiplies the first- stage image score, the first-stage text score, and two RZEn reranking scores. It reaches 66.42 mean average precision at 100 and 65.65 macro mean average precision at 100 while preserving the same candidate-set recall as the first stage. The results show that modern frozen embed- dings already contain useful instance information, but that high-performing instance-based composed image retrieval requires careful separation and recombination of identity and modi- fication evidence.
2025
Instance-Based Composed Image Retrieval with Vision-Language Models
Image retrieval systems increasingly need to support precise user intent rather than only broad semantic similarity. In instance-based composed image retrieval, the query combines a refer- ence image and a textual modification. A correct target must depict the same literal instance, such as the same product, character, landmark, or object, and must also satisfy the requested vi- sual change. This makes the task stricter than ordinary composed retrieval: a candidate is wrong if it shows the correct modification on a different instance, and it is also wrong if it preserves the instance but ignores the modification. This thesis studies the task as a conjunction of two verification signals: identity preservation and modification satisfaction. The central idea is to keep these two forms of evidence visible during scoring and combine them multiplicatively, so that a candidate is rewarded only when both are supported. I first audit several frozen embedding backbones and training-free retrieval components under the intended query-local hard-negative protocol. The strongest first-stage system uses SigLIP2 SO400M Patch16 384 with leave-one-instance-out image centering and text contextualization, reaching 59.64 mean average precision at 100 on the full benchmark. The thesis then introduces two second-stage reranking routes that preserve the first-stage im- age and text scores while adding more instance-focused evidence. Route A uses ILIAS, an image instance-level retrieval dataset, as an auxiliary source of supervision. Lightweight linear probes trained on frozen embeddings show that strong instance-level structure is already present in the representation; when used as a dual-encoder reranking signal together with a contextual- ized text branch, this route raises the ICIR score to 66.05 mean average precision at 100. Route B uses RZEnEmbed-7B as an instruction-aware multimodal embedder. Candidate images are embedded with an instance descriptor generated offline by open-source multimodal language models, encouraging the reranker to focus on the literal target instance rather than on generic scene content. The best final system is the descriptor-guided Route B product, which multiplies the first- stage image score, the first-stage text score, and two RZEn reranking scores. It reaches 66.42 mean average precision at 100 and 65.65 macro mean average precision at 100 while preserving the same candidate-set recall as the first stage. The results show that modern frozen embed- dings already contain useful instance information, but that high-performing instance-based composed image retrieval requires careful separation and recombination of identity and modification evidence.
Image retrieval
Instance Retrieval
Vision-Language Mode
File in questo prodotto:
File Dimensione Formato  
Mohammadamin_Almasi_Kolor.pdf

accesso aperto

Dimensione 12.88 MB
Formato Adobe PDF
12.88 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/110919