Composed image retrieval searches an image database with a query that pairs a reference im- age with a short modification text specifying how the desired result should differ from it. The task is challenging: a system must interpret the reference image, apply the requested modifi- cation, and pick out the matching items among many near-identical distractors. The recent i-CIR benchmark intensifies this difficulty by defining relevance at the level of a specific object instance rather than a general category, and by surrounding every query with hard negatives that penalise any method relying on a single modality. Its accompanying baseline, Basic, com- bines frozen vision–language features with careful, training-free post-processing and proves surprisingly competitive. Its weakness, however, is that it reads the modification text on its own, without the reference image: an instruction such as “make it red” tells a text encoder almost nothing, because the encoder cannot see what should become red. The idea of this thesis is to let language carry the composition. A frozen multimodal large language model reads the reference image and the modification instruction together and writes a short description of the imagined target, performing the requested edit in words rather than in pixels. This self-contained target caption is embedded by the frozen text encoder and fused with the image and text similarities through a clamped triplet product that behaves as a soft logical and: a candidate must agree with all three signals to rank highly. The whole pipeline is zero-shot – every encoder, the caption generator, and all fusion and post-processing steps are off- the-shelf and frozen. We choose this training-free setting deliberately: supervised CIR methods depend on large collections of (reference image, modification text, target image) triplets that are hard to obtain at scale and, however they are collected, cover far less visual and textual variety than the data behind modern vision–language pre-training; the reasons are developed in the introduction. The method is developed on an instance subset of i-CIR, and the successful configurations are re-run at full scale across several vision–language backbones. The target caption is consis- tently the single most effective component: it improves every backbone, and on its own, added to a plain image–text baseline, it is already competitive with the entire Basic post-processing stack it is designed to replace. On the standard CLIP ViT-L/14 backbone of the original bench- mark, the proposed method reaches 35.76 macro-mAP, surpassing both the 34.35 reported for the published Basic pipeline and our own reproduction of it (32.48 in an identical software environment); the best overall configuration attains 61.98 macro-mAP, the strongest result reported in this thesis.
Imagining the Target: Zero-Shot Instance-Level Composed Image Retrieval with MLLM-Generated Captions
HOSEINPOUR SIOUKI, ASMA
2025/2026
Abstract
Composed image retrieval searches an image database with a query that pairs a reference im- age with a short modification text specifying how the desired result should differ from it. The task is challenging: a system must interpret the reference image, apply the requested modifi- cation, and pick out the matching items among many near-identical distractors. The recent i-CIR benchmark intensifies this difficulty by defining relevance at the level of a specific object instance rather than a general category, and by surrounding every query with hard negatives that penalise any method relying on a single modality. Its accompanying baseline, Basic, com- bines frozen vision–language features with careful, training-free post-processing and proves surprisingly competitive. Its weakness, however, is that it reads the modification text on its own, without the reference image: an instruction such as “make it red” tells a text encoder almost nothing, because the encoder cannot see what should become red. The idea of this thesis is to let language carry the composition. A frozen multimodal large language model reads the reference image and the modification instruction together and writes a short description of the imagined target, performing the requested edit in words rather than in pixels. This self-contained target caption is embedded by the frozen text encoder and fused with the image and text similarities through a clamped triplet product that behaves as a soft logical and: a candidate must agree with all three signals to rank highly. The whole pipeline is zero-shot – every encoder, the caption generator, and all fusion and post-processing steps are off- the-shelf and frozen. We choose this training-free setting deliberately: supervised CIR methods depend on large collections of (reference image, modification text, target image) triplets that are hard to obtain at scale and, however they are collected, cover far less visual and textual variety than the data behind modern vision–language pre-training; the reasons are developed in the introduction. The method is developed on an instance subset of i-CIR, and the successful configurations are re-run at full scale across several vision–language backbones. The target caption is consis- tently the single most effective component: it improves every backbone, and on its own, added to a plain image–text baseline, it is already competitive with the entire Basic post-processing stack it is designed to replace. On the standard CLIP ViT-L/14 backbone of the original bench- mark, the proposed method reaches 35.76 macro-mAP, surpassing both the 34.35 reported for the published Basic pipeline and our own reproduction of it (32.48 in an identical software environment); the best overall configuration attains 61.98 macro-mAP, the strongest result reported in this thesis.| File | Dimensione | Formato | |
|---|---|---|---|
|
Asma_HoseinpourSiouki.pdf
accesso aperto
Dimensione
32.13 MB
Formato
Adobe PDF
|
32.13 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110926