Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
Yiyang ChenZhedong ZhengWei JiLeigang QuTat-Seng Chua
Proposes a multi-grained uncertainty regularization framework for composed image retrieval that unifies coarse-to-fine text feedback by modeling feature fluctuations, preventing premature candidate exclusion and significantly increasing recall across standard benchmarks.
Modern e-commerce and search platforms increasingly rely on multimodal image retrieval, where users search for products using an initial reference image modified by descriptive text feedback (such as requesting a shoe in a different color). In practice, users begin with broad, coarse-grained queries before narrowing their search to specific items. However, existing retrieval models are trained using strict one-to-one matching objectives borrowed from biometrics. This creates a critical misalignment: standard training penalizes models unless they match a single exact image, which inadvertently pushes away valid alternative candidates during the early search stages and significantly reduces overall retrieval recall.
The article evaluates a unified learning framework that simultaneously models both coarse-grained (one-to-many) and fine-grained (one-to-one) retrieval. Its primary objective is to demonstrate that introducing controlled feature fluctuations during training prevents model overfitting on ambiguous queries and improves candidate retrieval rates across benchmark datasets.
To address this challenge, the authors introduced an uncertainty modeling and regularization framework. Rather than modifying the underlying neural network architecture, they introduced a data-driven feature augmentation module that injects Gaussian noise based on the target image feature distribution to simulate natural candidate variations. They paired this with an adaptive uncertainty loss function that reduces penalty severity when feature fluctuation is high. A dynamic weighting schedule gradually shifts the model from coarse-grained matching early in training toward strict fine-grained matching as training progresses. The approach was evaluated across three standard public benchmarks: FashionIQ (containing over 46,000 training images), Fashion200k (with 172,000 training images), and Shoes (with 10,000 training pairs).
The evaluation produced several key findings. First, incorporating uncertainty regularization consistently improved retrieval recall across all benchmarks: on Recall@50, performance increased by 4.16 percentage points on FashionIQ (reaching 61.39%), by 3.38 percentage points on Shoes (reaching 79.84%), and by 2.40 percentage points on Fashion200k (reaching 70.20%). Second, a single standard ResNet-50 visual backbone trained with this method outperformed larger multi-model ensemble baselines, such as an ensemble using four ResNet-50 backbones that achieved 59.03% Recall@50 on FashionIQ. Third, ablation tests confirmed that targeted Gaussian feature noise outperforms standard regularization techniques like dropout, and that dynamic loss scheduling is necessary to avoid underfitting fine-grained details. Finally, the framework proved complementary to existing retrieval architectures, providing performance gains of roughly 2.5 to 5.0 percentage points when integrated with state-of-the-art multimodal vision-language backbones.
These findings indicate that search and recommendation platforms can significantly improve candidate discovery without expanding model parameter size or incurring higher inference costs. By avoiding the rigid one-to-one mapping trap, systems provide users with more stylistically relevant options in top search rankings, directly improving search user experience in commercial retail applications. Because the methodology modifies only the training loss and data augmentation process, it offers a low-risk, plug-and-play upgrade for existing retrieval pipelines.
Engineering and product teams developing text-guided visual search tools should consider integrating uncertainty regularization into their current training pipelines. When adopting this method, teams should calibrate the initial loss balance weight to ensure proper annealing and apply noise injection exclusively to target image representations rather than source query representations to prevent inverted matching dynamics.
While confidence in the benchmark results is high due to consistent gains across multiple datasets and backbones, the article's empirical scope is confined primarily to fashion and apparel datasets. Organizations deploying this method to broader, open-domain visual search scenarios should perform pilot validation to ensure the noise modeling generalizes across more diverse object categories.
- Paper: MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model, Yatai Ji et al. (2023). Provides foundational principles for probabilistic distribution encoding and uncertainty modeling in multimodal vision-language representations that motivate target feature fluctuation strategies.
- Paper: DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations, Ziwei Liu et al. (2016). Introduces the core concepts, benchmark characteristics, and retrieval architectures for fashion-domain visual search that underpin composed image retrieval datasets like FashionIQ.
- Paper: Noisy Correspondence Learning with Meta Similarity Correction, Haochen Han et al. (2023). Establishes techniques for handling noisy and ambiguous multimodal alignments in cross-modal retrieval, directly informing regularized training objectives for non-rigid matching.
- Paper: Negative-Aware Attention Framework for Image-Text Matching, Kun Zhang et al. (2022). Examines dissimilarity penalties and dynamic decision boundaries in vision-language matching, establishing essential background for balancing fine-grained versus coarse-grained alignment penalties.
- Paper: What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?, Alex Kendall et al. (2017). Formalizes aleatoric and epistemic uncertainty modeling in computer vision deep networks, providing the theoretical grounding for learning with data-dependent feature fluctuations.
- Paper: Deep Metric Learning via Lifted Structured Feature Embedding, Hyun Oh Song et al. (2015). Presents deep metric learning strategies and structured feature embeddings for visual product search that define standard pairwise and triplet training paradigms.
- Paper: A Simple Framework for Contrastive Learning of Visual Representations, Ting Chen et al. (2020). Demonstrates how controlled data and feature augmentations drive effective contrastive visual representation learning.
- Paper: SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger, Yuting Gao et al. (2024). Extends the principle of relaxing rigid one-to-one pairings to large-scale foundation vision-language models by constructing soft alignment targets via intra-modal similarities.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). Applies fine-grained multimodal late-interaction retrieval frameworks across diverse knowledge-intensive benchmarks, expanding multi-modal query-driven search beyond fashion domains.
- Paper: Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID, Wentao Tan et al. (2024). Leverages multimodal large language models and noise-aware token masking to overcome noisy textual-visual alignments in complex cross-modal retrieval tasks.
