Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization

Yiyang ChenZhedong ZhengWei JiLeigang QuTat-Seng Chua

article2024ICLR76 citations

Proposes a multi-grained uncertainty regularization framework for composed image retrieval that unifies coarse-to-fine text feedback by modeling feature fluctuations, preventing premature candidate exclusion and significantly increasing recall across standard benchmarks.

Listen

Modern e-commerce and search platforms increasingly rely on multimodal image retrieval, where users search for products using an initial reference image modified by descriptive text feedback (such as requesting a shoe in a different color). In practice, users begin with broad, coarse-grained queries before narrowing their search to specific items. However, existing retrieval models are trained using strict one-to-one matching objectives borrowed from biometrics. This creates a critical misalignment: standard training penalizes models unless they match a single exact image, which inadvertently pushes away valid alternative candidates during the early search stages and significantly reduces overall retrieval recall.

The article evaluates a unified learning framework that simultaneously models both coarse-grained (one-to-many) and fine-grained (one-to-one) retrieval. Its primary objective is to demonstrate that introducing controlled feature fluctuations during training prevents model overfitting on ambiguous queries and improves candidate retrieval rates across benchmark datasets.

To address this challenge, the authors introduced an uncertainty modeling and regularization framework. Rather than modifying the underlying neural network architecture, they introduced a data-driven feature augmentation module that injects Gaussian noise based on the target image feature distribution to simulate natural candidate variations. They paired this with an adaptive uncertainty loss function that reduces penalty severity when feature fluctuation is high. A dynamic weighting schedule gradually shifts the model from coarse-grained matching early in training toward strict fine-grained matching as training progresses. The approach was evaluated across three standard public benchmarks: FashionIQ (containing over 46,000 training images), Fashion200k (with 172,000 training images), and Shoes (with 10,000 training pairs).

The evaluation produced several key findings. First, incorporating uncertainty regularization consistently improved retrieval recall across all benchmarks: on Recall@50, performance increased by 4.16 percentage points on FashionIQ (reaching 61.39%), by 3.38 percentage points on Shoes (reaching 79.84%), and by 2.40 percentage points on Fashion200k (reaching 70.20%). Second, a single standard ResNet-50 visual backbone trained with this method outperformed larger multi-model ensemble baselines, such as an ensemble using four ResNet-50 backbones that achieved 59.03% Recall@50 on FashionIQ. Third, ablation tests confirmed that targeted Gaussian feature noise outperforms standard regularization techniques like dropout, and that dynamic loss scheduling is necessary to avoid underfitting fine-grained details. Finally, the framework proved complementary to existing retrieval architectures, providing performance gains of roughly 2.5 to 5.0 percentage points when integrated with state-of-the-art multimodal vision-language backbones.

These findings indicate that search and recommendation platforms can significantly improve candidate discovery without expanding model parameter size or incurring higher inference costs. By avoiding the rigid one-to-one mapping trap, systems provide users with more stylistically relevant options in top search rankings, directly improving search user experience in commercial retail applications. Because the methodology modifies only the training loss and data augmentation process, it offers a low-risk, plug-and-play upgrade for existing retrieval pipelines.

Engineering and product teams developing text-guided visual search tools should consider integrating uncertainty regularization into their current training pipelines. When adopting this method, teams should calibrate the initial loss balance weight to ensure proper annealing and apply noise injection exclusively to target image representations rather than source query representations to prevent inverted matching dynamics.

While confidence in the benchmark results is high due to consistent gains across multiple datasets and backbones, the article's empirical scope is confined primarily to fashion and apparel datasets. Organizations deploying this method to broader, open-domain visual search scenarios should perform pilot validation to ensure the noise modeling generalizes across more diverse object categories.

arXiv: 2211.07394
Cover for Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization

Abstract

We investigate composed image retrieval with text feedback. Users gradually look for the target of interest by moving from coarse to fine-grained feedback. However, existing methods merely focus on the latter, i.e., fine-grained search, by harnessing positive and negative pairs during training. This pair-based paradigm only considers the one-to-one distance between a pair of specific points, which is not aligned with the one-to-many coarse-grained retrieval process and compromises the recall rate. In an attempt to fill this gap, we introduce a unified learning approach to simultaneously modeling the coarse- and fine-grained retrieval by considering the multi-grained uncertainty. The key idea underpinning the proposed method is to integrate fine- and coarse-grained retrieval as matching data points with small and large fluctuations, respectively. Specifically, our method contains two modules: uncertainty modeling and uncertainty regularization. (1) The uncertainty modeling simulates the multi-grained queries by introducing identically distributed fluctuations in the feature space. (2) Based on the uncertainty modeling, we further introduce uncertainty regularization to adapt the matching objective according to the fluctuation range. Compared with existing methods, the proposed strategy explicitly prevents the model from pushing away potential candidates in the early stage, and thus improves the recall rate. On the three public datasets, i.e., FashionIQ, Fashion200k, and Shoes, the proposed method has achieved +4.03%, +3.38%, and +2.40% Recall@50 accuracy over a strong baseline, respectively.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 2.1 Composed Image Retrieval with Text Feedback
  • 2.2 Uncertainty Learning
  • 3 Method
  • 3.1 Problem Definition.
  • 3.2 Uncertainty Modeling
  • 3.3 Uncertainty Regularization
  • 4 Experiment
  • 4.1 Implementation Details
  • 4.2 Datasets
  • 4.3 Comparison with Competitive Methods
  • 5 Further Analysis and Discussions
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Feature-Level Uncertainty Modeling via Target Feature Jittering

    model/method

    In composed image retrieval with text feedback, queries can be coarse-grained or ambiguous, meaning that a single query (Is,Ts)(I_s, T_s) may legitimately match multiple candidate target images. Standard fine-grained metric learning strictly enforces one-to-one matching, which repels valid candidates and degrades recall. To model one-to-many coarse-grained matching in the representation space without altering network architectures, an uncertainty augmenter introduces intra-class jittering directly into the feature embedding ft∈RDf_t \in \mathbb{R}^D of the target image ItI_t.

    The mean μt∈R\mu_t \in \mathbb{R} and standard deviation σt∈R\sigma_t \in \mathbb{R} are calculated across the feature dimensions of the original target embedding ftf_t: μt=1D∑d=1Dft,d,σt=1D∑d=1D(ft,d−μt)2\mu_t = \frac{1}{D}\sum_{d=1}^D f_{t,d}, \quad \sigma_t = \sqrt{\frac{1}{D}\sum_{d=1}^D (f_{t,d} - \mu_t)^2}

    The target feature is first standardized (whitened): fˉt=ft−μtσt\bar{f}_t = \frac{f_t - \mu_t}{\sigma_t}

    The jittered target representation f^t\hat{f}_t is then synthesized via element-wise re-scaling and shifting using Gaussian noise vectors: f^t=α⋅fˉt+β\hat{f}_t = \alpha \cdot \bar{f}_t + \beta where α∼N(1,w1σtI)\alpha \sim \mathcal{N}(\mathbf{1}, w_1 \sigma_t \mathbf{I}) and β∼N(μt1,w2σtI)\beta \sim \mathcal{N}(\mu_t \mathbf{1}, w_2 \sigma_t \mathbf{I}), with 1\mathbf{1} denoting the vector of ones and w1,w2w_1, w_2 scaling hyperparameters (w1=1,w2=1w_1 = 1, w_2 = 1 by default). Applying this perturbation to the target feature ftf_t rather than the source feature fsf_s preserves one-to-many retrieval semantics.

  2. Knowl 2 — Multi-Grained Uncertainty Regularization Loss

    equation

    To unify fine-grained (one-to-one) and coarse-grained (one-to-many) retrieval in a single training objective, an uncertainty regularization loss derived from aleatoric uncertainty is formulated.

    For a mini-batch of BB samples with composed query features fsif_s^i and target features ftif_t^i, the standard InfoNCE loss is defined as: Linfo(fs,ft)=1B∑i=1B−log⁡exp⁡(κ(fsi,fti))∑j=1Bexp⁡(κ(fsi,ftj))\mathcal{L}_{info}(f_s, f_t) = \frac{1}{B} \sum_{i=1}^B -\log \frac{\exp\left(\kappa(f_s^i, f_t^i)\right)}{\sum_{j=1}^B \exp\left(\kappa(f_s^i, f_t^j)\right)} where κ(u,v)=u⋅v∥u∥∥v∥\kappa(u, v) = \frac{u \cdot v}{\|u\| \|v\|} is the cosine similarity.

    Given the jittered target feature f^t\hat{f}_t and its standard deviation σt\sigma_t, the uncertainty regularization loss Lu\mathcal{L}_u is defined as: Lu(fs,f^t,σt)=Linfo(fs,f^t)2σt2+12log⁡σt2\mathcal{L}_u(f_s, \hat{f}_t, \sigma_t) = \frac{\mathcal{L}_{info}(f_s, \hat{f}_t)}{2\sigma_t^2} + \frac{1}{2} \log \sigma_t^2 When the fluctuation σt\sigma_t is large, the penalty on the InfoNCE loss term is automatically attenuated, avoiding overly strict penalties for ambiguous queries.

    The overall training loss linearly balances the uncertainty regularized coarse-matching loss and the fine-grained InfoNCE loss using dynamic weight γ∈[0,1]\gamma \in [0, 1]: Ltotal=γLu(fs,f^t,σt)+(1−γ)Linfo(fs,ft)\mathcal{L}_{total} = \gamma \mathcal{L}_u(f_s, \hat{f}_t, \sigma_t) + (1 - \gamma)\mathcal{L}_{info}(f_s, f_t) which can be equivalently rewritten in unified form by absorbing the standard InfoNCE loss into Lu\mathcal{L}_u up to a constant gradient-free term: Ltotal=γLu(fs,f^t,σt)+(1−γ)Lu(fs,ft,12)\mathcal{L}_{total} = \gamma \mathcal{L}_u\left(f_s, \hat{f}_t, \sigma_t\right) + (1 - \gamma)\mathcal{L}_u\left(f_s, f_t, \frac{1}{\sqrt{2}}\right)

  3. Knowl 3 — Dynamic Weight Annealing Schedule for Multi-Grained Loss Balancing

    model/method

    Training with only coarse-grained uncertainty loss leads to early convergence on easy matches while failing on fine-grained discrimination; training with only fine-grained InfoNCE causes overfitting to strict 1-to-1 annotations. To balance both objectives throughout training, the loss weight γ∈[0,1]\gamma \in [0, 1] is dynamically decayed via an exponential annealing schedule: γ=exp⁡(−γ0⋅eE)\gamma = \exp\left(-\gamma_0 \cdot \frac{e}{E}\right) where e∈{1,2,…,E}e \in \{1, 2, \dots, E\} represents the current training epoch, EE is the total number of epochs (typically E=50E = 50), and γ0\gamma_0 is the initial weight hyperparameter (set to γ0=1.0\gamma_0 = 1.0).

    At the start of training (e≪Ee \ll E), γ≈1\gamma \approx 1, compelling the network to prioritize coarse-grained uncertainty regularization and avoid pushing away potential matching candidates. As training progresses (e→Ee \to E), γ\gamma decreases smoothly toward exp⁡(−γ0)\exp(-\gamma_0), refocusing optimization on fine-grained metric alignment.

  4. Knowl 4 — Experimental Setup and Implementation Details for Composed Image Retrieval

    experimental setup

    The uncertainty regularization framework for composed image retrieval is evaluated using the following benchmark environment and implementation setup:

    • Encoders: ImageNet-pretrained ResNet-50 as the visual backbone; RoBERTa as the text encoder.
    • Compositor: Follows the Content-Style Modulation (CosMo) architecture. A Content Modulator (CM) using a Disentangled Multi-modal Non-local (DMNL) block updates the source visual features based on text, while a Style Modulator (SM) with multi-modal attention extracts style information. CM and SM outputs are concatenated and projected by a linear fully connected layer to form composed query feature fsf_s.
    • Optimization: SGD optimizer with mini-batch size 32, base learning rate 2×10−22 \times 10^{-2}, trained for 50 epochs. A step learning rate scheduler decays the learning rate by a factor of 10 at epoch 45. Augmenter parameters are w1=1,w2=1w_1 = 1, w_2 = 1, and initial loss weight is γ0=1\gamma_0 = 1.
    • Datasets:
      • FashionIQ: 75,384 crawled images with 46,609 training pairs across three subsets: Dress, Shirt, Toptee.
      • Shoes: 10,751 pairs (10,000 train, 4,658 test) with relative text feedback describing visual changes.
      • Fashion200k: 172,000 training images across five subsets (dresses, jackets, pants, skirts, tops) and 33,480 evaluation queries.
    • Inference & Metrics: Gallery images ItI_t are ranked against composed query fsf_s by cosine similarity. Performance is reported using Recall@1, Recall@10, and Recall@50 (R@1, R@10, R@50).
  5. Knowl 5 — Performance Comparison on the FashionIQ Benchmark

    data/table

    The table below details retrieval accuracy on the FashionIQ benchmark across the Dress, Shirt, and Toptee subsets, as well as their average. Models trained with uncertainty regularization achieve substantial recall gains over the baseline and outperform prior methods using single and ensembled visual backbones.

    Method Visual Backbone Dress Shirt Toptee Average
    R@10 R@50 R@10 R@50 R@10 R@50 R@10 R@50
    MRN ResNet-152 12.32 32.18 15.88 34.33 18.11 36.33 15.44 34.28
    FiLM ResNet-50 14.23 33.34 15.04 34.09 17.30 37.68 15.52 35.04
    TIRG ResNet-17 14.87 34.66 18.26 37.89 19.08 39.62 17.40 37.39
    Pic2Word ViT-L/14 20.00 40.20 26.20 43.60 27.90 47.40 24.70 43.70
    VAL ResNet-50 21.12 42.19 21.03 43.44 25.64 49.49 22.60 45.04
    ARTEMIS ResNet-50 27.16 52.40 21.78 54.83 29.20 43.64 26.05 50.29
    CoSMo ResNet-50 25.64 50.30 24.90 49.18 29.21 57.46 26.58 52.31
    DCNet ResNet-50 28.95 56.07 23.95 47.30 30.44 58.29 27.78 53.89
    FashionViL ResNet-50 28.46 54.24 22.33 46.07 29.02 57.93 26.60 52.74
    FashionViL* ResNet-50 33.47 59.94 25.17 50.39 34.98 60.79 31.21 57.04
    Baseline ResNet-50 24.80 52.35 27.70 55.71 33.40 63.64 28.63 57.23
    Ours ResNet-50 30.60 57.46 31.54 58.29 37.37 68.41 33.17 61.39
    CLVC-Net ResNet-50×\times2 29.85 56.47 28.75 54.76 33.50 64.00 30.70 58.41
    Ours ResNet-50×\times2 31.25 58.35 31.69 60.65 39.82 71.07 34.25 63.36
    CLIP4Cir ResNet-50×\times4 31.63 56.67 36.36 58.00 38.19 62.42 35.39 59.03
    Ours ResNet-50×\times4 32.61 61.34 33.23 62.55 41.40 72.51 35.75 65.47

    FashionViL utilizes external pre-training datasets.

    With a single ResNet-50 backbone, uncertainty regularization improves average R@10 by +4.54% (from 28.63% to 33.17%) and R@50 by +4.16% (from 57.23% to 61.39%) over the baseline. The single-model result (61.39% R@50) surpasses the 4-model ensemble of CLIP4Cir (59.03% R@50). When ensembled across four initializations (4×4 \times ResNet-50), the method achieves 35.75% R@10 and 65.47% R@50.

  6. Knowl 6 — Performance Comparison on Shoes and Fashion200k Benchmarks

    data/table

    Evaluation of single-model composed image retrieval on the Shoes and Fashion200k benchmarks shows consistent gains from uncertainty regularization.

    Method Shoes Fashion200k
    R@1 R@10 R@50 R@1 R@10 R@50
    MRN 11.74 41.70 67.01 13.4 40.0 61.9
    FiLM 10.19 38.89 68.30 12.9 39.5 61.9
    TIRG 12.60 45.45 69.39 14.1 42.5 63.8
    VAL 16.49 49.12 73.53 21.2 49.0 68.8
    CoSMo 16.72 48.36 75.64 23.3 50.4 69.3
    DCNet - 53.82 79.33 - 46.9 67.6
    ARTEMIS 18.72 53.11 79.31 21.5 51.1 70.5
    Baseline 15.26 49.48 76.46 19.5 46.7 67.8
    Ours 18.41 53.63 79.84 21.8 52.1 70.2

    On the Shoes benchmark, the proposed method yields +3.15% R@1, +4.15% R@10, and +3.38% R@50 over the baseline. On the Fashion200k benchmark, it achieves +2.3% R@1, +5.4% R@10, and +2.40% R@50 over the baseline.

  7. Knowl 7 — Ablation of Uncertainty Regularization vs. Dropout and Source Feature Augmentation

    empirical result

    Ablation experiments on the Shoes dataset evaluate the specific contribution of target-side uncertainty regularization compared to feature dropout, direct mean/variance regression, and source-side augmentation (Average is calculated as (R@10+R@50)/2(\text{R@10} + \text{R@50})/2):

    • Baseline (InfoNCE loss Linfo\mathcal{L}_{info} only): 49.48% R@10, 76.46% R@50, Average 62.97%.
    • Linfo\mathcal{L}_{info} + Dropout (drop rate = 0.2): 49.74% R@10, 76.37% R@50, Average 63.06%.
    • Linfo\mathcal{L}_{info} + Dropout (drop rate = 0.5): 49.00% R@10, 75.83% R@50, Average 62.42%.
    • Learned Variance Prediction via Extra Branches: 50.14% R@10, 77.89% R@50, Average 64.01%.
    • Augmenting Source Visual Feature fsIf_s^I: 52.20% R@10, 77.75% R@50, Average 64.98%. Applying noise to fsIf_s^I rather than ftf_t drops performance by 1.43% on R@10 and 2.09% on R@50, because perturbing the source image creates a many-to-one retrieval dynamic conflicting with one-to-many retrieval.
    • Uncertainty Regularization Only (Lu\mathcal{L}_u alone): 50.83% R@10, 77.41% R@50, Average 64.12%.
    • Full Proposed Method (Linfo+Lu\mathcal{L}_{info} + \mathcal{L}_u): 53.63% R@10, 79.84% R@50, Average 66.74%.

    Dropout fails to provide meaningful recall improvements because its drop rate is fixed and independent of feature statistics. Combining fine-grained matching with feature-derived target uncertainty regularization yields the highest recall.

  8. Knowl 8 — Sensitivity Analysis of Noise Scaling and Dynamic Weight Annealing

    empirical result

    Hyperparameter sensitivity evaluated on the Shoes dataset reveals the behavior of target noise scaling factors (w1,w2)(w_1, w_2) in f^t=α′fˉt+β′\hat{f}_t = \alpha' \bar{f}_t + \beta' (where α′∼N(1,w1σt)\alpha' \sim \mathcal{N}(1, w_1 \sigma_t), β′∼N(μt,w2σt)\beta' \sim \mathcal{N}(\mu_t, w_2 \sigma_t)) and dynamic balancing weight γ\gamma:

    1. Noise Scale (w1,w2w_1, w_2):

      • Fixing w2=1w_2 = 1 and varying w1∈{0.1,0.2,0.5,0.7,1,2,5,7,10}w_1 \in \{0.1, 0.2, 0.5, 0.7, 1, 2, 5, 7, 10\} yields average recall scores of 62.00%, 63.85%, 64.94%, 65.45%, 66.74% (best at w1=1w_1 = 1), 62.76%, 64.68%, 62.57%, and 62.76%.
      • Fixing w1=1w_1 = 1 and varying w2∈{0.1,0.2,0.5,0.7,1,2,5,7,10}w_2 \in \{0.1, 0.2, 0.5, 0.7, 1, 2, 5, 7, 10\} yields average recall scores of 65.07%, 64.82%, 64.22%, 64.41%, 66.74% (best at w2=1w_2 = 1), 64.12%, 62.52%, 62.89%, and 60.87%.
      • Identical scaling w1=1,w2=1w_1 = 1, w_2 = 1 performs best as it faithfully mirrors the data distribution, while large noise scales remain stable due to the adaptive loss weighting term 2σt22\sigma_t^2.
    2. Initial Weight γ0\gamma_0 in Annealing Schedule γ=exp⁡(−γ0⋅eE)\gamma = \exp(-\gamma_0 \cdot \frac{e}{E}):

      • γ0=0.1\gamma_0 = 0.1: 50.80% R@10, 78.64% R@50, Average 64.72%.
      • γ0=0.5\gamma_0 = 0.5: 51.20% R@10, 79.12% R@50, Average 65.76%.
      • γ0=1.0\gamma_0 = 1.0: 53.63% R@10, 79.84% R@50, Average 66.74% (optimal).
      • γ0=2.0\gamma_0 = 2.0: 49.66% R@10, 77.63% R@50, Average 63.65%.
      • γ0=5.0\gamma_0 = 5.0: 50.14% R@10, 76.32% R@50, Average 63.23%.
      • γ0→∞\gamma_0 \to \infty (baseline): 49.48% R@10, 76.46% R@50, Average 62.97%.
    3. Static Balance Weight γ\gamma (without Annealing):

      • γ=0.2\gamma = 0.2: 30.90% R@10, 63.12% R@50, Average 47.01%.
      • γ=0.5\gamma = 0.5: 41.87% R@10, 73.11% R@50, Average 57.49%.
      • γ=0.8\gamma = 0.8: 41.87% R@10, 72.39% R@50, Average 57.13%. Constant weights perform substantially worse than dynamic annealing because maintaining a fixed coarse loss suppresses fine-grained discrimination in later epochs.
  9. Knowl 9 — Compatibility with Other Vision-Language Retrieval Frameworks

    empirical result

    Uncertainty regularization functions as an orthogonal, model-agnostic training strategy that enhances existing vision-language retrieval frameworks:

    • CLIP4Cir on FashionIQ:

      • CLIP4Cir (single ResNet-50 backbone): 32.36% R@10, 56.74% R@50.
      • CLIP4Cir + Uncertainty Regularization: 34.19% R@10 (+1.83%), 59.23% R@50 (+2.49%).
    • BLIP on Shoes:

      • BLIP: 43.93% R@10, 70.21% R@50.
      • BLIP + Uncertainty Regularization: 48.40% R@10 (+4.47%), 75.29% R@50 (+5.08%).
    • Evaluation on Ambiguous/Coarse Queries: When evaluated exclusively on short queries containing fewer than 5 words in the FashionIQ Dress split (6,246 out of 10,942 test pairs), uncertainty regularization improves R@10 by +1.46% and R@50 by +3.34% over the baseline, confirming enhanced robustness for underspecified user queries.

Coverage note — No substantial contributed material was omitted. Qualitative visual examples from Figure 3 were excluded in favor of comprehensive quantitative retrieval results, architectural specifications, and ablation analyses.

References

  1. 1.Alberto Baldrati, Marco Bertini, Tiberio Uricchio, and Alberto Del Bimbo. Conditioned and composed image retrieval combining and partially fine-tuning clip-based features. In CVPR Workshops, 2022.
  2. 2.Jie Chang, Zhonghao Lan, Changmao Cheng, and Yichen Wei. Data uncertainty learning in face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5710–5719, 2020.
  3. 3.Yanbei Chen, Shaogang Gong, and Loris Bazzani. Image search with text feedback by visiolinguistic attention learning. In CVPR, June 2020.
  4. 4.Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio De Rezende, Yannis Kalantidis, and Diane Larlus. Probabilistic embeddings for cross-modal retrieval. In CVPR, 2021.
  5. 5.Ginger Delmas, Rafael S Rezende, Gabriela Csurka, and Diane Larlus. Artemis: Attention-based retrieval with text-explicit matching and implicit similarity. In ICLR, 2022.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June 2019. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  7. 7.Zhaopeng Dou, Zhongdao Wang, Weihua Chen, Yali Li, and Shengjin Wang. Reliability-aware prediction via uncertainty learning for person image retrieval. In ECCV, 2022.
  8. 8.Yarin Gal. Uncertainty in Deep Learning. PhD thesis, University of Cambridge, 2016.
  9. 9.Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In ICML, pp. 1050–1059. PMLR, 2016.
  10. 10.Spyros Gidaris and Nikos Komodakis. Dynamic few-shot visual learning without forgetting. In CVPR, June 2018.
  11. 11.Jacob Goldberger, Geoffrey E Hinton, Sam Roweis, and Russ R Salakhutdinov. Neighbourhood components analysis. In NeurIPS, 2004.
  12. 12.Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, and Rogerio Feris. Dialog-based interactive image retrieval. In NeurIPS, 2018. URL http://papers.nips.cc/paper/7348-dialog-based-interactive-image-retrieval.pdf.
  13. 13.Xiao Han, Licheng Yu, Xiatian Zhu, Li Zhang, Yi-Zhe Song, and Tao Xiang. Fashionvil: Fashionfocused vision-and-language representation learning. In ECCV, 2022.
  14. 14.Xintong Han, Zuxuan Wu, Phoenix X. Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S. Davis. Automatic spatially-aware fashion concept discovery. In ICCV, 2017. doi: 10.1109/ICCV.2017.163.
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. CVPR, 2015.
  16. 16.Xuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, and Xin Eric Wang. CPL: Counterfactual prompt learning for vision and language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022.
  17. 17.Alexander Hermans, Lucas Beyer, and Bastian Leibe. In defense of the triplet loss for person re-identification. arXiv:1703.07737, 2017.
  18. 18.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 1997.
  19. 19.Michael Jordan, Jon Kleinberg, and Schölkopf Bernhard. Bayesian networks and decision graphs. Springer New York, 2007.
  20. 20.Alex Kendall and Yarin Gal. What uncertainties do we need in bayesian deep learning for computer vision? In NeurIPS, 2017. URL https://proceedings.neurips.cc/paper/2017/file/2650d6089a6d640c5e85b2b88265dc2b-Paper.pdf.
  21. 21.Jin-Hwa Kim, Sang-Woo Lee, Donghyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. Multimodal residual learning for visual qa. In NeurIPS, 2016. URL https://proceedings.neurips.cc/paper/2016/file/9b04d152845ec0a378394003c96da594-Paper.pdf.
  22. 22.Jongseok Kim, Youngjae Yu, Hoeseong Kim, and Gunhee Kim. Dual Compositional Learning in Interactive Image Retrieval. In AAAI, 2021.
  23. 23.S. Lee, D. Kim, and B. Han. Cosmo: Content-style modulation for image retrieval with text feedback. In CVPR, 2021.
  24. 24.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pretraining for unified vision-language understanding and generation. In International Conference on Machine Learning, pp. 12888–12900. PMLR, 2022.
  25. 25.Pan Li, Da Li, Wei Li, Shaogang Gong, Yanwei Fu, and Timothy M. Hospedales. A simple feature augmentation for domain generalization. In 2021ICCVICCV, pp. 8866–8875, 2021. doi: 10.1109/ICCV48922.2021.00876.
  26. 26.Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. On the variance of the adaptive learning rate and beyond. CoRR, abs/1908.03265, 2019a. URL http://arxiv.org/abs/1908.03265.
  27. 27.Shaoteng Liu, Jingjing Chen, Liangming Pan, Chong-Wah Ngo, Tat-Seng Chua, and Yu-Gang Jiang. Hyperbolic visual embedding learning for zero-shot recognition. In CVPR, 2020.
  28. 28.Si Liu, Zheng Song, Meng Wang, Changsheng Xu, Hanqing Lu, and Shuicheng Yan. Street-to-shop: Cross-scenario clothing retrieval via parts alignment and auxiliary set. In ACM Multimedia, pp. 1335–1336, 2012.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv:1907.11692, 2019b.
  30. 30.Zheyuan Liu, Cristian Rodriguez, Damien Teney, and Stephen Gould. Image retrieval on real-life images with pre-trained vision-and-language models. In ICCV, 2021.
  31. 31.Yair Movshovitz-Attias, Alexander Toshev, Thomas K. Leung, Sergey Ioffe, and Saurabh Singh. No fuss distance metric learning using proxies. In ICCV, 2017.
  32. 32.Antonio Norelli, Marco Fumero, Valentino Maiorca, Luca Moschella, Emanuele Rodolà, and Francesco Locatello. Asif: Coupled data turns unimodal models to multimodal without training. arXiv preprint arXiv:2210.01738, 2022.
  33. 33.Seong Joon Oh, Kevin Murphy, Jiyan Pan, Joseph Roth, Florian Schroff, and Andrew Gallagher. Modeling uncertainty with hedged instance embedding. arXiv:1810.00319, 2018.
  34. 34.Hyun Oh Song, Yu Xiang, Stefanie Jegelka, and Silvio Savarese. Deep metric learning via lifted structured feature embedding. In CVPR, 2016.
  35. 35.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv:1807.03748, 2018.
  36. 36.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS. 2019a.
  37. 37.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, volume 32, 2019b. URL https://proceedings.neurips.cc/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf.
  38. 38.Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville. Film: Visual reasoning with a general conditioning layer. In AAAI, 2018.
  39. 39.James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman. Object retrieval with large vocabularies and fast spatial matching. In CVPR, pp. 1–8. IEEE, 2007.
  40. 40.Leila Pishdad, Ran Zhang, Konstantinos G Derpanis, Allan Jepson, and Afsaneh Fazly. Uncertainty-based cross-modal retrieval with probabilistic representations. arXiv:2204.09268, 2022.
  41. 41.Janis Postels, Mattia Segù, Tao Sun, Luc Van Gool, Fisher Yu, and Federico Tombari. On the practicality of deterministic epistemic uncertainty. CoRR, abs/2107.00649, 2021. URL https://arxiv.org/abs/2107.00649.
  42. 42.Leigang Qu, Meng Liu, Jianlong Wu, Zan Gao, and Liqiang Nie. Dynamic modality interaction modeling for image-text retrieval. In ACM SIGIR, 2021.
  43. 43.Filip Radenović, Giorgos Tolias, and Ondřej Chum. Fine-tuning cnn image retrieval with no human annotation. IEEE transactions on pattern analysis and machine intelligence, 41(7):1655–1668, 2018.
  44. 44.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021.
  45. 45.Herbert Robbins and Sutton Monro. A Stochastic Approximation Method. The Annals of Mathematical Statistics, 22(3):400 – 407, 1951. doi: 10.1214/aoms/1177729586. URL https://doi.org/10.1214/aoms/1177729586.
  46. 46.Negar Rostamzadeh, Seyedarian Hosseini, Thomas Boquet, Wojciech Stokowiec, Ying Zhang, Christian Jauvin, and Chris Pal. Fashion-gen: The generative fashion dataset and challenge. arXiv:1806.08317, 2018.
  47. 47.Kuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li, Chen-Yu Lee, Kate Saenko, and Tomas Pfister. Pic2word: Mapping pictures to words for zero-shot composed image retrieval. arXiv preprint arXiv:2302.03084, 2023.
  48. 48.Minchul Shin, Yoonjae Cho, Byungsoo Ko, and Geonmo Gu. Rtic: Residual learning for text and image composition using graph convolutional network. arXiv preprint arXiv:2104.03015, 2021.
  49. 49.Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. In NeurIPS, 2017. URL https://proceedings.neurips.cc/paper/2017/file/cb8da6767461f2812ae4290eac7cbc42-Paper.pdf.
  50. 50.Yifan Sun, Yuke Zhu, Yuhan Zhang, Pengkun Zheng, Xi Qiu, Chi Zhang, and Yichen Wei. Dynamic metric learning: Towards a scalable metric space to accommodate multiple semantic scales. In CVPR, 2021.
  51. 51.Prune Truong, Martin Danelljan, Radu Timofte, and Luc Van Gool. Pdc-net+: Enhanced probabilistic dense correspondence network. CoRR, abs/2109.13912, 2021. URL https://arxiv.org/abs/2109.13912.
  52. 52.Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. Composing text and image for image retrieval - an empirical odyssey. CoRR, abs/1812.07119, 2018. URL http://arxiv.org/abs/1812.07119.
  53. 53.Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. Composing text and image for image retrieval-an empirical odyssey. In CVPR, 2019a.
  54. 54.Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. Composing text and image for image retrieval-an empirical odyssey. In CVPR, 2019b.
  55. 55.Frederik Warburg, Martin Jørgensen, Javier Civera, and Søren Hauberg. Bayesian triplet loss: Uncertainty quantification in image retrieval. ICCV, pp. 12138–12148, 2021.
  56. 56.Haokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan, and Liqiang Nie. Comprehensive linguisticvisual composition network for image retrieval. In ACM SIGIR, SIGIR ’21, 2021. ISBN 9781450380379. doi: 10.1145/3404835.3462967. URL https://doi.org/10.1145/3404835.3462967.
  57. 57.Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris. The fashion iq dataset: Retrieving images by combining side information and relative natural language feedback. CVPR, 2021.
  58. 58.Xuewen Yang, Heming Zhang, Di Jin, Yingru Liu, Chi-Hao Wu, Jianchao Tan, Dongliang Xie, Jue Wang, and Xin Wang. Fashion captioning: Towards generating accurate descriptions with semantic rewards. In ECCV, pp. 1–17. Springer, 2020.
  59. 59.Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. Stacked attention networks for image question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 21–29, 2016.
  60. 60.Tianyuan Yu, Da Li, Yongxin Yang, Timothy Hospedales, and Tao Xiang. Robust person reidentification by modelling feature uncertainty. In ICCV, pp. 552–561, 2019. doi: 10.1109/ICCV.2019.00064.
  61. 61.Xinyu Zhang, Dongdong Li, Zhigang Wang, Jian Wang, Errui Ding, Javen Qinfeng Shi, Zhaoxiang Zhang, and Jingdong Wang. Implicit sample extension for unsupervised person re-identification. In CVPR, 2022.
  62. 62.Ying Zhang, Tao Xiang, Timothy M. Hospedales, and Huchuan Lu. Deep mutual learning. In CVPR, pp. 4320–4328, 2018. doi: 10.1109/CVPR.2018.00454.
  63. 63.Yida Zhao, Yuqing Song, and Qin Jin. Progressive learning for image retrieval with hybrid-modality queries. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’22, pp. 1012–1021, New York, NY, USA, 2022. Association for Computing Machinery. ISBN 9781450387323. doi: 10.1145/3477495.3532047. URL https://doi.org/10.1145/3477495.3532047.
  64. 64.Liang Zheng, Yi Yang, and Qi Tian. Sift meets cnn: A decade survey of instance retrieval. IEEE transactions on pattern analysis and machine intelligence, 40(5):1224–1244, 2017a.
  65. 65.Zhedong Zheng and Yi Yang. Rectifying pseudo label learning via uncertainty estimation for domain adaptive semantic segmentation. International Journal of Computer Vision (IJCV), 2021. doi: 10.1007/s11263-020-01395-y. doi:10.1007/s11263-020-01395-y.
  66. 66.Zhedong Zheng, Liang Zheng, and Yi Yang. A discriminatively learned cnn embedding for person reidentification. ACM transactions on multimedia computing, communications, and applications (TOMM), 14(1):1–20, 2017b.
  67. 67.Zhedong Zheng, Yunchao Wei, and Yi Yang. University-1652: A multi-view multi-source benchmark for drone-based geo-localization. In ACM MM, 2020a.
  68. 68.Zhedong Zheng, Liang Zheng, Michael Garrett, Yi Yang, Mingliang Xu, and Yi-Dong Shen. Dualpath convolutional image-text embeddings with instance loss. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 16(2):1–23, 2020b.

Citation

MLA
Chen, Y., et al. “Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization”. arXiv, 2022, http://arxiv.org/abs/2211.07394v6.
APA
Chen, Y., Zheng, Z., Ji, W., Qu, L., & Chua, T.-S. (2022). Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization. arXiv. http://arxiv.org/abs/2211.07394v6
Chicago
Chen, Y., Z. Zheng, W. Ji, L. Qu, and T.-S. Chua. 2022. “Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization”. arXiv. http://arxiv.org/abs/2211.07394v6.
Harvard
Chen, Y. et al. (2022) “Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.07394v6.
Vancouver
1. Chen Y, Zheng Z, Ji W, Qu L, Chua T-S (2022) Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization. arXiv

BibTeX

@article{chen2022composed,
  title = {Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization},
  author = {Chen, Yiyang and Zheng, Zhedong and Ji, Wei and Qu, Leigang and Chua, Tat-Seng},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.07394v6},
  eprint = {2211.07394}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors