Sketching without Worrying: Noise-Tolerant Sketch-Based Image Retrieval

Ayan Kumar BhuniaSubhadeep KoleyAbdullah Faiz Ur Rahman KhiljiAneeshan SainPinaki Nath ChowdhuryTao XiangYi-Zhe Song

article2022CVPR56 citations

Proposes a reinforcement learning-based stroke subset selector that filters out noisy or detrimental strokes from amateur drawings to boost fine-grained sketch-based image retrieval accuracy without retraining underlying retrieval models.

Listen

Sketch-based image retrieval enables users to search for specific photos by drawing on interactive touchscreens. However, widespread adoption faces a major hurdle termed the fear-to-sketch barrier, where amateur users lack confidence in their drawing ability. The article investigates this issue and discovers that retrieval failures are rarely caused by an inability to sketch basic outlines. Instead, failures stem from irrelevant or inaccurate strokes that act as noise and severely degrade search accuracy. The main objective of the article is to demonstrate an intelligent preprocessing module that detects and removes these noisy strokes before image retrieval, enabling high-precision visual search from imperfect freehand drawings.

To achieve this, the authors designed a stroke subset selector operating directly on sequential sketch coordinates. The system pairs a hierarchical recurrent neural network, which models relationships between individual strokes and the overall drawing, with a reinforcement learning training scheme. Because drawing coordinates must be converted into raster images to query standard retrieval models, the process cannot use traditional derivative-based training. The selector is trained using an actor-critic algorithm where a fixed retrieval model serves as the critic, rewarding the selector whenever it picks a stroke combination that improves retrieval accuracy. The authors evaluated this framework on standard benchmark datasets comprising thousands of shoe and chair sketches.

Experimental results show that the proposed selector substantially improves search performance across all settings. When added to standard retrieval models, top-ranked retrieval accuracy increased by approximately 8% to 10%, achieving 43.7% accuracy on shoes and 64.8% on chairs, outperforming existing state-of-the-art systems. In simulated extreme conditions with heavy synthetic noise, the selector restored retrieval accuracy from 13.4% back up to 37.2%. Additionally, the evaluation score produced by the critic network reliably measures whether an incomplete sketch has enough information to search effectively, allowing systems to reduce unnecessary visual processing by 42.2% during interactive drawing.

These findings demonstrate that overcoming user drawing limitations does not require building massive, complicated retrieval architectures. Instead, lightweight preprocessing in the stroke coordinate domain offers a highly effective, plug-and-play solution. This approach improves system robustness and reduces user frustration while maintaining low computational overhead, adding only about 18.3% extra central processing unit time and roughly 22.4% more mathematical operations compared to baseline retrieval methods.

Organizations developing interactive visual search or sketch-based applications should adopt stroke subset selection as a standard front-end filter. Teams can also utilize the trained selector as a data augmentation tool to generate clean partial sketches for downstream training, avoiding unstable alternative methods. Future work should focus on extending this framework beyond simple product categories to multi-object complex scenes, as current testing was confined to isolated single-object benchmark datasets. Overall, the findings provide strong confidence that noisy stroke elimination is a practical and scalable solution for real-world sketch applications.

arXiv: 2203.14817
Cover for Sketching without Worrying: Noise-Tolerant Sketch-Based Image Retrieval

Abstract

Sketching enables many exciting applications, notably, image retrieval. The fear-to-sketch problem (i.e., “I can’t sketch”) has however proven to be fatal for its widespread adoption. This paper tackles this “fear” head on, and for the first time, proposes an auxiliary module for existing retrieval models that predominantly lets the users sketch without having to worry. We first conducted a pilot study that revealed the secret lies in the existence of noisy strokes, but not so much of the “I can’t sketch”. We consequently design a stroke subset selector that detects noisy strokes, leaving only those which make a positive contribution towards successful retrieval. Our Reinforcement Learning based formulation quantifies the importance of each stroke present in a given subset, based on the extent to which that stroke contributes to retrieval. When combined with pre-trained retrieval models as a pre-processing module, we achieve a significant gain of 8%-10% over standard baselines and in turn report new state-of-the-art performance. Last but not least, we demonstrate the selector once trained, can also be used in a plug-and-play manner to empower various sketch applications in ways that were not previously possible.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Pilot Study: What's Wrong with FG-SBIR?
  • 4. Noisy Stroke Tolerant FG-SBIR
  • 4.1. Stroke Subset Selector
  • 4.2. Training Procedure
  • 5. Applications of Stroke-Subset Selector
  • 6. Experiments
  • 6.1. Performance Analysis
  • 6.2. Further Analysis and Insights
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Two-Level Hierarchical Architecture for Sketch Stroke Subset Selection

    model/method

    To select a clean subset of strokes from an amateur sketch in vector modality V\mathcal{V}, a two-level hierarchical Recurrent Neural Network (RNN) processes the sketch sequentially. A sketch SVS_V with KK strokes is represented as SV=(s1,s2,…,sK)S_V = (s_1, s_2, \dots, s_K), where each stroke sis_i contains a sequence of NiN_i absolute 2D point coordinates si=(v1i,v2i,…,vNii)∈RNi×2s_i = (v_1^i, v_2^i, \dots, v_{N_i}^i) \in \mathbb{R}^{N_i \times 2}.

    The architecture consists of three components grouped as Xθ={Eθ,Rθ,Cθ}\mathcal{X}_\theta = \{E_\theta, R_\theta, C_\theta\}:

    1. Local Stroke-Embedding Network (EθE_\theta): A one-layer Long Short-Term Memory (LSTM) network with hidden size ds=128d_s = 128, shared across all strokes, processes the point sequence si∈RNi×2s_i \in \mathbb{R}^{N_i \times 2} to produce a localized representation fsil∈Rdsf_{s_i}^l \in \mathbb{R}^{d_s} from its final hidden state.
    2. Global Relational Network (RθR_\theta): A one-layer LSTM with hidden size ds=128d_s = 128 takes the sequence of local stroke embeddings (fs1l,fs2l,…,fsKl)∈RK×ds(f_{s_1}^l, f_{s_2}^l, \dots, f_{s_K}^l) \in \mathbb{R}^{K \times d_s} as input. Its final hidden state fg∈Rdsf^g \in \mathbb{R}^{d_s} captures the global semantic context of the full sketch.
    3. Compositional Fusion and Policy Head (CθC_\theta): Local and global features are fused via a residual connection followed by layer normalization: f^si=LayerNorm(fsil+fg)∈Rds\hat{f}_{s_i} = \text{LayerNorm}(f_{s_i}^l + f^g) \in \mathbb{R}^{d_s} A shared linear layer followed by softmax computes the categorical action probability distribution over two discrete actions a∈{select,ignore}a \in \{\text{select}, \text{ignore}\}: p(ai∣si)=softmax(WXf^si+bX)p(a_i \mid s_i) = \text{softmax}(W_{\mathcal{X}} \hat{f}_{s_i} + b_{\mathcal{X}}) where WX∈R2×dsW_{\mathcal{X}} \in \mathbb{R}^{2 \times d_s} and bX∈R2b_{\mathcal{X}} \in \mathbb{R}^2.
  2. Knowl 2 — Actor-Critic PPO Framework for Non-Differentiable Stroke Subset Selection

    algorithm

    Because rendering a selected vector stroke subset SV′⊂SVS_V' \subset S_V into a raster image SI′=R(SV′)S_I' = \mathcal{R}(S_V') involves a non-differentiable rasterization operator R:V→I\mathcal{R}: \mathcal{V} \to \mathcal{I}, gradient-based optimization cannot backpropagate directly from a pre-trained retrieval model F\mathcal{F} to the stroke selector Xθ\mathcal{X}_\theta. The stroke selection is therefore cast as a Markov Decision Process (MDP) optimized with the Actor-Critic variant of Proximal Policy Optimization (PPO) using a clipped surrogate objective.

    Input: Training dataset of sketch-photo pairs (S,P)(S, P), pre-trained FG-SBIR model F\mathcal{F}, pre-computed gallery photo embeddings G∈RM×DG \in \mathbb{R}^{M \times D}, episode length T=5T = 5, clipping factor ϵ=0.2\epsilon = 0.2, coefficients c1=0.5,c2=0.01c_1 = 0.5, c_2 = 0.01.
    Output: Optimized stroke subset selector policy Xθ\mathcal{X}_\theta.
    Initialize actor-critic parameters θ\theta, old policy parameters θold←θ\theta_{\text{old}} \leftarrow \theta, replay buffer B←∅\mathcal{B} \leftarrow \emptyset.
    for each training iteration do
        for each sketch SV=(s1,…,sK)S_V = (s_1, \dots, s_K) in batch do
            for step t=1t = 1 to TT do
                Compute stroke features f^si\hat{f}_{s_i} using EθE_\theta and RθR_\theta.
                Compute stroke selection distribution pθ(ai∣si)=softmax(WXf^si+bX)p_\theta(a_i \mid s_i) = \text{softmax}(W_{\mathcal{X}} \hat{f}_{s_i} + b_{\mathcal{X}}).
                Sample binary actions ai∼Categorical(pθ(ai∣si))a_i \sim \text{Categorical}(p_\theta(a_i \mid s_i)) for i∈{1,…,K}i \in \{1, \dots, K\}.
                Form subset SV′={si∣ai=select}S_V' = \{s_i \mid a_i = \text{select}\}.
                Rasterize subset SI′=R(SV′)S_I' = \mathcal{R}(S_V').
                Extract feature vector F(SI′)\mathcal{F}(S_I').
                Compute gallery ranking index rank\text{rank} of ground-truth photo PP using GG.
                Compute triplet loss LTriplet\mathcal{L}_{\text{Triplet}} on F(SI′)\mathcal{F}(S_I').
                Compute scalar reward R=ω11rank+ω2(−LTriplet)R = \omega_1 \frac{1}{\text{rank}} + \omega_2 (-\mathcal{L}_{\text{Triplet}}).
                Store transition (SV,{ai},R,pθold(ai∣si))(S_V, \{a_i\}, R, p_{\theta_{\text{old}}}(a_i \mid s_i)) in B\mathcal{B}.
            end for
        end for
        Sample transitions from B\mathcal{B}.
        Compute probability ratios ri(θ)=log⁡pθ(ai∣si)log⁡pθold(ai∣si)r_i(\theta) = \frac{\log p_\theta(a_i \mid s_i)}{\log p_{\theta_{\text{old}}}(a_i \mid s_i)}.
        Compute LCPI(θ)=−1K∑i=1Kri(θ)RL^{\text{CPI}}(\theta) = -\frac{1}{K} \sum_{i=1}^K r_i(\theta) R.
        Compute LCLIP(θ)=−1K∑i=1Kclip(ri(θ),1−ϵ,1+ϵ)RL^{\text{CLIP}}(\theta) = -\frac{1}{K} \sum_{i=1}^K \text{clip}(r_i(\theta), 1-\epsilon, 1+\epsilon) R.
        Compute actor surrogate LA(θ)=−1K∑i=1Kmin⁡(LCPI,LCLIP)L^A(\theta) = -\frac{1}{K} \sum_{i=1}^K \min(L^{\text{CPI}}, L^{\text{CLIP}}).
        Predict state value Vθ(S)V_\theta(S) via linear head on pooled stroke features.
        Compute full loss LAC(θ)=−1K∑i=1K(LA−c1(Vθ(S)−R)2+c2En)L^{\text{AC}}(\theta) = -\frac{1}{K} \sum_{i=1}^K (L^A - c_1(V_\theta(S) - R)^2 + c_2 E_n), where EnE_n is the policy entropy bonus.
        Update θ\theta using Adam optimizer.
        Every 20 iterations, set θold←θ\theta_{\text{old}} \leftarrow \theta.
    end for
  3. Knowl 3 — Combined Ranking and Feature-Space Triplet Reward Function

    equation

    The scalar reward RR driving the policy optimization of the stroke subset selector balances rank minimization in the retrieval gallery with distance minimization in the embedding space:

    R=ω1⋅1rank+ω2⋅(−LTriplet)R = \omega_1 \cdot \frac{1}{\text{rank}} + \omega_2 \cdot (-\mathcal{L}_{\text{Triplet}})

    where:

    • rank∈{1,2,…,M}\text{rank} \in \{1, 2, \dots, M\} is the 1-based ranking index of the paired ground-truth photo among all MM gallery photos, computed using Euclidean distances between the feature embedding of the rasterized subset sketch F(SI′)\mathcal{F}(S_I') and pre-computed gallery photo embeddings G∈RM×DG \in \mathbb{R}^{M \times D}.
    • LTriplet=max⁡{0,δ+−δ−+μ}\mathcal{L}_{\text{Triplet}} = \max\{0, \delta^+ - \delta^- + \mu\} is the margin triplet loss evaluated on normalized feature embeddings from the pre-trained retrieval network F\mathcal{F}, with positive pair distance δ+=∥F(aˉ)−F(pˉ)∥2\delta^+ = \|\mathcal{F}(\bar{a}) - \mathcal{F}(\bar{p})\|_2, negative pair distance δ−=∥F(aˉ)−F(nˉ)∥2\delta^- = \|\mathcal{F}(\bar{a}) - \mathcal{F}(\bar{n})\|_2, and margin hyperparameter μ>0\mu > 0.
    • ω1=1.0\omega_1 = 1.0 and ω2=1.0\omega_2 = 1.0 are weighting hyperparameters.
  4. Knowl 4 — Standard Fine-Grained Sketch-Based Image Retrieval Benchmark Performance

    data/table

    Evaluation of top-1 (Acc@1) and top-5 (Acc@5) retrieval accuracy on standard test splits of QMUL-Chair-V2 and QMUL-Shoe-V2 datasets under standard FG-SBIR settings demonstrates that adding the stroke subset selector as a pre-processing module to a pre-trained baseline Siamese model outperforms prior state-of-the-art architectures:

    Chair-V2 Shoe-V2
    Method Acc@1 Acc@5 Acc@1 Acc@5
    Triplet-SN 47.4% 71.4% 28.7% 63.5%
    Triplet-Attn-HOLEF 50.7% 73.6% 31.2% 66.6%
    Triplet-RL 51.2% 73.8% 30.8% 65.1%
    Mixed-Jigsaw 56.1% 75.3% 36.5% 68.9%
    Semi-Sup 60.2% 78.1% 39.1% 69.9%
    StyleMeUp 62.8% 79.6% 36.4% 68.1%
    Cross-Hier 62.4% 79.1% 36.2% 67.8%
    Baseline-Siamese 53.3% 74.3% 33.4% 67.8%
    Augmnt 54.1% 74.6% 33.9% 68.2%
    StyleMeUp+Augment 56.1% 76.9% 36.9% 69.9%
    Contrastive+Augment 58.8% 77.1% 37.6% 70.1%
    Upper-Limit (2K−12^K - 1 subsets) 78.6% 90.3% 66.3% 88.3%
    Linear-Limit (sequential drawing) 59.4% 77.3% 42.5% 73.2%
    Proposed (Selector + Baseline-Siamese) 64.8% 79.1% 43.7% 74.9%

    The proposed method improves the Baseline-Siamese model by 11.5% Acc@1 on Chair-V2 (53.3% to 64.8%) and by 10.3% Acc@1 on Shoe-V2 (33.4% to 43.7%), exceeding dedicated architectures without requiring fine-tuning of the underlying retrieval backbone.

  5. Knowl 5 — Pilot Study on Stroke Detriment and Stroke Subset Upper Limits

    empirical result

    In a preliminary empirical study evaluating stroke-by-stroke sequential rendering SkI=R([s1,s2,…,sk])S_k^I = \mathcal{R}([s_1, s_2, \dots, s_k]) on the QMUL-Shoe-V2 dataset using a pre-trained VGG-16 Siamese FG-SBIR model:

    1. Detrimental Stroke Frequency: In 43.44% of sequential drawing instances, adding a subsequent stroke causes the paired photo's retrieval rank to drop compared to the preceding stroke prefix Sk−1IS_{k-1}^I.
    2. Linear-Limit vs. Complete Sketch: Evaluating retrieval on the full sketch yields 33.43% Acc@1 and 67.81% Acc@5. Taking the best ranking achieved at any single prefix instant k∈{1,…,K}k \in \{1, \dots, K\} during the drawing episode (Linear-Limit) increases retrieval accuracy to 42.54% Acc@1 and 73.28% Acc@5.
    3. Combinatorial Upper-Limit: Evaluating all non-empty stroke subsets (2K−12^K - 1 possible combinations for a sketch of KK strokes, where KK averages 9 to 15) and recording whether any subset successfully retrieves the target yields an upper-bound performance of 66.37% Acc@1 and 88.31% Acc@5 on Shoe-V2 (and 78.6% Acc@1 / 90.3% Acc@5 on Chair-V2).
  6. Knowl 6 — Partial Sketch Retrievability Scoring and Inference Acceleration via Critic Network

    model/method

    The state-value critic network Vθ(S)V_\theta(S), trained to estimate the scalar reward 1/rank1/\text{rank}, functions as a computational gatekeeper for partial sketches during interactive sketching:

    1. Correlation with Retrieval Quality: Vector sketches evaluated at progressive 5% completion intervals demonstrate that higher critic-predicted scores correlate strongly with higher Average Ranking Percentile (ARP) of the ground-truth photo. Partial sketches exceeding a critic threshold of V(S)>1/5V(S) > 1/5 achieve 80.1% Acc@5 retrieval accuracy.
    2. Computational Savings in On-The-Fly FG-SBIR: By querying the lightweight vector-space critic network before executing rasterization R\mathcal{R} and CNN inference, and passing partial sketches only when V(S)>1/10V(S) > 1/10, the total number of expensive rasterization and image-retrieval passes is reduced by 42.2% with minimal change in early retrieval metrics: area under the ranking-percentile curve (r@Ar@A) changes from 85.78 to 85.07, and area under the reciprocal-rank curve (r@Br@B) changes from 21.10 to 20.98 on Shoe-V2.
  7. Knowl 7 — Stroke-Subset-Based Data Augmentation for Fine-Grained Retrieval

    model/method

    Instead of applying random stroke dropping (which frequently discards essential structural strokes and produces ambiguous partial sketches), the trained stroke selector policy pθ(ai∣si)p_\theta(a_i \mid s_i) provides a grounded sampling distribution to generate multiple valid partial versions of each training sketch.

    Sampling multiple subsets per training sketch via Categorical sampling from pθ(ai∣si)p_\theta(a_i \mid s_i) provides faithful synthetic training samples. When training a standard supervised triplet-loss model on these policy-augmented partial sketches for on-the-fly FG-SBIR, the model attains early-retrieval metrics of r@A=85.78r@A = 85.78 and r@B=21.10r@B = 21.10 on Shoe-V2, outperforming continuous reinforcement learning fine-tuning (r@A=85.38,r@B=21.24r@A = 85.38, r@B = 21.24) without incurring the instability of continuous RL training pipelines.

  8. Knowl 8 — Robustness to Synthetic Extreme Stroke Perturbations

    empirical result

    When test sketches from the QMUL-Shoe-V2 dataset are corrupted with synthetically injected noisy stroke patches, the baseline FG-SBIR model suffers severe performance degradation, dropping from 33.4% Acc@1 (67.8% Acc@5) on clean sketches down to 13.4% Acc@1 (44.9% Acc@5).

    Inserting the trained stroke subset selector as a vector-space pre-processing module upstream of the pre-trained retrieval model identifies and filters out the synthetic noise strokes, restoring retrieval performance to 37.2% Acc@1 and 68.2% Acc@5.

  9. Knowl 9 — Ablation Analysis on Architectural Components, RL Objectives, and Computational Overhead

    empirical result

    Ablation experiments conducted on the QMUL-Shoe-V2 dataset isolate the contribution of specific design choices:

    1. Hierarchical vs. Flat Point Modeling: Replacing the two-level hierarchical LSTM (local stroke LSTM + global relational LSTM) with a single-layer bidirectional LSTM processing raw coordinate points directly drops retrieval performance by 4.9% Acc@1 and 6.7% Acc@5. Replacing the LSTM with a Transformer architecture yields no measurable performance gain.
    2. RL Formulation: Optimizing with actor-critic PPO using a clipped surrogate objective outperforms the actor-only PPO variant by 1.7% in top-1 accuracy.
    3. Reward Function Composition: The combined dual-space reward (R=1rank−LTripletR = \frac{1}{\text{rank}} - \mathcal{L}_{\text{Triplet}}) outperforms a rank-only reward (R=1rankR = \frac{1}{\text{rank}}) by 1.2% in top-1 accuracy.
    4. Computational Cost: Pre-processing vector sketches with the stroke subset selector adds 22.4% multiply-add operations and 18.3% CPU runtime relative to the standard baseline FG-SBIR forward pass.

Coverage note — None was omitted; all primary contributions, models, pilot findings, algorithmic formulations, benchmark tables, application results, and ablation studies from the paper are fully covered.

References

  1. 1.Emre Aksan, Thomas Deselaers, Andrea Tagliasacchi, and Otmar Hilliges. Cose: Compositional stroke embeddings. In NeurIPS, 2021. 4
  2. 2.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. 5
  3. 3.Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Yongxin Yang, Tao Xiang, and Yi-Zhe Song. More photos are all you need: Semi-supervised learning for fine-grained sketch based image retrieval. In CVPR, 2021. 1, 2, 3, 4, 5, 6, 7
  4. 4.Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Yongxin Yang, Timothy M Hospedales, Tao Xiang, and Yi-Zhe Song. Vectorization and rasterization: Self-supervised learning for sketch and handwriting. In CVPR, 2021. 2, 3, 4
  5. 5.Ayan Kumar Bhunia, Ayan Das, Umar Riaz Muhammad, Yongxin Yang, Timothy M Hospedales, Tao Xiang, Yulia Gryaditskaya, and Yi-Zhe Song. Pixelor: A competitive sketching ai agent. so you think you can sketch? ACM TOG, 2020. 3, 4
  6. 6.Ayan Kumar Bhunia, Viswanatha Reddy Gajjala, Subhadeep Koley, Rohit Kundu, Aneeshan Sain, Tao Xiang, and Yi-Zhe Song. Doodle it yourself: Class incremental learning by drawing a few sketches. In CVPR, 2022. 2
  7. 7.Ayan Kumar Bhunia, Yongxin Yang, Timothy M Hospedales, Tao Xiang, and Yi-Zhe Song. Sketch less for more: On-the-fly fine-grained sketch-based image retrieval. In CVPR, 2020. 1, 2, 3, 4, 5, 6, 7, 8
  8. 8.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020. 7
  9. 9.Wengling Chen and James Hays. Sketchygan: Towards diverse and realistic sketch to image synthesis. In CVPR, 2018. 2
  10. 10.Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Viswanatha Reddy Gajjala, Aneeshan Sain, Tao Xiang, and Yi-Zhe Song. Partially does it: Towards scene-level fg-sbir with partial input. In CVPR, 2022. 1
  11. 11.John Collomosse, Tu Bui, and Hailin Jin. Livesketch: Query perturbations for guided sketch-based visual search. In CVPR, 2019. 1, 2
  12. 12.Sounak Dey, Pau Riba, Anjan Dutta, Josep Llados, and Yi-Zhe Song. Doodle to search: Practical zero-shot sketch-based image retrieval. In CVPR, 2019. 1, 2
  13. 13.Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Nghia Nguyen, Eric Patterson, Tien D Bui, and Ngan Le. Automatic face aging in videos via deep reinforcement learning. In CVPR, 2019. 3
  14. 14.Anjan Dutta and Zeynep Akata. Semantically tied paired cycle consistency for zero-shot sketch-based image retrieval. In CVPR, 2019. 1, 2
  15. 15.Arnab Ghosh, Richard Zhang, Puneet K Dokania, Oliver Wang, Alexei A Efros, Philip HS Torr, and Eli Shechtman. Interactive sketch & fill: Multiclass sketch-to-image translation. In ICCV, 2019. 3
  16. 16.David Ha and Douglas Eck. A neural representation of sketch drawings. In ICLR, 2018. 3
  17. 17.Bo Han, Jiangchao Yao, Gang Niu, Mingyuan Zhou, Ivor Tsang, Ya Zhang, and Masashi Sugiyama. Masking: A new perspective of noisy supervision. In NeurIPS, 2018. 3
  18. 18.Xiaoguang Han, Zhaoxuan Zhang, Dong Du, Mingdai Yang, Jingming Yu, Pan Pan, Xin Yang, Ligang Liu, Zixiang Xiong, and Shuguang Cui. Deep reinforcement learning of volume-guided progressive view inpainting for 3d point scene completion from a single depth image. In CVPR, 2019. 3
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 4
  20. 20.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. In ICLR, 2017. 5
  21. 21.Youngjoo Jo and Jongyoul Park. Sc-fegan: Face editing generative adversarial network with user’s sketch and color. In ICCV, 2019. 2
  22. 22.Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore. Reinforcement learning: A survey. JAIR, 1996. 3, 5, 6
  23. 23.Andrew Zisserman Karen Simonyan. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015. 3
  24. 24.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2014. 6
  25. 25.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching. In ECCV, 2018. 1
  26. 26.Xiaodan Liang, Lisa Lee, and Eric P Xing. Deep variation-structured reinforcement learning for visual relationship and attribute detection. In CVPR, 2017. 3
  27. 27.Fang Liu, Xiaoming Deng, Yu-Kun Lai, Yong-Jin Liu, Cuixia Ma, and Hongan Wang. Sketchgan: Joint sketch completion and recognition with generative adversarial network. In CVPR, 2019. 3
  28. 28.Li Liu, Fumin Shen, Yuming Shen, Xianglong Liu, and Ling Shao. Deep sketch hashing: Fast free-hand sketch-based image retrieval. In CVPR, 2017. 2
  29. 29.Runtao Liu, Qian Yu, and Stella X Yu. Unsupervised sketch to photo synthesis. In ECCV, 2020. 7, 8
  30. 30.Umar Riaz Muhammad, Yongxin Yang, Timothy M Hospedales, Tao Xiang, and Yi-Zhe Song. Goal-driven sequential data abstraction. In ICCV, 2019. 3
  31. 31.Umar Riaz Muhammad, Yongxin Yang, Yi-Zhe Song, Tao Xiang, and Timothy M Hospedales. Learning deep sketch abstraction. In CVPR, 2018. 3
  32. 32.Radford M Neal. Annealed importance sampling. Statistics and Computing, 2001. 5, 6
  33. 33.Kaiyue Pang, Ke Li, Yongxin Yang, Honggang Zhang, Timothy M Hospedales, Tao Xiang, and Yi-Zhe Song. Generalising fine-grained sketch-based image retrieval. In CVPR, 2019. 1, 2, 6, 7
  34. 34.Kaiyue Pang, Yi-Zhe Song, Tony Xiang, and Timothy M Hospedales. Cross-domain generative learning for fine-grained sketch-based image retrieval. In BMVC, 2017. 2
  35. 35.Kaiyue Pang, Yongxin Yang, Timothy M Hospedales, Tao Xiang, and Yi-Zhe Song. Solving mixed-modal jigsaw puzzle for fine-grained sketch-based image retrieval. In CVPR, 2020. 3, 6, 7
  36. 36.Leo Sampaio Ferraz Ribeiro, Tu Bui, John Collomosse, and Moacir Ponti. Sketchformer: Transformer-based representation for sketched structure. In CVPR, 2020. 1, 2
  37. 37.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 2015. 6
  38. 38.Aneeshan Sain, Ayan Kumar Bhunia, Vaishnav Potlapalli, Pinaki Nath Chowdhury, Tao Xiang, and Yi-Zhe Song. Sketch3t: Test-time training for zero-shot sbir. In CVPR, 2022. 1
  39. 39.Aneeshan Sain, Ayan Kumar Bhunia, Yongxin Yang, Tao Xiang, and Yi-Zhe Song. Cross-modal hierarchical modelling for fine-grained sketch based image retrieval. In BMVC, 2020. 2, 3, 6, 7
  40. 40.Aneeshan Sain, Ayan Kumar Bhunia, Yongxin Yang, Tao Xiang, and Yi-Zhe Song. Stylemeup: Towards style-agnostic sketch-based image retrieval. In CVPR, 2021. 2, 3, 7
  41. 41.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017. 5, 7, 8
  42. 42.Yuming Shen, Li Liu, Fumin Shen, and Ling Shao. Zero-shot sketch-image hashing. In CVPR, 2018. 1, 2
  43. 43.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 6
  44. 44.Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee. Learning from noisy labels with deep neural networks: A survey. arXiv preprint arXiv:2007.08199, 2021. 3
  45. 45.Jifei Song, Yi-Zhe Song, Tony Xiang, and Timothy M Hospedales. Fine-grained image retrieval: the text/sketch input dilemma. In BMVC, 2017. 2, 4, 7
  46. 46.Jifei Song, Qian Yu, Yi-Zhe Song, Tao Xiang, and Timothy M Hospedales. Deep spatial-semantic attention for fine-grained sketch-based image retrieval. In CVPR, 2017. 1, 2, 7
  47. 47.Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour. Policy gradient methods for reinforcement learning with function approximation. In NeurIPS, 2000. 5
  48. 48.Ryutaro Tanno, Ardavan Saeedi, Swami Sankaranarayanan, Daniel C Alexander, and Nathan Silberman. Learning from noisy labels by regularized estimation of annotator confusion. In CVPR, 2019. 3
  49. 49.Giorgos Tolias and Ondrej Chum. Asymmetric feature maps with application to sketch based retrieval. In CVPR, 2017. 2
  50. 50.Sheng-Yu Wang, David Bau, and Jun-Yan Zhu. Sketch your own gan. In ICCV, 2021. 2
  51. 51.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In CVPR, 2019. 3
  52. 52.Kilian Q Weinberger and Lawrence K Saul. Distance metric learning for large margin nearest neighbor classification. JMLR, 2009. 3
  53. 53.Peng Xu, Yongye Huang, Tongtong Yuan, Kaiyue Pang, Yi-Zhe Song, Tao Xiang, Timothy M Hospedales, Zhanyu Ma, and Jun Guo. Sketchmate: Deep hashing for million-scale human sketch retrieval. In CVPR, 2018. 2
  54. 54.Peng Xu, Chaitanya K Joshi, and Xavier Bresson. Multi-graph transformer for free-hand sketch recognition. IEEE T-NNLS, 2021. 3
  55. 55.Sasi Kiran Yelamarthi, Shiva Krishna Reddy, Ashish Mishra, and Anurag Mittal. A zero-shot framework for sketch based image retrieval. In ECCV, 2018. 2
  56. 56.Qian Yu, Feng Liu, Yi-Zhe Song, Tao Xiang, Timothy M Hospedales, and Chen-Change Loy. Sketch me that shoe. In CVPR, 2016. 1, 2, 3, 4, 6, 7
  57. 57.Qian Yu, Jifei Song, Yi-Zhe Song, Tao Xiang, and Timothy M Hospedales. Fine-grained instance-level sketch-based image retrieval. IJCV, 2021. 2
  58. 58.Qian Yu, Yongxin Yang, Feng Liu, Yi-Zhe Song, Tao Xiang, and Timothy M Hospedales. Sketch-a-net: A deep neural network that beats humans. IJCV, 2017. 2, 6
  59. 59.Jingyi Zhang, Fumin Shen, Li Liu, Fan Zhu, Mengyang Yu, Ling Shao, Heng Tao Shen, and Luc Van Gool. Generative domain-migration hashing for sketch-to-image retrieval. In ECCV, 2018. 1, 2
  60. 60.Tianyi Zhou, Shengjie Wang, and Jeff Bilmes. Robust curriculum learning: from clean label detection to noisy label self-correction. In ICLR, 2021. 3

Citation

MLA
Bhunia, A. K., et al. “Sketching Without Worrying: Noise-Tolerant Sketch-Based Image Retrieval”. arXiv, 2022, http://arxiv.org/abs/2203.14817v1.
APA
Bhunia, A. K., Koley, S., Khilji, A. F. U. R., Sain, A., Chowdhury, P. N., Xiang, T., & Song, Y.-Z. (2022). Sketching without Worrying: Noise-Tolerant Sketch-Based Image Retrieval. arXiv. http://arxiv.org/abs/2203.14817v1
Chicago
Bhunia, A. K., S. Koley, A. F. U. R. Khilji, et al. 2022. “Sketching Without Worrying: Noise-Tolerant Sketch-Based Image Retrieval”. arXiv. http://arxiv.org/abs/2203.14817v1.
Harvard
Bhunia, A.K. et al. (2022) “Sketching without Worrying: Noise-Tolerant Sketch-Based Image Retrieval”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.14817v1.
Vancouver
1. Bhunia AK, Koley S, Khilji AFUR, Sain A, Chowdhury PN, Xiang T, Song Y-Z (2022) Sketching without Worrying: Noise-Tolerant Sketch-Based Image Retrieval. arXiv

BibTeX

@article{bhunia2022sketching,
  title = {Sketching without Worrying: Noise-Tolerant Sketch-Based Image Retrieval},
  author = {Bhunia, Ayan Kumar and Koley, Subhadeep and Khilji, Abdullah Faiz Ur Rahman and Sain, Aneeshan and Chowdhury, Pinaki Nath and Xiang, Tao and Song, Yi-Zhe},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.14817v1},
  eprint = {2203.14817}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE