Retrieval Augmented Classification for Long-Tail Visual Recognition

Alexander LongWei YinThalaiyasingam AjanthanVu NguyenPulak PurkaitRavi GargAlan BlairChunhua ShenAnton van den Hengel

article2022CVPR142 citations

Proposes a dual-branch architecture that combines a standard image encoder with non-parametric multimodal retrieval to handle rare visual classes, significantly outperforming existing long-tail recognition baselines without requiring model fine-tuning.

Listen

Real-world visual recognition systems frequently struggle with skewed data distributions where a small number of frequent classes dominate the training set, while the majority of classes contain very few examples. Standard deep learning architectures typically store knowledge implicitly within their model weights, causing common classes to overshadow rare categories and leading to poor accuracy on infrequent classes. Addressing this performance gap is essential for deploying reliable visual recognition systems in practical environments where rare classes are common.

The article demonstrates that augmenting a standard vision backbone with an external retrieval module—a framework called Retrieval Augmented Classification (RAC)—significantly improves classification accuracy on heavily imbalanced image datasets without requiring expensive fine-tuning of large models.

To evaluate this approach, the authors constructed a two-branch architecture combining a standard image encoder with a retrieval branch that queries an external memory bank of pre-encoded images and text labels using approximate nearest-neighbor search. They evaluated RAC across standard long-tail benchmarks, including Places365-LT (62,500 scene images across 365 classes) and iNaturalist2018 (437,000 fine-grained wildlife images across 8,142 classes). The experiments compared RAC against prior state-of-the-art methods and isolated the specific contributions of the retrieval branch, encoder architectures, index size, and text encodings.

The findings show that RAC establishes a new state of the art on long-tail image recognition. First, RAC achieved overall classification accuracy of 47.17% on Places365-LT and 80.24% on iNaturalist2018 at standard resolutions, outperforming previous top models by 14.5% and 6.7% in relative accuracy gains. Second, the evaluation demonstrated an emergent division of labor: without explicit prompting, the retrieval module achieved high accuracy on rare categories, freeing the primary base encoder to focus on frequent classes. Third, nearest-neighbor searches over external memory banks exceeding 10 million samples introduced negligible query latency, with computational overhead confined to the secondary text encoder and increasing training runtime by 1.5 to 2 times. Fourth, increasing the number of unique classes in the external index produced larger performance gains than simply adding more images per existing class.

These results indicate that separating world knowledge into an external memory reduces the need to fine-tune massive neural networks, substantially lowering computational costs while improving accuracy on rare categories. Organizations can dynamically add or remove information from the retrieval index without retraining the primary model weights, simplifying system updates and maintenance.

Stakeholders deploying vision systems in long-tailed operational environments should consider adopting retrieval-augmented architectures rather than relying entirely on parameter adjustments or loss modifications. When implementing such pipelines, teams should prioritize pairing high-capacity vision backbones with fast approximate nearest-neighbor index structures. Future development should explore expanding retrieved metadata beyond basic text labels to richer text descriptions, such as captions or paragraphs, and testing the framework across broader multi-class and balanced domains.

The authors note limitations regarding dataset scope and text complexity, as the evaluation was restricted to two primary long-tail datasets and constrained by a 76-token input limit on the text encoder. Nonetheless, the high consistency of the empirical benchmarks supports strong confidence in the findings for long-tail image recognition tasks.

arXiv: 2202.11233
Cover for Retrieval Augmented Classification for Long-Tail Visual Recognition

Abstract

We introduce Retrieval Augmented Classification (RAC), a generic approach to augmenting standard image classification pipelines with an explicit retrieval module. RAC consists of a standard base image encoder fused with a parallel retrieval branch that queries a non-parametric external memory of pre-encoded images and associated text snippets. We apply RAC to the problem of long-tail classification and demonstrate a significant improvement over previous state-of-the-art on Places365-LT and iNaturalist-2018 (14.5% and 6.7% respectively), despite using only the training datasets themselves as the external information source. We demonstrate that RAC’s retrieval module, without prompting, learns a high level of accuracy on tail classes.This, in turn, frees the base encoder to focus on common classes, and improve its performance thereon. RAC represents an alternative approach to utilizing large, pretrained models without requiring fine-tuning, as well as a first step towards more effectively making use of external memory within common computer vision architectures.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. LACE Loss
  • 3.3. Retrieval Augmented Classification
  • 3.4. Retrieval Module
  • 4. Experiments
  • 4.1. Places365-LT
  • 4.2. iNaturalist-2018
  • 4.3. Ablation
  • 4.4. Retrieval
  • 4.5. Importance of the Text Encoder
  • 4.6. Effect of k
  • 4.7. Impact of Index Content
  • 4.8. Runtime Consideration
  • 5. Limitations
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Retrieval Augmented Classification (RAC) Architecture

    model/method

    Retrieval Augmented Classification (RAC) is an image classification architecture designed for long-tail visual recognition that pairs a parametric base image encoder with a non-parametric multi-modal external memory module.

    The framework consists of two parallel branches:

    1. Base Branch: A trainable image encoder B(⋅)B(\cdot) (such as a Vision Transformer ViT-B-16) that maps an input query image xqx^q to classification logits fbase(xq)∈RL\mathbf{f}^{\text{base}}(x^q) \in \mathbb{R}^L, where LL is the number of target classes.
    2. Retrieval Branch: A module containing a fixed, pre-trained image encoder E(⋅)E(\cdot) and an external memory index of images I={ij}j=1J\mathcal{I} = \{i_j\}_{j=1}^J with associated text snippets T={tj}j=1J\mathcal{T} = \{t_j\}_{j=1}^J (e.g., class names or descriptions). For a query xqx^q, latent features zq=E(xq)z^q = E(x^q) are extracted and used to query an approximate nearest neighbor index over stored image keys Z={zj=E(ij)}\mathcal{Z} = \{z_j = E(i_j)\}. The text snippets associated with the top-kk nearest neighbors are retrieved, tokenized, and passed into a trainable text encoder T(⋅)T(\cdot) followed by a linear classification head to produce retrieval logits fret(xq)∈RL\mathbf{f}^{\text{ret}}(x^q) \in \mathbb{R}^L.

    The outputs of both branches are combined via unit-norm normalization, linear addition, and rescaling before being trained jointly end-to-end under a common classification loss.

  2. Knowl 2 — Logit Normalization and Fusion Operation in RAC

    equation

    To combine the predictions of the base image branch and the retrieval branch without allowing either branch to dominate or alter logit magnitudes relative to standard baselines, RAC computes the combined class logit vector f(x)∈RL\mathbf{f}(x) \in \mathbb{R}^L as:

    f(x)=L2(fret(x)∥fret(x)∥2+fbase(x)∥fbase(x)∥2)\mathbf{f}(x) = \frac{L}{2} \left( \frac{\mathbf{f}^{\text{ret}}(x)}{\|\mathbf{f}^{\text{ret}}(x)\|_2} + \frac{\mathbf{f}^{\text{base}}(x)}{\|\mathbf{f}^{\text{base}}(x)\|_2} \right)

    where:

    • xx is the input image.
    • LL is the total number of target classes.
    • fbase(x)∈RL\mathbf{f}^{\text{base}}(x) \in \mathbb{R}^L is the unnormalized logit vector from the base image encoder B(⋅)B(\cdot).
    • fret(x)∈RL\mathbf{f}^{\text{ret}}(x) \in \mathbb{R}^L is the unnormalized logit vector from the retrieval text encoder T(⋅)T(\cdot).
    • ∥⋅∥2\|\cdot\|_2 denotes the Euclidean (L2L_2) norm.
    • The factor L2\frac{L}{2} rescales the combined unit-norm vectors to match expected logit variance under Xavier initialization (which depends on LL).

    Weighting between the branches is performed implicitly via the learned sharpness of the respective logit distributions.

  3. Knowl 3 — Retrieval Module Index Construction and Query Procedure

    model/method

    The RAC retrieval module indexes external visual and textual knowledge through the following construction and query pipeline:

    1. Index Initialization: An external image set I={ij}j=1J\mathcal{I} = \{i_j\}_{j=1}^J is embedded using a frozen pre-trained image encoder E(⋅)E(\cdot) to produce key representations Z={zj=E(ij)}j=1J\mathcal{Z} = \{z_j = E(i_j)\}_{j=1}^J. These vectors are stored in a Hierarchical Navigable Small World (HNSW) approximate kk-NN index (via FAISS) using cosine similarity as the metric and construction hyperparameter M=32M = 32 bidirectional links per node.
    2. Feature Extraction and Search: For an input batch of query images xqx^q, latent query representations zq=E(xq)z^q = E(x^q) are extracted. The HNSW index is searched to obtain the indices of the kk closest keys in Z\mathcal{Z}.
    3. Training Self-Match Filtering: When the training set is included in the index, the top-1 retrieved match is discarded during training to prevent the text encoder from overfitting to identical self-retrievals.
    4. Text Aggregation and Encoding: The text snippets {tj}\{t_j\} corresponding to the top-kk indices are concatenated, truncated at 76 tokens, and zero-padded. All kk snippets for a query image are processed in a single batched forward pass through a BERT-like text transformer T(⋅)T(\cdot) (from CLIP), followed by a linear projection layer to output class logits fret(xq)\mathbf{f}^{\text{ret}}(x^q).
  4. Knowl 4 — Spontaneous Specialization of Base and Retrieval Branches on Head and Tail Classes

    empirical result

    When RAC is trained under a single unified Logit Adjusted Cross-Entropy (LACE) loss without explicit loss routing or class-specific constraints:

    • The retrieval branch fret(x)\mathbf{f}^{\text{ret}}(x) spontaneously learns to achieve high classification accuracy on tail (few-shot) classes while showing low accuracy on head classes.
    • The base image branch fbase(x)\mathbf{f}^{\text{base}}(x) focuses on many-shot (head) classes, achieving higher performance on frequent classes.
    • The non-parametric memory of the retrieval module absorbs the burden of memorizing sparse tail classes, allowing the parametric base network to allocate its representational capacity to complex head classes.
    • When combined, the joint model f(x)\mathbf{f}(x) attains balanced, high accuracy across all class frequencies (many-, medium-, and few-shot).
  5. Knowl 5 — Long-Tail Benchmark Evaluation on Places365-LT and iNaturalist-2018

    data/table

    RAC significantly outperforms prior state-of-the-art methods across standard long-tail recognition benchmarks. Classes are split into Many (>100 samples), Medium (20 to 100 samples), and Few (<20 samples) shots.

    Method Backbone Many Med Few All
    iNaturalist-2018 (224×224224 \times 224)
    OLTR ResNet-50 59.0 64.1 64.9 63.9
    Dec. LWS ResNet-50 65.0 66.3 65.5 65.9
    LADE ResNet-50 - - - 70.0
    ALA ResNet-50 71.3 70.8 70.4 70.7
    LACE ResNet-50 - - - 71.9
    RIDE ResNet-50 70.9 72.4 73.1 72.6
    TADE ResNet-50 74.4 72.5 73.1 72.9
    DisAlign ResNet-152 - - - 74.1
    PaCo ResNet-152 75.0 75.5 74.7 75.2
    RAC (ours) ViT-B-16 75.92 80.47 81.07 80.24
    iNaturalist-2018 (384×384384 \times 384)
    Grafit RegNetY - - - 81.2
    RAC (ours) ViT-B-16 82.91 85.71 86.06 85.56
    Places365-LT (256×256256 \times 256)
    Focal Loss ResNet-152 41.1 34.8 22.4 34.6
    Range Loss ResNet-152 41.1 35.4 23.2 35.1
    OLTR ResNet-152 44.7 37.0 25.3 35.9
    Dec. LWS ResNet-152 40.6 39.1 28.6 37.6
    LADE ResNet-152 42.8 39.0 31.2 38.8
    DisAlign ResNet-152 40.4 42.4 30.1 39.3
    ALA ResNet-152 43.9 40.1 32.9 40.1
    TADE ResNet-152 43.1 42.4 33.2 40.9
    PaCo ResNet-152 36.1 47.9 35.3 41.2
    RAC (ours) ViT-B-16 48.69 48.31 41.76 47.17

    On Places365-LT, RAC achieves 47.17% overall top-1 accuracy (a 5.97% absolute increase over PaCo and a 14.5% relative error reduction over prior state-of-the-art). On iNaturalist-2018, RAC improves overall accuracy from 75.2% to 80.24% at 224×224224 \times 224 and to 85.56% at 384×384384 \times 384.

  6. Knowl 6 — Ablation against Controlled Baselines on Identical ViT-B-16 Backbones

    data/table

    To isolate the benefit of RAC's architecture from modernized Vision Transformer backbones and training recipes, RAC is compared against class-balanced softmax cross-entropy (BalCE) and logit-adjusted cross-entropy (LACE, designated as 'Base only') using the same ViT-B-16 backbone:

    Dataset Method / Branch Many Med Few All
    Places365-LT CE (ResNet-50) - - - 32.14
    Balanced CE (ResNet-50) - - - 38.31
    CE (ViT-B-16) 50.81 33.83 19.51 37.16
    Balanced CE (ViT-B-16) 49.03 45.72 29.05 43.67
    Retrieval only 43.50 41.99 26.83 39.58
    Base only (LACE) 44.57 45.06 40.77 44.05
    RAC (Full) 48.69 48.31 41.76 47.17
    iNaturalist 2018 CE (ResNet-50) - - - 61.70
    Balanced CE (ResNet-50) - - - 69.80
    CE (ViT-B-16) 81.53 76.62 69.82 74.44
    Balanced CE (ViT-B-16) 72.39 76.06 73.05 74.49
    Retrieval only 50.10 52.77 52.45 52.37
    Base only (LACE) 74.41 78.95 78.55 78.32
    RAC (Full) 75.92 80.48 81.07 80.24

    Against the strongest ViT-B-16 baseline ('Base only' trained under LACE with label smoothing), RAC yields a +3.12% overall improvement on Places365-LT (from 44.05% to 47.17%) and a +1.92% overall improvement on iNaturalist-2018 (from 78.32% to 80.24%).

  7. Knowl 7 — Impact of Number of Retrieved Neighbors $k$ and Text Encoder Architecture

    empirical result

    Analysis of the retrieval branch's hyperparameters and architecture demonstrates:

    1. Effect of kk: Increasing the number of retrieved text snippets kk in the kk-NN search on Places365-LT monotonically improves top-1 accuracy (from ∼29%\sim 29\% at k=3k=3 up to ∼40%\sim 40\% at k=30k=30) before saturating near the 76-token truncation limit. Even when k=30k=30 exceeds the sample count of few-shot classes (which forces incorrect classes into the retrieved set), the text encoder T(⋅)T(\cdot) learns to filter out irrelevant labels and maintain high accuracy.
    2. Text Encoder Capacity: Comparing the 63M parameter CLIP BERT-like text encoder against Bag-of-Words (BoW) GloVe embeddings (300-d) and BoW cached random word embeddings (300-d with uniform random vectors per word):
      • The CLIP text encoder provides the best performance, particularly boosting few-shot accuracy on Places365-LT where class labels correspond to natural language scene descriptions.
      • Simpler encoders and even random embeddings perform moderately well (retaining significant retrieval utility), demonstrating that exact semantic language comprehension is helpful but not strictly required when querying simple class labels.
  8. Knowl 8 — Impact of External Index Composition: Class Diversity vs. Instance Density

    empirical result

    Evaluating retrieval branch accuracy on Places365-LT using subsets of ImageNet21k in the index reveals distinct scaling behaviors:

    • Total Index Size: Naively scaling the number of indexed images increases accuracy up to a point, after which returns diminish due to redundant visual representations.
    • Samples per Class vs. Total Classes: Holding total index size constant, increasing the number of distinct semantic classes (Ni=10N_i = 10 constant per class while varying total classes from 10110^1 to 10410^4) yields significantly larger accuracy improvements than increasing the number of samples per class (L=10L=10 classes constant while varying samples from 10110^1 to 10310^3).

    Adding diverse semantic categories increases the probability of retrieving relevant textual information, whereas adding more visual instances of existing classes provides diminishing informational gain to the text encoder.

  9. Knowl 9 — Training Computational Scalability and Index Query Latency

    data/table

    RAC scales to large external memory indexes by utilizing HNSW indexing and per-node index replication:

    Index Dataset Index Size Text Encoder Speed (s/epoch)
    None (Base only) None None 23.6
    Places365-LT 184K Random (Rand.) 28.3
    Places365-LT 184K CLIP 44.3
    Places + ImageNet21k 11.2M CLIP 47.0

    Key observations:

    • Querying an index containing over 11.2 million samples adds only 2.7 seconds per epoch over an index of 184K samples (44.3 s to 47.0 s), confirming logarithmic search complexity with HNSW.
    • The primary computational overhead arises from backpropagation through the 63M-parameter text transformer T(⋅)T(\cdot) (increasing training time by 1.5×−2.0×1.5\times - 2.0\times), not from the nearest neighbor lookups.
    • Multi-GPU and multi-node training bottlenecks are avoided by keeping full copies of the HNSW index in host CPU memory on each node.
  10. Knowl 10 — Limitations of Retrieval Augmented Classification

    limitation

    The RAC framework has several identified limitations:

    1. Pretraining Dependency: RAC requires a pre-trained image encoder E(⋅)E(\cdot) to construct the index. It cannot be fairly evaluated in standard from-scratch benchmarks (e.g., CIFAR-LT or ImageNet-LT trained without pretraining).
    2. Token Truncation Limit: The text encoder enforces a 76-token truncation length, limiting the amount and complexity of external text that can be returned (e.g., restricting retrieval to simple labels rather than detailed multi-sentence captions or paragraphs).
    3. Restricted Domain Evaluation: The approach was analyzed primarily on long-tailed classification distributions (Places365-LT and iNaturalist-2018); its characteristics and performance on balanced classification datasets remain unexamined.

Coverage note — None was omitted; all key contributions including framework architecture, logit normalization equation, index construction/query algorithms, head/tail specialization analysis, empirical benchmark results, ablation comparisons, index content experiments, and scalability analyses are captured.

References

  1. 1.Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma. Learning imbalanced datasets with label-distribution-aware margin loss. arXiv: Comp. Res. Repository, 2019. 3
  2. 2.Claire Cardie and Nicholas Howe. Improving minority class prediction using case-specific feature weights. In Proc. Int. Conf. Mach. Learn., 1997. 1
  3. 3.Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer. Smote: synthetic minority over-sampling technique. J. Artificial Intelligence Research, 16:321–357, 2002. 1, 2
  4. 4.Hila Chefer, Shir Gur, and Lior Wolf. Transformer interpretability beyond attention visualization. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 782–791, 2021. 1
  5. 5.Peng Chu, Xiao Bian, Shaopeng Liu, and Haibin Ling. Feature space augmentation for long-tailed data. In Proc. Eur. Conf. Comp. Vis., pages 694–710, 2020. 2
  6. 6.Jiequan Cui, Zhisheng Zhong, Shu Liu, Bei Yu, and Jiaya Jia. Parametric contrastive learning. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 715–724, 2021. 1, 2, 5
  7. 7.Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song, and Serge Belongie. Class-balanced loss based on effective number of samples. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 9268–9277, 2019. 6
  8. 8.Nicola De Cao, Wilker Aziz, and Ivan Titov. Editing factual knowledge in language models. arXiv: Comp. Res. Repository, 2021. 1
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 248–255. Ieee, 2009. 5
  10. 10.Zongyong Deng, Hao Liu, Yaoxing Wang, Chenyang Wang, Zekuan Yu, and Xuehong Sun. PML: Progressive margin loss for long-tailed age classification. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 10503–10512, 2021. 1
  11. 11.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv: Comp. Res. Repository, 2020. 1, 4, 5
  12. 12.Chris Drumnond. Class imbalance and cost sensitivity: Why undersampling beats oversampling. In ICML-KDD Workshop: Learning from Imbalanced Datasets, volume 3, 2003. 1, 2
  13. 13.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proc. Int. Conf. Artificial Intelligence and Statistics, pages 249–256, 2010. 4
  14. 14.Dong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh, and Anton van den Hengel. Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 1705–1714, 2019. 3
  15. 15.Alex Graves, Greg Wayne, and Ivo Danihelka. Neural turing machines. arXiv: Comp. Res. Repository, 2014. 3
  16. 16.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. In Proc. Int. Conf. Mach. Learn., pages 1321–1330, 2017. 3
  17. 17.Hao Guo and Song Wang. Long-tailed multi-label visual recognition by collaborative training on uniform and re-balanced samplings. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 15089–15098, 2021. 1
  18. 18.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. Realm: Retrieval-augmented language model pre-training. arXiv: Comp. Res. Repository, 2020. 3
  19. 19.Hui Han, Wen-Yuan Wang, and Bing-Huan Mao. Borderline-SMOTE: A new over-sampling method in imbalanced data sets learning. In Proc. Int. Conf. Intelligent Computing, pages 878–887, 2005. 1, 2
  20. 20.Haibo He and Edwardo A. Garcia. Learning from imbalanced data. IEEE Trans. Knowledge & Data Engineering, 21(9):1263–1284, 2009. 1
  21. 21.Youngkyu Hong, Seungju Han, Kwanghee Choi, Seokjun Seo, Beomsu Kim, and Buru Chang. Disentangling label distribution for long-tailed visual recognition. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 6626–6636, June 2021. 1, 2, 5
  22. 22.Chen Huang, Yining Li, Chen Change Loy, and Xiaoou Tang. Learning deep representation for imbalanced classification. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 5375–5384, 2016. 1, 2
  23. 23.Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with GPUs. arXiv: Comp. Res. Repository, 2017. 5
  24. 24.Bingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan, Albert Gordo, Jiashi Feng, and Yannis Kalantidis. Decoupling representation and classifier for long-tailed recognition. arXiv: Comp. Res. Repository, 2019. 1, 5
  25. 25.Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense passage retrieval for open-domain question answering. arXiv: Comp. Res. Repository, 2020. 3
  26. 26.Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. Transformers in vision: A survey. arXiv: Comp. Res. Repository, 2021. 1
  27. 27.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. Generalization through memorization: Nearest neighbor language models. In Proc. Int. Conf. Learn. Representations, 2020. 3
  28. 28.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. In Proc. Eur. Conf. Comp. Vis., pages 491–507. Springer, 2020. 5
  29. 29.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. Latent retrieval for weakly supervised open domain question answering. arXiv: Comp. Res. Repository, 2019. 3
  30. 30.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. arXiv: Comp. Res. Repository, 2020. 3
  31. 31.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In Proc. IEEE Int. Conf. Comp. Vis., pages 2980–2988, 2017. 5
  32. 32.Jialun Liu, Yifan Sun, Chuchu Han, Zhaopeng Dou, and Wenhui Li. Deep representation learning on long-tailed data: A learnable embedding augmentation perspective. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 2970–2979, 2020. 2
  33. 33.Ziwei Liu, Zhongqi Miao, Xiaohang Zhan, Jiayun Wang, Boqing Gong, and Stella X Yu. Large-scale long-tailed recognition in an open world. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 2537–2546, 2019. 3, 5, 6
  34. 34.Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens Van Der Maaten. Exploring the limits of weakly supervised pretraining. In Proc. Eur. Conf. Comp. Vis., pages 181–196, 2018. 1
  35. 35.Yu A Malkov and Dmitry A Yashunin. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Trans. Pattern Anal. Mach. Intell., 42(4):824–836, 2018. 5
  36. 36.Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain, Andreas Veit, and Sanjiv Kumar. Long-tail learning via logit adjustment. arXiv: Comp. Res. Repository, 2020. 1, 2, 3, 5, 6
  37. 37.Rafael Müller, Simon Kornblith, and Geoffrey Hinton. When does label smoothing help? arXiv: Comp. Res. Repository, 2019. 3
  38. 38.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv: Comp. Res. Repository, 2018. 2
  39. 39.Yoon-Joo Park and Alexander Tuzhilin. The long tail of recommender systems and how to leverage it. In Proc. ACM Conf. Recommender Systems, pages 11–18, 2008. 1
  40. 40.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. Glove: Global vectors for word representation. In Proc. Conf. Empirical Methods in Natural Language Process., pages 1532–1543, 2014. 7
  41. 41.Joaquin Quiñonero-Candela, Masashi Sugiyama, Neil D. Lawrence, and Anton Schwaighofer. Dataset shift in machine learning. Mit Press, 2009. 1
  42. 42.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. arXiv: Comp. Res. Repository, 2021. 1, 4
  43. 43.Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. How to train your vit? Data, augmentation, and regularization in vision transformers. arXiv: Comp. Res. Repository, 2021. 4
  44. 44.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In Proc. IEEE Int. Conf. Comp. Vis., pages 843–852, 2017. 1
  45. 45.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 2818–2826, 2016. 3
  46. 46.Hugo Touvron, Alexandre Sablayrolles, Matthijs Douze, Matthieu Cord, and Hervé Jégou. Grafit: Learning fine-grained image representations with coarse labels. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 874–884, 2021. 3, 6
  47. 47.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 8769–8778, 2018. 6
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proc. Advances in Neural Inf. Process. Syst., pages 5998–6008, 2017. 1
  49. 49.Pat Verga, Haitian Sun, Livio Baldini Soares, and William Cohen. Adaptable and interpretable neural MemoryOver symbolic knowledge. In Proc. Conf. North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3678–3691, June 2021. 3
  50. 50.Xudong Wang, Long Lian, Zhongqi Miao, Ziwei Liu, and Stella X Yu. Long-tailed recognition by routing diverse distribution-aware experts. arXiv: Comp. Res. Repository, 2020. 3, 5
  51. 51.Ross Wightman, Hugo Touvron, and Hervé Jégou. Resnet strikes back: An improved training procedure in timm. arXiv: Comp. Res. Repository, 2021. 6
  52. 52.Liuyu Xiang, Guiguang Ding, and Jungong Han. Learning from multiple experts: Self-paced knowledge distillation for long-tailed classification. In Proc. Eur. Conf. Comp. Vis., pages 247–263. Springer, 2020. 1
  53. 53.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. arXiv: Comp. Res. Repository, 2021. 1
  54. 54.Songyang Zhang, Zeming Li, Shipeng Yan, Xuming He, and Jian Sun. Distribution alignment: A unified framework for long-tail visual recognition. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 2361–2370, 2021. 2, 5
  55. 55.Xiao Zhang, Zhiyuan Fang, Yandong Wen, Zhifeng Li, and Yu Qiao. Range loss for deep face recognition with long-tailed training data. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 5409–5418, 2017. 5
  56. 56.Yifan Zhang, Bryan Hooi, Lanqing Hong, and Jiashi Feng. Test-agnostic long-tailed recognition by test-time aggregating diverse experts with self-supervision. arXiv: Comp. Res. Repository, 2021. 2, 5, 6
  57. 57.Yan Zhao, Weicong Chen, Xu Tan, Kai Huang, Jin Xu, Changhu Wang, and Jihong Zhu. Adaptive logit adjustment loss for long-tailed visual recognition. arXiv: Comp. Res. Repository, 2021. 1, 5
  58. 58.Yan Zhao, Weicong Chen, Xu Tan, Kai Huang, Jin Xu, Changhu Wang, and Jihong Zhu. Improving long-tailed classification from instance level. arXiv: Comp. Res. Repository, 2021. 6
  59. 59.Boyan Zhou, Quan Cui, Xiu-Shen Wei, and Zhao-Min Chen. Bbn: Bilateral-branch network with cumulative learning for long-tailed visual recognition. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 9719–9728, 2020. 1
  60. 60.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Trans. Pattern Anal. Mach. Intell., 40(6):1452–1464, 2017. 6

Citation

MLA
Long, A., et al. “Retrieval Augmented Classification for Long-Tail Visual Recognition”. arXiv, 2022, http://arxiv.org/abs/2202.11233v1.
APA
Long, A., Yin, W., Ajanthan, T., Nguyen, V., Purkait, P., Garg, R., Blair, A., Shen, C., & Hengel, A. van . den . (2022). Retrieval Augmented Classification for Long-Tail Visual Recognition. arXiv. http://arxiv.org/abs/2202.11233v1
Chicago
Long, A., W. Yin, T. Ajanthan, et al. 2022. “Retrieval Augmented Classification for Long-Tail Visual Recognition”. arXiv. http://arxiv.org/abs/2202.11233v1.
Harvard
Long, A. et al. (2022) “Retrieval Augmented Classification for Long-Tail Visual Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2202.11233v1.
Vancouver
1. Long A, Yin W, Ajanthan T, Nguyen V, Purkait P, Garg R, Blair A, Shen C, Hengel A van den (2022) Retrieval Augmented Classification for Long-Tail Visual Recognition. arXiv

BibTeX

@article{long2022retrieval,
  title = {Retrieval Augmented Classification for Long-Tail Visual Recognition},
  author = {Long, Alexander and Yin, Wei and Ajanthan, Thalaiyasingam and Nguyen, Vu and Purkait, Pulak and Garg, Ravi and Blair, Alan and Shen, Chunhua and Hengel, Anton van den},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2202.11233v1},
  eprint = {2202.11233}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE