Exploring the Limits of Weakly Supervised Pretraining

Dhruv MahajanRoss GirshickVignesh RamanathanKaiming HeManohar PaluriYixuan LiAshwin BharambeLaurens van der Maaten

article2018ECCV1,485 citations

Demonstrates that pretraining convolutional networks on billions of weakly supervised hashtagged social media images sets new state-of-the-art benchmarks across standard image classification and object detection tasks, including an 85.4% top-1 accuracy on ImageNet.

Listen

State-of-the-art computer vision systems rely heavily on pretraining deep neural networks on manually labeled datasets before adapting them to specific tasks. For nearly a decade, standard pretraining has depended on ImageNet, a benchmark containing roughly one million curated images. However, manually labeling datasets is expensive, time-consuming, and difficult to scale, leaving open the question of how models behave when pretrained on datasets that are orders of magnitude larger. Social media images provide an abundant, continuously growing data source, but they rely on user-generated hashtags that are uncurated, highly biased, and noisy.

The article evaluates whether training convolutional networks on billions of public social media images with noisy hashtag annotations can serve as an effective pretraining strategy for standard visual recognition tasks. It demonstrates the transferability of these weakly supervised models to standard image classification and object detection benchmarks, examining scaling behavior, label noise, and model capacity.

To conduct this evaluation, the authors trained standard deep residual networks (ResNeXt architectures) on up to 3.5 billion public Instagram images tagged with vocabularies ranging from 1,500 to 17,000 canonical hashtags. Training was distributed across hundreds of graphics processing units using synchronous stochastic gradient descent with large minibatches of approximately 8,000 images. The pretrained models were then evaluated on downstream benchmarks—including standard ImageNet classification, fine-grained bird recognition (CUB2011), scene recognition (Places365), and common object detection (COCO)—via full network fine-tuning or by training a simple linear classifier on fixed representations. The authors also applied rigorous image deduplication to ensure test sets were not contaminated by overlapping training images.

The findings establish that large-scale hashtag pretraining yields exceptional performance across multiple tasks. First, the approach achieved a new state-of-the-art single-crop top-1 accuracy of 85.4% (and 97.6% top-5) on the standard ImageNet-1k benchmark, outperforming the previous state of the art by 2.7 percentage points and surpassing standard ImageNet-only training by 5.8 percentage points. Second, the learned features are so effective that training only a simple linear classifier on fixed features achieved 83.6% top-1 accuracy on ImageNet-1k, nearly matching fully fine-tuned models. Third, model accuracy scales log-linearly with data volume, though gains become limited by model capacity, indicating that existing architectures underfit when trained on billions of images. Fourth, hashtag training showed remarkable resilience to noise: artificially corrupting 10% of labels reduced final accuracy by less than 1 percentage point, and 25% noise reduced accuracy by only about 2 percentage points. Fifth, data resampling techniques (such as square-root or uniform sampling) that rebalance the heavy-tailed hashtag distribution improved downstream classification accuracy by 5 to 6 percentage points compared to natural distribution sampling.

These results demonstrate that massive, uncurated web data can bypass the expensive requirement for manual data annotation in vision systems, dramatically lowering data collection costs while improving model performance. The findings also reveal critical nuances: while hashtag pretraining substantially enhances visual classification, its benefits for spatial localization tasks (such as keypoint detection and bounding-box precision) are mixed. Additionally, standard fine-tuning recipes developed for ImageNet models do not work out of the box; models pretrained at this scale require significantly lower fine-tuning learning rates (about 4 to 10 times lower).

Organizations developing computer vision applications should consider adopting large-scale weakly supervised pretraining to improve baseline performance. When deploying this approach, practitioners should carefully engineer the pretraining hashtag vocabulary to align with target tasks, employ data resampling to handle imbalanced label frequencies, and expand model capacities (such as exploring wider architectures or mixtures of experts) to avoid underfitting. For downstream transfer, teams must re-tune fine-tuning schedules with lower learning rates.

Confidence in these findings is high due to consistent trends observed across multiple standard benchmarks and extensive ablation experiments. However, key limitations remain. Because the underlying repository of billions of public social media images cannot be redistributed en masse, direct external replication of the exact pretraining dataset is constrained. Furthermore, pretraining on hashtag classification may not inherently optimize features for fine-grained spatial localization, highlighting the need for future research into pretraining objectives tailored for detection and dense prediction tasks.

arXiv: 1805.00932
Cover for Exploring the Limits of Weakly Supervised Pretraining

Abstract

State-of-the-art visual perception models for a wide range of tasks rely on supervised pretraining. ImageNet classification is the de facto pretraining task for these models. Yet, ImageNet is now nearly ten years old and is by modern standards "small". Even so, relatively little is known about the behavior of pretraining with datasets that are multiple orders of magnitude larger. The reasons are obvious: such datasets are difficult to collect and annotate. In this paper, we present a unique study of transfer learning with large convolutional networks trained to predict hashtags on billions of social media images. Our experiments demonstrate that training for large-scale hashtag prediction leads to excellent results. We show improvements on several image classification and object detection tasks, and report the highest ImageNet-1k single-crop, top-1 accuracy to date: 85.4% (97.6% top-5). We also perform extensive experiments that provide novel empirical data on the relationship between large-scale pretraining and transfer learning performance.

Table of Contents

  • 1 Introduction
  • 2 Scaling up Supervised Pretraining
  • 2.1 Instagram Datasets
  • 2.2 ImageNet Datasets
  • 2.3 Models
  • 2.4 Pretraining Details
  • 3 Experiments
  • 3.1 Image Classification Experiments
  • 3.2 Object Detection
  • 4 Related Work
  • 5 Discussion
  • References
  • 0.A Supplemental Material
  • 0.A.1 Hashtag Selection
  • 0.A.2 Image Deduplication
  • 0.A.3 Training Details
  • 0.A.4 Data Resampling
  • 0.A.5 Comparison with the State of the Art on ImageNet-1k

Knowls

  1. Knowl 1 — State-of-the-Art ImageNet-1k Classification Performance from Large-Scale Hashtag Pretraining

    data/table

    Pretraining high-capacity convolutional neural networks (ResNeXt-101) on hundreds of millions to billions of public Instagram images annotated with social media hashtags, followed by supervised finetuning on the ImageNet-1k training set, sets a new state of the art in single-crop top-1 and top-5 accuracy on the ImageNet-1k validation set (evaluated on 224×224224 \times 224 center crops).

    Model Image size Parameters Mult-adds Top-1 Acc. (%) Top-5 Acc. (%)
    Inception V2 224 11.2M 1.94B 74.8 92.2
    NASNet-A (5 @ 1538) 299 10.9M 2.35B 78.6 94.2
    Inception V3 299 23.8M 5.72B 78.0 93.9
    Xception 299 22.8M 8.38B 79.0 94.5
    Inception ResNet V2 299 55.8M 13.2B 80.4 95.3
    NASNet-A (7 @ 1920) 299 22.6M 4.93B 80.8 95.3
    ResNeXt-101 64×464 \times 4 320 83.6M 31.5B 80.9 95.6
    PolyNet 331 92M 34.7B 81.3 95.8
    DPN-131 320 79.5M 32.0B 81.5 95.8
    SENet 320 145.8M 42.3B 82.7 96.2
    NASNet-A (6 @ 4032) 331 88.9M 23.8B 82.7 96.2
    IG-3.5B-17k ResNeXt-101 32×16d32 \times 16\text{d} 224 194M 36B 84.2 97.2
    IG-940M-1.5k ResNeXt-101 32×32d32 \times 32\text{d} 224 466M 87B 85.1 97.5
    IG-940M-1.5k ResNeXt-101 32×48d32 \times 48\text{d} 224 829M 153B 85.4 97.6

    The largest model, ResNeXt-101 32×48d32 \times 48\text{d} (829M parameters, 153B FLOPs) pretrained on 940 million Instagram images with 1.5k ImageNet-aligned hashtags, achieves 85.4% top-1 accuracy (97.6% top-5), outperforming the prior state of the art by 2.7% absolute top-1 accuracy. When evaluating rectified top-1 accuracy (where predictions verified as correct by at least 4 out of 5 human annotators are accepted), the IG-3.5B-17k pretrained model achieves 90.4% compared to 87.5% for an ImageNet-only trained baseline.

  2. Knowl 2 — Log-Linear Performance Scaling and Capacity Bottlenecks in Billion-Scale Pretraining

    empirical result

    When varying the scale of weakly supervised pretraining data across three orders of magnitude (from 3.5 million to 3.5 billion images), target task classification accuracy follows a log-linear relationship: multiplying the volume of pretraining data by a constant factor xx yields an approximately constant additive gain yy in target classification accuracy.

    However, this scaling behavior is governed by two critical constraints:

    1. Model Capacity Dependency: Higher-capacity network architectures display substantially steeper log-linear scaling slopes. ResNeXt-101 32×16d32 \times 16\text{d} gains more performance per order of magnitude of pretraining data than 32×8d32 \times 8\text{d} and 32×4d32 \times 4\text{d}, indicating that standard convolutional networks underfit when exposed to billions of training images.
    2. Breakdowns of Log-Linear Scaling: The log-linear scaling regime breaks down due to two distinct factors: (a) ceiling effects on bounded benchmarks (e.g., ImageNet-1k and CUB2011), and (b) an observed sub-log-linear deviation in the 1B to 3.5B image regime on large-vocabulary benchmarks (ImageNet-5k and ImageNet-9k) where saturation is absent.

    Furthermore, while training directly from scratch on ImageNet-1k saturates at approximately 79.6% top-1 accuracy regardless of increasing model capacity beyond ResNeXt-101 32×16d32 \times 16\text{d}, Instagram-pretrained models scale from 84.2% (32×16d32 \times 16\text{d}) to 85.1% (32×32d32 \times 32\text{d}) and 85.4% (32×48d32 \times 48\text{d}). This indicates that transfer learning from billions of weakly supervised images is bottlenecked by model capacity rather than data volume.

  3. Knowl 3 — Softmax Cross-Entropy Objective for Multi-Hashtag Weak Supervision

    model/method

    Although social media images typically contain multiple user-assigned hashtags (averaging approximately 2 hashtags per image), the pretraining objective is formulated as a multi-class softmax cross-entropy loss over the full hashtag vocabulary rather than independent per-hashtag binary logistic (sigmoid) losses.

    For an image associated with k≥1k \ge 1 ground-truth hashtags from a vocabulary of CC total classes, the target probability vector y∈RC\mathbf{y} \in \mathbb{R}^C distributes probability mass uniformly over the present tags:

    yc={1kif hashtag c is assigned to the image0otherwisey_c = \begin{cases} \frac{1}{k} & \text{if hashtag } c \text{ is assigned to the image} \\ 0 & \text{otherwise} \end{cases}

    Given predicted unnormalized logits z∈RC\mathbf{z} \in \mathbb{R}^C, the network computes class probabilities via the softmax function:

    pc=exp⁡(zc)∑j=1Cexp⁡(zj)p_c = \frac{\exp(z_c)}{\sum_{j=1}^C \exp(z_j)}

    The training loss is the cross-entropy between p\mathbf{p} and y\mathbf{y}:

    L=−∑c=1Cyclog⁡pc=−1k∑c∈tagslog⁡pc\mathcal{L} = - \sum_{c=1}^C y_c \log p_c = - \frac{1}{k} \sum_{c \in \text{tags}} \log p_c

    Empirically, using per-hashtag sigmoid activations with binary logistic loss yields substantially worse transfer learning performance compared to this normalized softmax cross-entropy formulation.

  4. Knowl 4 — Label-Space Alignment and the Impact of Hashtag Vocabulary Selection

    empirical result

    The effectiveness of weakly supervised pretraining depends heavily on the alignment between the pretraining label space and the target evaluation domain (referred to as hashtag engineering or label-space engineering):

    • Target-Task Alignment: On ImageNet-1k, pretraining on an aligned vocabulary of 1,500 hashtags matching ImageNet-1k synsets achieves higher transfer accuracy (84.2% top-1 on 1B images) than pretraining on broader vocabularies of 8,500 hashtags (83.4%) or 17,000 hashtags (83.6%).
    • Generalization to Diverse Label Spaces: Conversely, when transferring to broad or fine-grained recognition tasks with larger visual variety—such as ImageNet-9k, CUB-2011 bird classification (89.2% vs. 87.5%), and Places365 scene recognition (58.0% vs. 56.2%)—the broader 17,000 hashtag vocabulary significantly outperforms the 1,500 vocabulary. On ImageNet-9k, the 17k hashtag model outperforms the 1.5k model by approximately 7% absolute accuracy.
    • Concreteness Correlation: Visual predictability of hashtags strongly correlates with noun concreteness. Matching 17,000 hashtags against 40,000 noun concreteness values reveals a Pearson correlation of ρ=0.43\rho = 0.43 between the concreteness of a word and the model's ability to predict the corresponding hashtag (measured by AUC on balanced validation sets). Visually concrete concepts (e.g., #eiffeltower) are substantially easier to learn than abstract concepts (e.g., #party).
  5. Knowl 5 — Zipfian Resampling Strategy for Weakly Supervised Hashtag Pretraining

    algorithm

    Hashtags on social media follow a heavy-tailed Zipfian distribution, where the most frequent tags appear over one million times more often than tail tags. Pretraining on the natural hashtag distribution harms downstream transfer. Resampling the training set using square-root or uniform sampling yields a 5% to 6% top-1 accuracy improvement across ImageNet target benchmarks.

    Input: Training images D={Ij}j=1nD = \{I_j\}_{j=1}^n where image IjI_j has hashtags {hj,i}\{h_{j,i}\}; hashtag frequencies f(h)f(h); target schedule length TT; sampling mode mode∈{square-root,uniform}\text{mode} \in \{\text{square-root}, \text{uniform}\}
    Output: Resampled and permuted training sequence SS
    if mode==square-root\text{mode} == \text{square-root} then
        ϕ(x)←x\phi(x) \leftarrow \sqrt{x}
    else if mode==uniform\text{mode} == \text{uniform} then
        ϕ(x)←x\phi(x) \leftarrow x
    end if
    Find scaling threshold tt such that the final dataset length matches TT
    for each unique hashtag hh do
        r(h)←max⁡(1,ϕ(t/f(h)))r(h) \leftarrow \max(1, \phi(t / f(h)))
    end for
    S←[]S \leftarrow []
    for each image IjI_j with hashtags {hj,1,…,hj,m}\{h_{j,1}, \dots, h_{j,m}\} do
        r(Ij)←max⁡i=1mr(hj,i)r(I_j) \leftarrow \max_{i=1}^m r(h_{j,i})
        for k←1k \leftarrow 1 to r(Ij)r(I_j) do
            Create duplicated instance Ij′I'_j retaining only hashtags hj,ih_{j,i} satisfying k≤r(hj,i)k \le r(h_{j,i})
            Append Ij′I'_j to SS
        end for
    end for
    Randomly permute SS
    return SS
  6. Knowl 6 — Robustness of Billion-Scale Pretraining to Label Noise

    empirical result

    Convolutional networks trained on billion-image weakly supervised datasets exhibit extreme resilience to supervisory label noise. In an injection experiment using ResNeXt-101 32×16d32 \times 16\text{d} pretrained on 1 billion Instagram images with 17,000 hashtags (IG-1B-17k), p%p\% of the hashtags were artificially replaced with tags sampled from the marginal hashtag distribution (excluding the true tag). Downstream ImageNet performance using a fixed-feature linear classifier demonstrated:

    • At p=10%p = 10\% injected noise, top-1 accuracy drops by less than 1.0 percentage point across all target sets: ImageNet-1k drops from 82.1% to 81.5%, ImageNet-5k from 52.6% to 51.7%, and ImageNet-9k from 42.7% to 41.9%.
    • At p=25%p = 25\% injected noise, top-1 accuracy drops by approximately 2.0 percentage points (ImageNet-1k: 80.2%, ImageNet-5k: 50.3%, ImageNet-9k: 40.6%).
    • At p=50%p = 50\% injected noise (where half of all labels are randomized), top-1 accuracy degrades by only ~6 percentage points (ImageNet-1k: 76.1%, ImageNet-5k: 46.1%, ImageNet-9k: 36.6%).

    This demonstrates that massive dataset scale can compensate for substantial label noise inherent in uncurated web data.

  7. Knowl 7 — Transfer Learning to Object Detection and the Classification-Localization Tradeoff

    empirical result

    Finetuning Mask R-CNN with ResNeXt-101 FPN backbones pretrained on Instagram hashtags on the COCO dataset demonstrates distinct transfer properties compared to ImageNet pretraining:

    • Detection Improvements: ResNeXt-101 32×16d32 \times 16\text{d} pretrained on IG-1B-17k or IG-3.5B-17k improves COCO test-dev bounding box AP to 45.2% and mask AP to 39.7%, compared to 43.7% box AP and 38.6% mask AP for ImageNet-1k pretraining.
    • Capacity Dependence: Detection gains are model-capacity bound. For the smaller 32×4d32 \times 4\text{d} backbone, pretraining on larger datasets yields negligible or negative AP changes, whereas the larger 32×16d32 \times 16\text{d} backbone consistently benefits from larger pretraining corpora.
    • AP vs. AP@50 Discrepancy: Pretraining gains are predominantly concentrated in AP@50 (box AP@50 rises from 65.6% to 68.3% for 32×16d32 \times 16\text{d}), which allows looser spatial localization, whereas strict AP (averaged over IoU 0.5:0.95) shows smaller relative gains.
    • Degradation on Keypoint Detection: On COCO keypoint estimation, IG-1B-1.5k pretraining underperforms ImageNet-1k pretraining (65.3% vs. 67.0% keypoint AP). Weakly supervised hashtag classification encourages spatial and translation invariance that aids semantic categorization but can degrade precise spatial localization.
    • Finetuning Learning Rate: Models pretrained on Instagram require learning rates that are 4×4\times to 10×10\times lower during COCO finetuning (0.0025 for IG-1B and 0.00075 for IG-3.5B) compared to ImageNet-pretrained defaults (0.01).
  8. Knowl 8 — Efficacy of Linear Feature Transfer without Backbone Finetuning

    empirical result

    Features learned through large-scale hashtag prediction on billions of images are linearly separable to the extent that training a simple L2L_2-regularized linear logistic regression classifier on fixed, frozen representations achieves performance competitive with full end-to-end network finetuning:

    • On ImageNet-1k validation, a linear classifier trained on fixed features from a ResNeXt-101 32×16d32 \times 16\text{d} model pretrained on IG-3.5B-17k achieves 83.6% top-1 accuracy, closely approaching the 84.2% top-1 accuracy obtained by full network finetuning.
    • A linear classifier trained on fixed features from IG-940M-1.5k pretraining achieves 83.3% top-1 accuracy on ImageNet-1k.
    • Pretraining representations exhibit high statistical consistency: evaluating linear classifiers on ImageNet-1k, ImageNet-5k, and ImageNet-9k using two separate, independently sampled 1-billion-image Instagram datasets yields top-1 accuracy differences of less than 0.1% across all target benchmarks.
  9. Knowl 9 — Two-Stage Scalable Image Deduplication Pipeline

    algorithm

    To prevent data contamination between 3.5 billion pretraining images and evaluation test sets, a two-stage retrieval and verification pipeline identifies near-duplicate images across billion-scale datasets:

    Input: Pretraining database DD (3.5B images), Validation image collection QQ, Truncated ResNet-50 feature extractor (excluding final 5 conv layers), Distance threshold τ=0.6\tau = 0.6
    Output: Duplicate match pairs M⊂Q×DM \subset Q \times D
    Stage 1: Approximate Nearest Neighbor Index Construction
    for each image in DD do
        Resize image so that max⁡(height,width)=400\max(\text{height}, \text{width}) = 400 pixels
        Compute 2048-dimensional Regional Maximum Activation of Convolutions (R-MAC) features
        Whiten features using PCA to 512 dimensions and scalar-quantize to 8 bits/dimension
        L2-normalize and apply Optimized Product Quantization (OPQ) to reduce to 256 dimensions
    end for
    Build Faiss inverted index using coarse quantizer (2×142 \times 14 bits) and residual quantizer (32×832 \times 8 bits) trained on 2M images
    Stage 2: Target Image Querying and Exact Verification
    M←∅M \leftarrow \emptyset
    for each query image q∈Qq \in Q do
        Extract and quantize R-MAC descriptor for qq
        Query Faiss index to retrieve 128 nearest neighbor candidates from DD
        for each candidate dd in the 128 neighbors do
            Compute exact Euclidean distance e(q,d)=∥vq−vd∥22e(q, d) = \|\mathbf{v}_q - \mathbf{v}_d\|_2^2 between uncompressed 2048-dim R-MAC descriptors
            if e(q,d)<τe(q, d) < \tau then
                Flag (q,d)(q, d) for human inspection
            end if
        end for
        Manually inspect flagged candidate pairs (up to 21 nearest neighbors) to confirm visual duplicates
        Add confirmed pairs to MM
    end for
    return MM

    This pipeline identified duplicate rates of 0.30% in ImageNet-1k validation (150 images), 0.17% in CUB-200 (10 images), 0.41% in Places365 (151 images), and 0.12% in COCO validation (6 images).

  10. Knowl 10 — Synchronous Distributed Pretraining Infrastructure and Optimization Hyperparameters

    experimental setup

    Pretraining ResNeXt-101 architectures on up to 3.5 billion images employs a large-scale synchronous distributed SGD framework:

    • Hardware and Distributed Batching: Training runs on 336 GPUs across 42 machines with a global minibatch size of 8,064 images (24 images per GPU). Batch normalization statistics are computed independently per GPU across each 24-image local batch.
    • Optimizer and Initialization: Optimization utilizes SGD with Nesterov momentum of 0.9 and weight decay of 1×10−41 \times 10^{-4} (excluding batch normalization scale γ\gamma and bias β\beta). Residual block final batch normalization scale parameters γ\gamma are initialized to 0. Final fully connected weights are initialized from N(0,0.012)\mathcal{N}(0, 0.01^2).
    • Learning Rate Schedule: Uses the linear scaling rule with gradual warmup from an initial learning rate of 0.1 up to 0.1×8064256=3.150.1 \times \frac{8064}{256} = 3.15. Following warmup, the learning rate is multiplied by 0.5 at 20 equally spaced intervals across training (for IG-3.5B-17k, 40 intervals with decay factor 0.5\sqrt{0.5}). ImageNet pretraining baselines use 128 GPUs (minibatch 3,072) with 3 step reductions by a factor of 0.1.
    • Schedule Duration: The schedule length in terms of total images processed is determined by linearly interpolating between two boundary configurations: 120 epochs on 1.2M images (144M images processed) and 2 epochs on 3.5B images (7.0B images processed). Training ResNeXt-101 32×16d32 \times 16\text{d} on 3.5B images requires approximately 22 days on 336 GPUs.
  11. Knowl 11 — Hashtag Selection and WordNet-Based Canonical Merging

    model/method

    To derive structured supervision from public Instagram post captions, raw hashtags are filtered and merged using WordNet synsets:

    1. Hashtag-to-Synset Matching: Given a set of WordNet synsets SS and a hashtag string hh, a matching function s(S,h)s(S, h) returns the union of WordNet synsets matching hh directly or matching any bigram query formed by inserting a single space character at each position in hh.
    2. Canonical Merging: Two distinct hashtags hh and h′h' are defined as semantic duplicates and merged into a single canonical tag if and only if their WordNet synset mappings are identical:

    s(S,h)=s(S,h′)s(S, h) = s(S, h')

    When SS is set to all WordNet synsets, this conservative merging policy unites hashtags only when they coincide across all WordNet senses (e.g., #brownbear and #ursusarctos are merged into one canonical label). 3. Hashtag Vocabulary Sets: Three primary vocabulary sets are constructed from public captions:

    • 1.5k set: ~1,500 canonical hashtags matching the 1,000 synsets in ImageNet-1k.
    • 17k set: ~17,000 canonical hashtags matching any noun synset in WordNet.
    • 8.5k set: ~8,500 most frequent canonical hashtags from the 17k set.

Coverage note — Omitted specific grid-search learning rate tables for secondary classification finetuning benchmarks and conservative lower-bound accuracy tables that yielded identical scientific conclusions to the primary results.

References

  1. 1.Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: CVPR. (2014)
  2. 2.Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., Darrell, T.: DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition. arXiv:1310.1531 (2013)
  3. 3.Zeiler, M.D., Fergus, R.: Visualizing and understanding convolutional neural networks. In: ECCV. (2014)
  4. 4.He, K., Gkioxari, G., Dollar, P., Girshick, R.: Mask R-CNN. In: ICCV. (2017)
  5. 5.Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: CVPR. (2015)
  6. 6.Zhao, H., Shi, J., Qi, X., Wang, X., Jia, J.: Pyramid scene parsing network. In: CVPR. (2017)
  7. 7.Cao, Z., Simon, T., Wei, S.E., Sheikh, Y.: Realtime multi-person 2d pose estimation using part affinity fields. In: CVPR. (2017)
  8. 8.Papandreou, G., Zhu, T., Kanazawa, N., Toshev, A., Tompson, J., Bregler, C., Murphy, K.: Towards accurate multi-person pose estimation in the wild. In: CVPR. (2017)
  9. 9.Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: CVPR. (2017)
  10. 10.Eigen, D., Fergus, R.: Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. In: ICCV. (2015)
  11. 11.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: ImageNet Large Scale Visual Recognition Challenge. IJCV (2015)
  12. 12.Agrawal, P., Girshick, R., Malik, J.: Analyzing the performance of multilayer neural networks for object recognition. In: ECCV. (2014)
  13. 13.Huh, M., Agrawal, P., Efros, A.: What makes ImageNet good for transfer learning? arXiv:1608.08614 (2016)
  14. 14.Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., Torralba, A.: Places: A 10 million image database for scene recognition. PAMI (2017)
  15. 15.Xie, S., Girshick, R., Dollar, P., Tu, Z., He, K.: Aggregated residual transformations for deep neural networks. In: CVPR. (2017)
  16. 16.Joulin, A., van der Maaten, L., Jabri, A., Vasilache, N.: Learning visual features from large weakly supervised data. In: Proceedings of the European Conference on Computer Vision (ECCV), Springer (2016) 67–84
  17. 17.Sun, C., Shrivastava, A., Singh, S., Gupta, A.: Revisiting unreasonable effectiveness of data in deep learning era. In: Proc. ICCV. (2017)
  18. 18.Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll´ar, P., Zitnick, C.L.: Microsoft COCO: Common objects in context. In: European conference on computer vision, Springer (2014) 740–755
  19. 19.Goyal, P., Dollar, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., He, K.: Accurate, large minibatch SGD: Training ImageNet in 1 hour. In: arXiv:1706.02677. (2017)
  20. 20.WordNet: About WordNet. http://wordnet.princeton.edu (2010)
  21. 21.Welinder, P., Branson, S., Mita, T., Wah, C., Schroff, F., Belongie, S., Perona, P.: Caltech-UCSD Birds 200. Technical report, Caltech (2010)
  22. 22.Gordo, A., Almazan, J., Revaud, J., Larlus, D.: Deep image retrieval: Learning global representations for image search. In: arXiv:1604.01325. (2016)
  23. 23.Tolias, G., Sicre, R., , Jegou, H.: Particular object retrieval with integral maxpooling of cnn activations. In: ICLR. (2016)
  24. 24.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR. (2016)
  25. 25.Huang, G., Liu, Z., Weinberger, K., van der Maaten, L.: Densely connected convolutional networks. In: CVPR. (2017)
  26. 26.Szegedy, C., Ioffe, S., Vanhoucke, V., Alemi, A.: Inception-v4, inception-resnet and the impact of residual connections on learning. In: arXiv:1602.07261. (2016)
  27. 27.Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: ICML. (2015)
  28. 28.He, K., Zhang, X., Ren, S., Sun, J.: Delving deep into rectifiers: Surpassing humanlevel performance on imagenet classification. In: ICCV. (2015)
  29. 29.Pathak, D., Girshick, R., Doll´ar, P., Darrell, T., Hariharan, B.: Learning features by watching objects move. In: CVPR. (2017)
  30. 30.Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: CVPR. (2009)
  31. 31.Zoph, B., Vasudevan, V., Shlens, J., Le, Q.: Learning transferable architectures for scalable image recognition. In: arXiv:1707.07012. (2017)
  32. 32.Storck, P., Cisse, M.: Convnets and imagenet beyond accuracy: Explanations, bias detection, adversarial examples and model criticism. In: arXiv:1711.11443. (2017)
  33. 33.Misra, I., Zitnick, C.L., Mitchell, M., Girshick, R.: Seeing through the human reporting bias: Visual classifiers from noisy human-centric labels. In: CVPR. (2016)
  34. 34.Mikolov, T., Sutskever, I., Chen, K., Corrado, G., Dean, J.: Distributed representations of words and phrases and their compositionality. In: NIPS. (2013)
  35. 35.Brysbaert, M., Warriner, A.B., Kuperman, V.: Concreteness ratings for 40 thousand generally known english word lemmas. Behavior Research Methods (2014)
  36. 36.Girshick, R., Radosavovic, I., Gkioxari, G., Doll´ar, P., He, K.: Detectron. https://github.com/facebookresearch/detectron (2018)
  37. 37.Lin, T.Y., Doll´ar, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: CVPR. (2017)
  38. 38.Li, A., Jabri, A., Joulin, A., van der Maaten, L.: Learning visual n-grams from web data. In: Proc. ICCV. (2017)
  39. 39.Thomee, B., Shamma, D.A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., Li, L.J.: Yfcc100m: The new data in multimedia research. Communications of the ACM 59(2) (2016) 64–73
  40. 40.Veit, A., Nickel, M., Belongie, S., van der Maaten, L.: Separating self-expression and visual content in hashtag supervision. In: arXiv 1711.09825. (2017)
  41. 41.Gross, S., Ranzato, M., Szlam, A.: Hard mixtures of experts for large scale weakly supervised vision. In: CVPR. (2017)
  42. 42.Denton, E., Weston, J., Paluri, M., Bourdev, L., Fergus, R.: User conditional hashtag prediction for images. In: Proc. KDD. (2015) 1731–1740
  43. 43.Schroff, F., Kalenichenko, D., Philbin, J.: Facenet: A unified embedding for face recognition and clustering. In: CVPR. (2015)
  44. 44.Taigman, Y., Yang, M., Ranzato, M., Wolf, L.: Web-scale training for face identification. In: CVPR. (2015)
  45. 45.Johnson, J., Douze, M., J´egou, H.: Billion-scale similarity search with GPUs. arXiv preprint arXiv:1702.08734 (2017)
  46. 46.Stewenius, H., Gunderson, S., Pilet, J.: Size matters: Exhaustive geometric verification for image retrieval. In: Proceedings of the European Conference on Computer Vision (ECCV), Springer (2012)
  47. 47.Jegou, H., Douze, M., Schmid, C.: Hamming embedding and weak geometry consistency for large scale image search. In: Proceedings of the European Conference on Computer Vision (ECCV). (2008)
  48. 48.Ge, T., He, K., Ke, Q., Sun, J.: Optimized product quantization. PAMI 36(4) (2013) 744–755
  49. 49.Jegou, H., Douze, M., Schmid, C.: Product quantization for nearest neighbor search. PAMI 33(1) (2011) 117–128
  50. 50.Nesterov, Y.: Introductory lectures on convex optimization: A basic course. Springer (2004)
  51. 51.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., Wojna, Z.: Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). (2016)
  52. 52.Chollet, F.: Xception: Deep learning with depthwise separable convolutions. In: arXiv:1610.02357. (2016)
  53. 53.Szegedy, C., Ioffe, S., Vanhoucke, V., Alemi, A.: Inceptionv4, inception-resnet and the impact of residual connections on learning. In: Proceedings of the International Conference on Learning Representations (ICLR) Workshop. (2016)
  54. 54.Zhang, X., Li, Z., Loy, C., Lin, D.: Polynet: A pursuit of structural diversity in very deep networks. In: arXiv:1611.05725. (2016)
  55. 55.Chen, Y., Li, J., Xiao, H., Jin, X., Yan, S., Feng, J.: Dual path networks. In: arXiv:1707.01629. (2017)
  56. 56.Hu, J., Shen, L., Sun, G.: Squeeze-and-excitation networks. In: arXiv:1709.01507. (2017)

Citation

MLA
Mahajan, D., et al. “Exploring the Limits of Weakly Supervised Pretraining”. arXiv, 2018, http://arxiv.org/abs/1805.00932v1.
APA
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., & Maaten, L. van . der . (2018). Exploring the Limits of Weakly Supervised Pretraining. arXiv. http://arxiv.org/abs/1805.00932v1
Chicago
Mahajan, D., R. Girshick, V. Ramanathan, et al. 2018. “Exploring the Limits of Weakly Supervised Pretraining”. arXiv. http://arxiv.org/abs/1805.00932v1.
Harvard
Mahajan, D. et al. (2018) “Exploring the Limits of Weakly Supervised Pretraining”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1805.00932v1.
Vancouver
1. Mahajan D, Girshick R, Ramanathan V, He K, Paluri M, Li Y, Bharambe A, Maaten L van der (2018) Exploring the Limits of Weakly Supervised Pretraining. arXiv

BibTeX

@article{mahajan2018exploring,
  title = {Exploring the Limits of Weakly Supervised Pretraining},
  author = {Mahajan, Dhruv and Girshick, Ross and Ramanathan, Vignesh and He, Kaiming and Paluri, Manohar and Li, Yixuan and Bharambe, Ashwin and Maaten, Laurens van der},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1805.00932v1},
  eprint = {1805.00932}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF