Revisiting Unreasonable Effectiveness of Data in Deep Learning Era

Chen SunAbhinav ShrivastavaSaurabh SinghAbhinav Gupta

article2017ICCV2,771 citations

Demonstrates through experiments on the 300-million-image JFT dataset that visual task performance scales logarithmically with data volume, proving that scaling pre-training data directly improves downstream accuracy across classification, detection, and segmentation.

Listen

The article addresses a key question in computer vision: while model capacity and GPU power have grown substantially since 2012, the largest widely used training dataset has stayed fixed at roughly one million ImageNet images. This stagnation raises the practical issue of whether performance on core tasks would continue to improve if training data were scaled by factors of ten or one hundred.

The article set out to measure how visual representations learned from a much larger, automatically labeled collection affect downstream performance. Researchers pre-trained standard ResNet architectures on the JFT-300M dataset of 300 million images carrying 375 million noisy labels across 18,000 categories, then evaluated the resulting models on image classification, object detection, semantic segmentation, and human pose estimation.

Experiments compared models trained from scratch on JFT-300M, models initialized from ImageNet and then trained on JFT-300M, and conventional ImageNet-only baselines. Subsets of 10 million, 30 million, and 100 million images were also tested, along with variations in model depth and label vocabulary size.

Performance on every task rose steadily with data volume, following a clear logarithmic relationship. Pre-training on the full 300 million images produced new state-of-the-art numbers, including a 3.1-point gain in COCO average precision (37.4 versus 34.3) and comparable lifts on PASCAL VOC detection and segmentation. Larger-capacity networks captured more benefit from the added data, while label noise of roughly 20 percent and a long-tailed category distribution did not prevent convergence or gains. Increasing the number of images mattered more than expanding the label vocabulary.

These results indicate that representation learning remains a high-leverage direction and that further scaling of training data can still deliver measurable improvements even after models have grown deeper. Because the observed gains are logarithmic, organizations can expect continued but diminishing returns that must be weighed against the cost of data collection and training.

The article recommends renewed collective investment in larger, diverse datasets and renewed attention to unsupervised or self-supervised methods that could exploit similar scale without exhaustive labeling. It also notes that the reported numbers likely underestimate potential gains, because training schedules were carried over from the 1-million-image regime without extensive retuning. Readers should treat the precise magnitudes as directional rather than definitive until larger-scale hyper-parameter studies are performed.

Cover for Revisiting Unreasonable Effectiveness of Data in Deep Learning Era

Abstract

The success of deep learning in vision can be attributed to: (a) models with high capacity; (b) increased computational power; and (c) availability of large-scale labeled data. Since 2012, there have been significant advances in representation capabilities of the models and computational capabilities of GPUs. But the size of the biggest dataset has surprisingly remained constant. What will happen if we increase the dataset size by 10x or 100x? This paper takes a step towards clearing the clouds of mystery surrounding the relationship between `enormous data' and visual deep learning. By exploiting the JFT-300M dataset which has more than 375M noisy labels for 300M images, we investigate how the performance of current vision tasks would change if this data was used for representation learning. Our paper delivers some surprising (and some expected) findings. First, we find that the performance on vision tasks increases logarithmically based on volume of training data size. Second, we show that representation learning (or pre-training) still holds a lot of promise. One can improve performance on many vision tasks by just training a better base model. Finally, as expected, we present new state-of-the-art results for different vision tasks including image classification, object detection, semantic segmentation and human pose estimation. Our sincere hope is that this inspires vision community to not undervalue the data and develop collective efforts in building larger datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The JFT-300M Dataset
  • 4 Training and Evaluation Framework
  • 4.1 Training on JFT-300M Data
  • 4.2 Monitoring Training Progress
  • 4.3 Evaluating the Visual Representations
  • 5 Experiments
  • 5.1 Image Classification
  • 5.2 Object Detection
  • 5.3 Semantic Segmentation
  • 5.4 Human Pose Estimation
  • 6 Discussions
  • References

Knowls

  1. Knowl 1 — Logarithmic Scaling of Downstream Performance with Pre-training Data Volume

    empirical result

    When pre-training convolutional networks (such as ResNet-101) on subsets of web-scale data (10M, 30M, 100M, and 300M images from JFT-300M), the transfer performance on downstream computer vision tasks increases logarithmically as a function of the pre-training dataset size.

    This logarithmic relationship holds consistently across diverse downstream benchmarks:

    • COCO Object Detection: Mean average precision (mAP@[0.5, 0.95]) increases steadily with dataset size. This behavior is evident both when fine-tuning all model weights and when freezing the pre-trained feature extractor (all layers prior to the conv5 block).
    • PASCAL VOC 2007 Object Detection: [email protected] follows the same logarithmic growth trajectory.
    • PASCAL VOC 2012 Semantic Segmentation: Mean intersection-over-union (mIOU) improves continuously as pre-training data scales up to 300M images.

    Unlike prior studies on large-scale weakly supervised datasets that observed performance plateauing around 100M images, representation quality exhibits no signs of plateauing up to 300M images when paired with high-capacity neural networks.

  2. Knowl 2 — Model Capacity Dependency in Web-Scale Visual Representation Learning

    empirical result

    Fully exploiting massive pre-training datasets requires deep neural networks with sufficient representational capacity. When evaluating ResNet architectures of varying depths (ResNet-50, ResNet-101, and ResNet-152) pre-trained on either ImageNet (1.2M images) or JFT-300M (300M images) and transferred to the COCO object detection minival* benchmark:

    Architecture ImageNet Pre-training JFT-300M Pre-training Absolute Gain ()
    ResNet-50 31.6% 33.5% +1.9%
    ResNet-101 34.5% 36.8% +2.3%
    ResNet-152 34.7% 37.7% +3.0%

    While deeper models show signs of plateauing when trained solely on ImageNet (ResNet-152 only outperforms ResNet-101 by 0.2% mAP@[0.5, 0.95]), the performance gap widens significantly when trained on JFT-300M (ResNet-152 outperforms ResNet-101 by 0.9% mAP@[0.5, 0.95], achieving a 3.0% gain over its ImageNet baseline). High model capacity is essential to capture the visual diversity of 300 million images.

  3. Knowl 3 — JFT-300M Pre-training Framework and Multi-Label Loss Formulation

    model/method

    The JFT-300M pre-training dataset consists of over 300 million images associated with 375 million non-mutually exclusive labels across K=18,291K = 18{,}291 categories (averaging 1.26 labels per image with approximately 20% label noise). The category space forms a directed hierarchy with a maximum depth of 12 levels and parent nodes having up to 2,876 children. The category distribution is long-tailed, containing over 3,000 classes with fewer than 100 images each.

    To pre-train a ResNet-101 architecture on this data:

    1. Label Hierarchy Completion: Missing parent labels are populated according to the taxonomy (e.g., an image labeled apple is augmented with the label fruit).
    2. Multi-Label Objective: The final layer has K=18,291K = 18{,}291 outputs. A per-label binary logistic cross-entropy loss is computed over all categories, treating non-present classes as negatives: L=c=1K[yclogσ(zc)+(1yc)log(1σ(zc))]\mathcal{L} = -\sum_{c=1}^{K} \left[ y_c \log \sigma(z_c) + (1 - y_c) \log (1 - \sigma(z_c)) \right] where yc{0,1}y_c \in \{0, 1\} denotes the ground-truth binary presence for class cc, zcz_c is the unnormalized logit output for class cc, and σ(zc)=11+ezc\sigma(z_c) = \frac{1}{1 + e^{-z_c}}.
    3. Training & Optimization: Images are resized to 340×340340 \times 340, randomly cropped to 299×299299 \times 299, and augmented with random horizontal reflections. Training uses the RMSProp optimizer with momentum 0.90.9, batch size 32, weight decay 10410^{-4}, and initial learning rate 10310^{-3} decayed by 0.90.9 every 3M steps.
    4. Distributed Infrastructure: Optimization is executed via Downpour asynchronous SGD across 50 NVIDIA K80 GPUs and 17 parameter servers. The 36M-parameter final classification layer (2048 input units ×\times 18,291 output units) is split vertically into 50 equal sub-layers distributed across parameter servers.
  4. Knowl 4 — Pre-training Data Scale Dominates Label Vocabulary Granularity

    empirical result

    To isolate the effect of label vocabulary diversity from image dataset volume, a controlled experiment compares ResNet-101 models pre-trained on two distinct 30-million-image subsets of JFT-300M:

    1. A 30M-image subset restricted to 941 classes that directly map to the 1,000 ImageNet categories.
    2. A 30M-image subset sampled across the full 18,291 JFT category space.

    Both models are pre-trained for 4 epochs and fine-tuned for object detection on COCO using Faster R-CNN:

    Pre-training Class Vocabulary COCO minival* mAP@[0.5, 0.95]
    1K ImageNet-aligned classes (30M images) 31.2%
    18K JFT classes (30M images) 31.9%

    The two configurations achieve comparable performance (difference 0.7%\le 0.7\%), demonstrating that the primary driver of visual representation quality in large-scale pre-training is the number of distinct training images rather than the granularity or breadth of the label vocabulary.

  5. Knowl 5 — Object Detection Transfer Performance on COCO and PASCAL VOC

    data/table

    Transferring ResNet-101 representations pre-trained on JFT-300M (either trained from random initialization, denoted 300M, or initialized from ImageNet before JFT pre-training, denoted ImageNet+300M) to the Faster R-CNN detection framework yields large improvements over standard ImageNet-trained baselines across COCO and PASCAL VOC benchmarks:

    Initialization Checkpoint COCO test-dev [email protected] COCO test-dev mAP@[0.5, 0.95] VOC 2007 Test [email protected]
    ResNet-101 (He et al.) 53.3% 32.2% -
    ResNet-101 (ImageNet) 53.6% 34.3% 76.3%
    Inception-ResNet-v2 (ImageNet) 56.3% 35.5% -
    ResNet-101 (300M) 56.9% 36.7% 81.4%
    ResNet-101 (ImageNet+300M) 58.0% 37.4% 81.3%

    JFT-300M pre-training provides a +3.1% boost on COCO test-dev mAP@[0.5, 0.95] and a +5.1% boost on PASCAL VOC 2007 [email protected] compared to the competitive ResNet-101 ImageNet baseline. The gain surpasses the improvements achieved by switching to more complex architectures like Inception-ResNet-v2 trained on ImageNet.

  6. Knowl 6 — Image Classification Transfer Performance on ImageNet

    data/table

    When a ResNet-101 pre-trained on JFT-300M for 36M iterations (4 epochs) is fine-tuned on the standard ImageNet ILSVRC 2012 classification training set (1.2M images, 1000 classes) for 4M iterations, it substantially outperforms models trained directly from scratch on ImageNet:

    Initialization Checkpoint Top-1 Accuracy (%) Top-5 Accuracy (%)
    MSRA Checkpoint (He et al.) 76.4 92.9
    Random Initialization (from scratch) 77.5 93.9
    Fine-tuned from JFT-300M 79.2 94.7

    Evaluations use single-model, single-crop inference on the 50,000-image ImageNet validation set. Pre-training on 300M web images provides an absolute gain of +1.7% in top-1 accuracy and +0.8% in top-5 accuracy over the scratch baseline.

  7. Knowl 7 — Semantic Segmentation Transfer Performance on PASCAL VOC 2012

    data/table

    Visual representations pre-trained on JFT-300M transfer effectively to dense pixel-level tasks. Using the DeepLab-ASPP-L framework (which adds four parallel atrous convolution branches with dilation rates r{6,12,8,24}r \in \{6, 12, 8, 24\} after the ResNet-101 conv5 block), models are fine-tuned on the augmented PASCAL VOC 2012 trainaug set (10,582 images) and evaluated on the VOC 2012 validation set (1,449 images):

    Initialization Checkpoint Mean Intersection-over-Union (mIOU %)
    ImageNet Pre-trained 73.6
    JFT-300M Pre-trained (from scratch) 75.3
    JFT-300M Pre-trained (ImageNet-initialized) 76.5

    JFT-300M pre-training provides an absolute improvement of up to +2.9% mIOU over ImageNet initialization without requiring multi-scale inference, CRF post-processing, or segmentation data augmentation. Per-category analysis shows improvements exceeding 7 percentage points for challenging classes such as boat and horse.

  8. Knowl 8 — Human Pose Estimation Transfer Performance on COCO

    data/table

    When applied to keypoint detection using the fully-convolutional G-RMI pose estimation framework with a ResNet-101 backbone, representations pre-trained on JFT-300M yield superior multi-person keypoint localization on the COCO test-dev split compared to ImageNet-initialized counterparts:

    Method / Initialization Average Precision (AP) [email protected] Average Recall (AR) [email protected]
    CMU Pose (Cao et al.) 61.8 84.9 66.5 87.2
    ImageNet Pre-trained (G-RMI) 62.4 84.0 66.7 86.6
    JFT-300M Pre-trained (from scratch) 64.8 85.8 69.4 88.4
    JFT-300M Pre-trained (ImageNet-initialized) 64.4 85.7 69.1 88.2

    All models are evaluated in a single-model setting without ensembling on identical person bounding box detections. JFT-300M pre-training delivers a +2.4 point improvement in Average Precision (AP) and +2.7 points in Average Recall (AR) over the ImageNet baseline.

  9. Knowl 9 — Impact of Pre-training Epochs and Step Budgets on Downstream Performance

    empirical result

    Downstream transfer performance scales with the number of training iterations and epochs completed during JFT-300M pre-training. Evaluating ResNet-101 on COCO minival* object detection across training intervals demonstrates consistent gains:

    JFT-300M Iterations JFT-300M Epochs COCO minival* mAP@[0.5, 0.95]
    12M 1.3 35.0%
    24M 2.6 36.1%
    36M 4.0 36.8%

    In contrast, pre-training on ImageNet requires far more epochs (150 epochs, 5M iterations) to reach 34.5% mAP@[0.5, 0.95]. Furthermore, tracking training progress on the FastEval14k validation benchmark reveals that while initializing JFT-300M models from ImageNet weights accelerates convergence in the first 15M iterations, training from random initialization reaches identical asymptotic performance by 36M iterations.

  10. Knowl 10 — Robustness of Transfer Gains to Dataset De-duplication

    empirical result

    Because web-scale datasets may contain images that duplicate downstream test or validation sets, a rigorous de-duplication analysis was conducted using visual deep feature embeddings to identify and remove near-duplicate images across evaluation sets:

    • ImageNet validation: 5,536 of 50,000 images had near-duplicates in JFT-300M.
    • COCO minival*: 1,648 of 8,000 images had near-duplicates in JFT-300M.
    • PASCAL VOC 2007 test: 201 of 4,952 images had near-duplicates in JFT-300M.
    • PASCAL VOC 2012 validation: 84 of 1,449 images had near-duplicates in JFT-300M.

    Re-evaluating models after purging all duplicate images from the validation and test sets resulted in negligible performance changes:

    • ImageNet Top-1 Accuracy: 79.2% (original) vs. 79.3% (de-duplicated).
    • COCO minival* mAP@[0.5, 0.95]: 37.8% (original) vs. 37.7% (de-duplicated).
    • PASCAL VOC 2007 Detection [email protected]: 81.3% (original) vs. 81.2% (de-duplicated).
    • PASCAL VOC 2012 Segmentation mIOU: 76.5% (original) vs. 76.5% (de-duplicated).

    These results confirm that the empirical gains from JFT-300M pre-training are driven by generalizable representation learning rather than nearest-neighbor memorization or validation set contamination.

  11. Knowl 11 — Underestimation of Pre-training Impact Due to Fixed Hyperparameter Regimes

    limitation

    Due to immense computational demands (training a single ResNet-101 on JFT-300M for 4 epochs required two months of continuous execution on 50 NVIDIA K80 GPUs), the optimization hyperparameters, learning rate decay schedules, and batch sizes were directly adopted from standard ImageNet (1M image) conventions without extensive hyperparameter search.

    Because these optimization schedules were engineered and tuned for dataset sizes two orders of magnitude smaller, the reported quantitative transfer metrics likely represent a conservative lower bound on the true representation learning capacity of 300-million-image pre-training.

Coverage note — None omitted. All major experimental setups, transfer learning benchmarks (ImageNet classification, COCO/PASCAL detection, PASCAL segmentation, COCO pose estimation), scaling analyses (data volume, model capacity, vocabulary size, epoch count), de-duplication audits, and stated computational limitations are fully represented.

References

  1. 1.P. Agrawal, R. B. Girshick, and J. Malik. Analyzing the performance of multilayer neural networks for object recognition. In ECCV, 2014.
  2. 2.A. Bergamo and L. Torresani. Exploiting weakly-labeled web images to improve object classification: a domain adaptation approach. In NIPS. 2010.
  3. 3.Z. Cao, T. Simon, S. Wei, and Y. Sheikh. Realtime multi-person 2d pose estimation using part affinity fields. arXiv:1611.08050, 2016.
  4. 4.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. arXiv preprint arXiv:1606.00915, 2016.
  5. 5.X. Chen and A. Gupta. Webly supervised learning of convolutional networks. In ICCV, 2015.
  6. 6.X. Chen, A. Shrivastava, and A. Gupta. Neil: Extracting visual knowledge from web data. In ICCV, 2013.
  7. 7.F. Chollet. Xception: Deep learning with depthwise separable convolutions. arXiv:1610.02357, 2016.
  8. 8.J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. W. Senior, P. A. Tucker, K. Yang, and A. Y. Ng. Large scale distributed deep networks. In NIPS, 2012.
  9. 9.S. Divvala, A. Farhadi, and C. Guestrin. Learning everything about anything: Webly-supervised visual concept learning. In CVPR, 2014.
  10. 10.C. Doersch, A. Gupta, and A. A. Efros. Unsupervised visual representation learning by context prediction. In ICCV, 2015.
  11. 11.J. Donahue, P. Krähenbühl, and T. Darrell. Adversarial feature learning. arXiv:1605.09782, 2016.
  12. 12.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The Pascal Visual Object Classes (VOC) Challenge. IJCV, 2010.
  13. 13.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The pascal visual object classes (voc) challenge. IJCV, 2010.
  14. 14.R. B. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. arXiv:1311.2524, 2013.
  15. 15.B. Hariharan, P. Arbeláez, L. Bourdev, S. Maji, and J. Malik. Semantic contours from inverse detectors. In ICCV, 2011.
  16. 16.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  17. 17.G. Hinton, O. Vinyals, and J. Dean. Distilling the knowledge in a neural network. In NIPS, 2014.
  18. 18.J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama, and K. Murphy. Speed/accuracy trade-offs for modern convolutional object detectors. In CVPR, 2017.
  19. 19.M. Huh, P. Agrawal, and A. A. Efros. What makes imagenet good for transfer learning? arXiv:1608.08614, 2016.
  20. 20.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv:1502.03167, 2015.
  21. 21.H. Izadinia, B. C. Russell, A. Farhadi, M. D. Hoffman, and A. Hertzmann. Deep classifiers from image tags in the wild. In ACM MM, 2015.
  22. 22.M. Jain, J. C. van Gemert, and C. G. Snoek. What do 15,000 object categories tell us about classifying and localizing actions? In CVPR, 2015.
  23. 23.A. Joulin, L. van der Maaten, A. Jabri, and N. Vasilache. Learning visual features from large weakly supervised data. arXiv:1511.02251, 2015.
  24. 24.J. Krause, B. Sapp, A. Howard, H. Zhou, A. Toshev, T. Duerig, J. Philbin, and F. Li. The unreasonable effectiveness of noisy data for fine-grained recognition. arXiv:1511.06789, 2015.
  25. 25.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  26. 26.T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft COCO: common objects in context. In ECCV, 2014.
  27. 27.K. Ni, R. A. Pearce, K. Boakye, B. V. Essen, D. Borth, B. Chen, and E. X. Wang. Large-scale deep learning on the YFCC100M dataset. arXiv:1502.03409, 2015.
  28. 28.M. Oquab, L. Bottou, I. Laptev, and J. Sivic. Learning and transferring mid-level image representations using convolutional neural networks. In CVPR, 2014.
  29. 29.G. Papandreou, T. Zhu, N. Kanazawa, A. Toshev, J. Tompson, C. Bregler, and K. Murphy. Towards accurate multi-person pose estimation in the wild. arXiv:1701.01779, 2017.
  30. 30.F. Pereira, P. Norvig, and A. Halev. The unreasonable effectiveness of data. IEEE Intelligent Systems, 2009.
  31. 31.L. Pinto, D. Gandhi, Y. Han, Y. Park, and A. Gupta. The curious robot: Learning visual representations via physical interactions. arXiv:1604.01360, 2016.
  32. 32.L. Pinto and A. Gupta. Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours. arXiv:1509.06825, 2015.
  33. 33.S. Ren, K. He, R. Girshick, and J. Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015.
  34. 34.M. Rubinstein, A. Joulin, J. Kopf, and C. Liu. Unsupervised joint object discovery and segmentation in internet images. CVPR, 2013.
  35. 35.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li. Imagenet large scale visual recognition challenge. arXiv:1409.0575, 2014.
  36. 36.K. Simonyan and A. Zisserman. Two-stream convolutional networks for action recognition in videos. arXiv:1406.2199, 2014.
  37. 37.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556, 2014.
  38. 38.C. Szegedy, S. Ioffe, and V. Vanhoucke. Inception-v4, inception-resnet and the impact of residual connections on learning. arXiv:1602.07261, 2016.
  39. 39.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015.
  40. 40.B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L. Li. The new data and new challenges in multimedia research. arXiv:1503.01817, 2015.
  41. 41.A. Torralba and A. Efros. Unbiased look at dataset bias. CVPR, 2011.
  42. 42.C. Vondrick, H. Pirsiavash, and A. Torralba. Generating videos with scene dynamics. In NIPS, 2016.
  43. 43.X. Wang and A. Gupta. Unsupervised learning of visual representations using videos. arXiv:1505.00687, 2015.
  44. 44.T. Weyand, I. Kostrikov, and J. Philbin. Planet - photo geolocation with convolutional neural networks. arXiv:1602.05314, 2016.

Citation

MLA
Sun, C., et al. “Revisiting Unreasonable Effectiveness of Data in Deep Learning Era”. 2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp. 843–52, https://doi.org/10.1109/ICCV.2017.97.
APA
Sun, C., Shrivastava, A., Singh, S., & Gupta, A. (2017). Revisiting Unreasonable Effectiveness of Data in Deep Learning Era. 2017 IEEE International Conference on Computer Vision (ICCV), 843–852. https://doi.org/10.1109/ICCV.2017.97
Chicago
Sun, C., A. Shrivastava, S. Singh, and A. Gupta. 2017. “Revisiting Unreasonable Effectiveness of Data in Deep Learning Era”. 2017 IEEE International Conference on Computer Vision (ICCV), 843–52. https://doi.org/10.1109/ICCV.2017.97.
Harvard
Sun, C. et al. (2017) “Revisiting Unreasonable Effectiveness of Data in Deep Learning Era”, 2017 IEEE International Conference on Computer Vision (ICCV). IEEE, pp. 843–852. Available at: https://doi.org/10.1109/ICCV.2017.97.
Vancouver
1. Sun C, Shrivastava A, Singh S, Gupta A (2017) Revisiting Unreasonable Effectiveness of Data in Deep Learning Era. In: 2017 IEEE International Conference on Computer Vision (ICCV). IEEE, pp 843–852

BibTeX

@inproceedings{Sun_2017, title={Revisiting Unreasonable Effectiveness of Data in Deep Learning Era}, url={http://dx.doi.org/10.1109/ICCV.2017.97}, DOI={10.1109/iccv.2017.97}, booktitle={2017 IEEE International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Sun, Chen and Shrivastava, Abhinav and Singh, Saurabh and Gupta, Abhinav}, year={2017}, month=Oct, pages={843–852} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE