Object Detectors Emerge in Deep Scene CNNs

Bolei ZhouAditya KhoslaAgata LapedrizaAude OlivaAntonio Torralba

article2014ICLR1,331 citations

Demonstrates that convolutional neural networks trained solely on scene classification spontaneously develop internal object detectors, enabling simultaneous scene recognition and object localization without object-level supervision.

Listen

Deep neural networks achieve state-of-the-art performance across computer vision benchmarks, yet their internal decision-making processes often operate as an uninterpretable "black box." Understanding what internal network layers actually learn is essential for validating system reliability, improving efficiency, and designing better vision architectures. The article addresses this challenge by evaluating whether neural networks implicitly discover interpretable visual concepts when trained on broad contextual tasks.

The article demonstrates that training a convolutional neural network solely on whole-image scene classification naturally forces internal units to emerge as dedicated, localized object detectors. It proves that this internal emergence occurs without requiring explicit object-level annotations or bounding boxes during training.

To demonstrate this behavior, the authors trained identical network architectures on two distinct datasets: 2.4 million scene-centric images from the Places database across 205 categories and 1.3 million object-centric images from ImageNet across 1,000 categories. They quantified internal representations by creating minimal image representations, measuring empirical receptive fields using dense patch occlusions across 200,000 test images, and crowdsourcing concept annotations from human workers. The emergent detectors were then quantitatively benchmarked against fully annotated ground-truth images from the SUN database.

The investigation produced several key findings. First, object detectors emerge automatically in scene-trained networks at a significantly higher rate than in object-trained networks, despite the scene network receiving zero object-level supervision. Second, empirical receptive fields are substantially more localized and compact than theoretical limits, expanding into complex semantic concepts as network depth increases; for instance, the final convolutional pooling layer shows empirical receptive field sizes around 72 pixels compared to the theoretical 195 pixels. Third, the specific object detectors that emerge strongly match the most discriminative elements needed to separate scene categories, exhibiting a high 0.84 correlation with scene informativeness compared to a 0.54 correlation with raw object frequency. Fourth, multiple units often specialize in different visual appearances or subcategories of the same object, such as separate units detecting desk lamps versus ceiling lamps. Finally, a single scene-trained network can perform both overall scene recognition and multi-object localization simultaneously in a single computational forward-pass.

These findings indicate that complex visual systems do not always require expensive, multi-stage pipelines or separate models for scene classification and object localization. Instead, high-level task objectives naturally organize internal representations into shared, hierarchical components ranging from low-level edges to high-level objects. This can significantly reduce computational overhead, annotation costs, and system latency in operational computer vision deployments.

Practitioners should leverage scene-classification networks to extract dual outputs—both global scene labels and localized object detections—within a single execution pass. Research teams should also explore what additional high-level visual tasks can induce unsupervised learning of other semantic concepts. However, decision-makers should recognize that about 115 out of 256 units in the final pooling layer of the scene network did not map to discrete objects, pointing to complementary texture- or part-based representations. Because the emergent detectors are strictly limited to objects that are informative for distinguishing target scenes, users should maintain high confidence in the network's discriminative features while noting that non-discriminative background objects will not be discovered without further supervision.

Cover for Object Detectors Emerge in Deep Scene CNNs

Abstract

With the success of new computational architectures for visual processing, such as convolutional neural networks (CNN) and access to image databases with millions of labeled examples (e.g., ImageNet, Places), the state of the art in computer vision is advancing rapidly. One important factor for continued progress is to understand the representations that are learned by the inner layers of these deep architectures. Here we show that object detectors emerge from training CNNs to perform scene classification. As scenes are composed of objects, the CNN for scene classification automatically discovers meaningful objects detectors, representative of the learned scene categories. With object detectors emerging as a result of learning to recognize scenes, our work demonstrates that the same network can perform both scene recognition and object localization in a single forward-pass, without ever having been explicitly taught the notion of objects.

Table of Contents

  • 1 Introduction
  • 2 ImageNet-CNN and Places-CNN
  • 3 Uncovering the CNN representation
  • 3.1 Simplifying the input images
  • 3.2 Visualizing the receptive fields of units and their activation patterns
  • 3.3 Identifying the semantics of internal units
  • 4 Emergence of objects as the internal representation
  • 4.1 What object classes emerge?
  • 4.2 Object Localization within the inner Layers
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Emergence of Object Detectors in Scene-Trained Convolutional Networks

    empirical result

    Training a convolutional neural network (CNN) on scene categorization (the Places database with 205 scene categories) without any object-level labels or bounding-box annotations causes units in the higher convolutional layers (specifically conv4 and pool5) to emerge as explicit, localized object detectors. Despite receiving no object-level supervision, the scene-trained network (Places-CNN) discovers more full object-selective units in its late layers than an identical network architecture trained directly on object categorization (ImageNet-CNN), where intermediate units specialize more in object parts than in whole objects.

  2. Knowl 2 — Discriminative Utility Dictates the Emergence of Learned Object Detectors

    empirical result

    The distribution of object detector units that emerge in layer pool5 of a scene-trained CNN (Places-CNN) correlates strongly with how informative those object categories are for discriminating scene classes, rather than merely their raw frequency in natural images.

    When evaluated against 8,220 fully annotated images across 205 scene categories in the SUN database:

    1. The correlation between natural object occurrence frequency (which follows Zipf's law) and the count of pool5 units discovering that object class is r=0.54r = 0.54.
    2. The correlation between the discriminative utility of an object class for scene classification (quantified as the number of scene categories for which that object class is the single most informative predictor) and the count of pool5 units discovering that object class is r=0.84r = 0.84.

    This demonstrates that the unsupervised emergence of internal object representations is primarily driven by task-specific discriminative selection.

  3. Knowl 3 — Data-Driven Estimation of CNN Empirical Receptive Fields

    algorithm
    Input: Unit uu in a convolutional layer, image dataset D\mathcal{D}, number of top images KK, occluder patch size s×ss \times s (set to 11×1111 \times 11), sliding stride rr (set to 3).
    Output: Empirical receptive field spatial activation profile RuR_u.
    1. For each image I∈DI \in \mathcal{D}, compute the network feature activations for unit uu.
    2. Select the top KK images {I1,…,IK}\{I_1, \dots, I_K\} yielding the highest activation values for unit uu.
    3. For each selected image IkI_k (k∈{1,…,K}k \in \{1, \dots, K\}):
        a. Record original maximum activation a0=max⁡x,yu(Ik)x,ya_0 = \max_{x,y} u(I_k)_{x,y} and its spatial coordinate (xk∗,yk∗)(x_k^*, y_k^*).
        b. Generate a dense grid of occluded images by placing random s×ss \times s pixel patches across IkI_k with stride rr (yielding approximately 5000 occluded variants per image).
        c. For each occluder position (i,j)(i, j), evaluate the perturbed image Ik(i,j)I_k^{(i,j)} and record a(i,j)=max⁡x,yu(Ik(i,j))x,ya(i, j) = \max_{x,y} u(I_k^{(i,j)})_{x,y}.
        d. Construct a discrepancy map Dk(i,j)=max⁡(0,a0−a(i,j))D_k(i, j) = \max(0, a_0 - a(i, j)).
        e. Shift and center DkD_k relative to the unit peak activation coordinate (xk∗,yk∗)(x_k^*, y_k^*) in image space to obtain calibrated discrepancy map D~k\tilde{D}_k.
    4. Average the calibrated discrepancy maps across all KK top images: Ru=1K∑k=1KD~kR_u = \frac{1}{K} \sum_{k=1}^K \tilde{D}_k.
    5. Return empirical receptive field RuR_u.

    The algorithm determines the effective spatial region influencing a unit's activation by measuring sensitivity to local image occlusion.

  4. Knowl 4 — Theoretical versus Empirical Receptive Field Sizes Across CNN Layers

    data/table

    Theoretical receptive fields (RF) calculated from convolutional filter sizes, strides, and pooling operations grow rapidly with network depth. However, the empirical RF (the actual image region that modulates a neuron's activation) is substantially more localized. The empirical RF side length (in pixels, assuming a square RF) is measured across layers of an AlexNet architecture trained on Places (Places-CNN) and ImageNet (ImageNet-CNN):

    Layer pool1 pool2 conv3 conv4 pool5
    Theoretical RF size (px) 19 67 99 131 195
    Places-CNN empirical RF (px) 17.8±1.617.8 \pm 1.6 37.4±5.937.4 \pm 5.9 52.1±10.652.1 \pm 10.6 60.0±13.760.0 \pm 13.7 72.0±20.072.0 \pm 20.0
    ImageNet-CNN empirical RF (px) 17.9±1.617.9 \pm 1.6 36.7±5.436.7 \pm 5.4 51.1±9.951.1 \pm 9.9 60.4±16.060.4 \pm 16.0 70.3±21.670.3 \pm 21.6

    In early layers (pool1), empirical RF size matches theoretical calculations (17.817.8 vs 1919 pixels). With increasing depth, the empirical RF expands much slower than theoretical sizes: at pool5, empirical RF side length is approximately 70–7270\text{--}72 pixels compared to the theoretical size of 195195 pixels.

  5. Knowl 5 — Layer-Wise Hierarchy of Semantic Concept Abstraction in CNNs

    empirical result

    Human semantic annotation of units across network layers reveals a systematic progression in the level of visual abstraction from early to late layers:

    • Early layers (pool1, pool2): Units predominantly respond to low-level visual primitives (simple elements, colors, materials, and textures).
    • Intermediate layers (conv3, conv4): Units transition to detecting surfaces and object parts. In ImageNet-CNN, object part detectors peak around layer conv4.
    • Late layers (conv4, pool5): High-level semantic units dominate. In Places-CNN, units are heavily biased toward full objects and scene layouts, whereas ImageNet-CNN maintains a higher fraction of object part detectors.

    When filtering for units where naive annotators reach ≥75%\ge 75\% precision (accounting for ≈60%\approx 60\% of units in each layer), Places-CNN consistently yields a higher proportion of whole-object and scene detectors in layers conv4 and pool5 compared to ImageNet-CNN.

  6. Knowl 6 — Simultaneous Single Forward-Pass Scene Classification and Object Localization

    model/method

    A standard convolutional neural network trained exclusively for scene classification with image-level labels (Places-CNN) performs simultaneous scene classification and semantic object localization in a single forward pass without requiring external region proposals, multi-scale sliding windows, or bounding-box supervision during training.

    During inference:

    1. The input image passes through the network to generate scene category probabilities at the final softmax classification layer.
    2. For each unit in intermediate convolutional layers (e.g., pool5), the unit spatial activation map is thresholded.
    3. Activated spatial regions are mapped to image coordinates using the unit's empirical receptive field, yielding localized bounding boxes or segmentation masks corresponding to the unit's discovered semantic object tag.
  7. Knowl 7 — Distributed Multi-Unit Encoding of Object Appearance Variations

    empirical result

    In the pool5 layer of a scene-classification CNN (Places-CNN, containing 256 units), semantic object categories are represented not by a single unit, but across multiple distinct units specialized to different visual appearances, sub-types, viewpoints, or scales:

    • Buildings are detected by 15 separate units in pool5.
    • People are detected by 9 separate units, each selective to specific scales, poses, or activities.
    • Lamps are represented by 6 distinct units, each tuned to a specific sub-category (such as ceiling lamps, desk lamps, and chandeliers).
    • In ImageNet-CNN pool5, dogs occupy 15 distinct units alongside multiple separate units tuned to dog parts (heads, legs, bodies).
  8. Knowl 8 — Minimal Image Representation via Iterative Diagnostic Region Extraction

    model/method

    The diagnostic elements utilized by a scene-classification network are identified by simplifying an input image while preserving its correct category prediction score. Starting from an image correctly classified by Places-CNN, image segments are removed iteratively in gradient space. At each step, the segment whose deletion causes the minimal decrease in the target category score is removed until the image is incorrectly classified.

    Applying this procedure using ground-truth object annotations from the SUN database demonstrates that scene recognition relies on specific key objects:

    • In bedroom scenes, the minimal representation retains the bed segment in 87%87\% of cases (followed by walls at 28%28\% and windows at 21%21\%).
    • In art gallery scenes, paintings (81%81\%) and pictures (58%58\%) are retained.
    • In amusement park scenes, carousels (75%75\%) and rides (64%64\%) are retained.
    • In bookstore scenes, bookcases (96%96\%) and books (68%68\%) are retained.
  9. Knowl 9 — Experimental Network Architecture and Baseline Classification Benchmark

    experimental setup

    Both ImageNet-CNN and Places-CNN share an 8-layer AlexNet architecture:

    • Layer structure: conv1 (96 units, 55×5555 \times 55), pool1 (96 units, 27×2727 \times 27), conv2 (256 units, 27×2727 \times 27), pool2 (256 units, 13×1313 \times 13), conv3 (384 units, 13×1313 \times 13), conv4 (384 units, 13×1313 \times 13), conv5 (256 units, 13×1313 \times 13), pool5 (256 units, 6×66 \times 6), fc6 (4096 units), fc7 (4096 units).
    • ImageNet-CNN: Trained from scratch on 1.3 million images from 1,000 object categories of ILSVRC 2012, achieving top-1 accuracy of 57.4%57.4\%.
    • Places-CNN: Trained from scratch on 2.4 million images from 205 scene categories of the Places Database, achieving top-1 accuracy of 50.0%50.0\%.
    • Scene classification transfer: When tested on the 205 scene test set, Places-CNN achieves 50.0%50.0\%, whereas features from ImageNet-CNN paired with a linear SVM achieve 40.8%40.8\%.

Coverage note — No substantial contributed material was omitted. All key models, algorithms, empirical results, and analytical experiments from the paper are represented.

References

  1. 1.Agrawal, Pulkit, Girshick, Ross, and Malik, Jitendra. Analyzing the performance of multilayer neural networks for object recognition. ECCV, 2014.
  2. 2.Bergamo, Alessandro, Bazzani, Loris, Anguelov, Dragomir, and Torresani, Lorenzo. Self-taught object localization with deep networks. arXiv preprint arXiv:1409.3964, 2014.
  3. 3.Biederman, Irving. Visual object recognition, volume 2. MIT press Cambridge, 1995.
  4. 4.Deng, Jia, Dong, Wei, Socher, Richard, Li, Li-Jia, Li, Kai, and Fei-Fei, Li. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  5. 5.Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., and Zisserman, A. The pascal visual object classes challenge. IJCV, 2010.
  6. 6.Farabet, Clement, Couprie, Camille, Najman, Laurent, and LeCun, Yann. Learning hierarchical features for scene labeling. TPAMI, 2013.
  7. 7.Girshick, R., Donahue, J., Darrell, T., and Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. 2014.
  8. 8.Grangier, D., Bottou, L., and Collobert, R. Deep convolutional networks for scene parsing. TPAMI, 2009.
  9. 9.Jia, Yangqing. Caffe: An open source convolutional architecture for fast feature embedding, 2013.
  10. 10.Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  11. 11.Lazebnik, Svetlana, Schmid, Cordelia, and Ponce, Jean. Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. In CVPR, 2006.
  12. 12.Le, Quoc V. Building high-level features using large scale unsupervised learning. In ICASSP, 2013.
  13. 13.Li, Li-Jia, Su, Hao, Fei-Fei, Li, and Xing, Eric P. Object bank: A high-level image representation for scene classification & semantic feature sparsification. In NIPS, pp. 1378–1386, 2010.
  14. 14.Lin, Tg-Yi, Maire, Michael, Belongie, Serge, Hays, James, Perona, Pietro, Ramanan, Deva, Dollr, Piotr, and Zitnick, C. Lawrence. Microsoft COCO: Common objects in context. In ECCV, 2014.
  15. 15.Long, Jonathan, Zhang, Ning, and Darrell, Trevor. Do convnets learn correspondence? In NIPS, 2014.
  16. 16.Oliva, A. and Torralba, A. Building the gist of a scene: The role of global image features in recognition. Progress in Brain Research, 2006.
  17. 17.Oquab, Maxime, Bottou, Leon, Laptev, Ivan, Sivic, Josef, et al. Weakly supervised object recognition with convolutional neural networks. In NIPS. 2014.
  18. 18.Pandey, M. and Lazebnik, S. Scene recognition and weakly supervised object localization with deformable part-based models. In ICCV, 2011.
  19. 19.Perez, Patrick, Gangnet, Michel, and Blake, Andrew. Poisson image editing. ACM Trans. Graph., 2003.
  20. 20.Renninger, Laura Walker and Malik, Jitendra. When is scene identification just texture recognition? Vision research, 44(19):2301–2311, 2004.
  21. 21.Russakovsky, Olga, Deng, Jia, Su, Hao, Krause, Jonathan, Satheesh, Sanjeev, Ma, Sean, Huang, Zhiheng, Karpathy, Andrej, Khosla, Aditya, Bernstein, Michael, Berg, Alexander C., and Fei-Fei, Li. ImageNet Large Scale Visual Recognition Challenge, 2014.
  22. 22.Szegedy, Christian, Zaremba, Wojciech, Sutskever, Ilya, Bruna, Joan, Erhan, Dumitru, Goodfellow, Ian, and Fergus, Rob. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013.
  23. 23.Tanaka, Keiji. Neuronal mechanisms of object recognition. Science, 262(5134):685–688, 1993.
  24. 24.Tang, Yichuan, Srivastava, Nitish, and Salakhutdinov, Ruslan R. Learning generative models with visual attention. In NIPS. 2014.
  25. 25.Xiao, J, Ehinger, K A., Hays, J, Torralba, A, and Oliva, A. SUN database: Exploring a large collection of scene categories. IJCV, 2014.
  26. 26.Yosinski, Jason, Clune, Jeff, Bengio, Yoshua, and Lipson, Hod. How transferable are features in deep neural networks? In NIPS, 2014.
  27. 27.Zeiler, M. and Fergus, R. Visualizing and understanding convolutional networks. In ECCV, 2014.
  28. 28.Zhou, Bolei, Lapedriza, Agata, Xiao, Jianxiong, Torralba, Antonio, and Oliva, Aude. Learning deep features for scene recognition using places database. In NIPS, 2014.

Citation

MLA
Zhou, B., et al. “Object Detectors Emerge in Deep Scene CNNs”. arXiv, 2014, http://arxiv.org/abs/1412.6856v2.
APA
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., & Torralba, A. (2014). Object Detectors Emerge in Deep Scene CNNs. arXiv. http://arxiv.org/abs/1412.6856v2
Chicago
Zhou, B., A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba. 2014. “Object Detectors Emerge in Deep Scene CNNs”. arXiv. http://arxiv.org/abs/1412.6856v2.
Harvard
Zhou, B. et al. (2014) “Object Detectors Emerge in Deep Scene CNNs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1412.6856v2.
Vancouver
1. Zhou B, Khosla A, Lapedriza A, Oliva A, Torralba A (2014) Object Detectors Emerge in Deep Scene CNNs. arXiv

BibTeX

@article{zhou2014object,
  title = {Object Detectors Emerge in Deep Scene CNNs},
  author = {Zhou, Bolei and Khosla, Aditya and Lapedriza, Agata and Oliva, Aude and Torralba, Antonio},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1412.6856v2},
  eprint = {1412.6856}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors