Convolutional neural networks at constrained time cost

Kaiming HeJian Sun

article2014CVPR1,394 citations

Demonstrates how to optimize convolutional neural network design trade-offs across depth, filter count, and filter size to significantly improve ImageNet accuracy while running 20% faster than AlexNet under a strict computation budget.

Listen

Modern computer vision systems rely heavily on deep convolutional neural networks, but recent performance gains have come at the cost of significantly greater model complexity and runtime. In production environments—such as real-time search engines, cloud platforms processing thousands of images per second, or mobile devices with limited computing power—unconstrained models are often impractical or prohibitively expensive. In addition, training large models can require weeks of compute cluster time, creating severe bottlenecks for product development.

The article investigates how to design vision models that maximize recognition accuracy while strictly adhering to a fixed computation time budget. Specifically, it demonstrates how systematically trading off structural network parameters—such as depth, layer width, and filter size—allows models to achieve superior image classification accuracy without increasing computational costs.

To conduct this evaluation, the authors performed controlled empirical comparisons on the standard 1,000-category ImageNet dataset using a single graphics processing unit. Starting from an efficient eight-layer baseline model, the authors applied a "layer replacement" strategy. Under this approach, individual layers were progressively modified or substituted while keeping the theoretical convolutional time complexity constant, allowing the direct isolation and measurement of specific architectural trade-offs.

The study yielded several critical findings regarding efficient model design. First, increasing network depth takes clear priority over layer width and filter size; replacing larger filters with sequences of smaller filters (such as 2x2 filters) allowed the network to grow deeper, markedly cutting error rates at the same computational budget. Second, depth cannot be increased indefinitely, as adding excessive layers eventually degraded training and validation accuracy even when complexity was not constrained. Third, delaying the spatial downsampling in pooling layers—by setting pooling stride to 1 and moving the step size to the following convolutional layer—consistently improved accuracy at no extra computational cost. Finally, the resulting optimized model achieved an 11.8% top-5 error rate on ImageNet, making it 20% faster in actual runtime than the standard AlexNet architecture while reducing top-5 error by 4.2 percentage points and computational complexity by 40%.

These findings indicate that organizations do not need to choose between rapid processing speed and competitive accuracy. By strategically prioritizing depth and smaller filter dimensions rather than wider layers or large filters, engineering teams can deploy models that run within strict latency and cost budgets while outperforming previous industry baselines. The results also show that highly complex multi-path designs or massive 16-layer architectures, which require up to 23 times more runtime, may be unnecessary for systems constrained by practical operating budgets.

Engineering and infrastructure teams should adopt these layer-replacement and delayed subsampling techniques when designing or upgrading vision pipelines subject to fixed latency limits. Organizations should also evaluate test-time acceleration methods on top of these efficient architectures to gain additional speedups. For further work, teams should investigate memory usage optimizations, as the proposed deeper models consume more operating memory during training, which warrants future analysis for environments facing tight memory ceilings.

arXiv: 1412.1710
  • Paper: ImageNet Classification with Deep Convolutional Neural Networks, Alex Krizhevsky et al. (2012). It introduces the foundational AlexNet architecture that serves as the explicit accuracy and time-complexity baseline modified by the source paper.
  • Paper: Visualizing and Understanding Convolutional Networks, Matthew D. Zeiler et al. (2014). It provides crucial insights into how adjusting filter sizes and layer strides alters convolutional representations, directly informing the structural trade-off experiments conducted in the source.
  • Paper: Network In Network, Min Lin et al. (2014). It introduces micro-networks and 1x1 convolutions within CNN layers, establishing a design principle utilized in time-constrained architecture modifications.
Cover for Convolutional neural networks at constrained time cost

Abstract

Though recent advanced convolutional neural networks (CNNs) have been improving the image recognition accuracy, the models are getting more complex and time-consuming. For real-world applications in industrial and commercial scenarios, engineers and developers are often faced with the requirement of constrained time budget. In this paper, we investigate the accuracy of CNNs under constrained time cost. Under this constraint, the designs of the network architectures should exhibit as trade-offs among the factors like depth, numbers of filters, filter sizes, etc. With a series of controlled comparisons, we progressively modify a baseline model while preserving its time complexity. This is also helpful for understanding the importance of the factors in network designs. We present an architecture that achieves very competitive accuracy in the ImageNet dataset (11.8% top-5 error, 10-view test), yet is 20% faster than "AlexNet" (16.0% top-5 error, 10-view test).

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Prerequisites
  • 3.1 A Baseline Model
  • 3.2 Time Complexity of Convolutions
  • 4 Model Designs by Layer Replacement
  • 4.1 Trade-offs between Depth and Filter Sizes
  • 4.2 Trade-offs between Depth and Width
  • 4.3 Trade-offs between Width and Filter Sizes
  • 4.4 Is Deeper Always Better?
  • 4.5 Adding a Pooling Layer
  • 4.6 Delayed Subsampling of Pooling Layers
  • 4.7 Summary
  • 5 Implementation Details
  • 6 Comparisons
  • 7 Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — Theoretical Time Complexity of Convolutional Layers

    equation

    The theoretical time complexity of all convolutional layers in a deep convolutional neural network is quantified as:

    O(∑l=1dnl−1⋅sl2⋅nl⋅ml2)\mathcal{O}\left( \sum_{l=1}^d n_{l-1} \cdot s_l^2 \cdot n_l \cdot m_l^2 \right)

    where:

    • l∈{1,…,d}l \in \{1, \dots, d\} is the index of a convolutional layer, and dd is the total number of convolutional layers (network depth).
    • nln_l is the number of filters (width) in the ll-th layer, with nl−1n_{l-1} representing the number of input channels to the ll-th layer.
    • sls_l is the spatial length (height and width) of the square filter in the ll-th layer.
    • mlm_l is the spatial length of the output feature map of the ll-th layer.

    This complexity metric governs both feedforward testing time and training time per image, with training time requiring roughly three times the computation of testing time (one forward pass and two backward passes). Fully connected and pooling layers are excluded from this calculation as they generally account for only 5% to 10% of total computational time.

  2. Knowl 2 — Stage-Wise Layer Replacement under Constant Complexity

    model/method

    To systematically explore network architecture variations under a fixed computational time budget, networks are decomposed into modular "stages"—subnetworks situated between consecutive pooling layers. Modifications are executed via "layer replacement": substituting a group of convolutional layers within a stage with an alternative set of layers that preserves the stage's theoretical computational time complexity:

    O(nl−1⋅sl2⋅nl⋅ml2)\mathcal{O}\left(n_{l-1} \cdot s_l^2 \cdot n_l \cdot m_l^2\right)

    To isolate architectural trade-offs and prevent cascading computational changes across the network, the number of input channels feeding into the stage and the number of output filters exiting the stage are held strictly constant. By varying only two architectural factors at a time within a stage (such as depth vs. filter size, or depth vs. width) while keeping feature map dimensions and boundary interfaces fixed, the relative significance of individual structural hyperparameters can be directly isolated.

  3. Knowl 3 — Priority of Depth over Spatial Filter Size under Constant Time Cost

    empirical result

    Under a fixed theoretical computational time budget, replacing larger convolutional filters with cascades of smaller filters to increase network depth consistently improves classification accuracy on ImageNet 2012 (10-view evaluation, 75 training epochs):

    • Replacing one 3×33\times 3 layer (256 channels) with two 2×22\times 2 layers preserves time complexity (22+22≈322^2 + 2^2 \approx 3^2, relative complexity 8/9) and reduces validation top-5 error from 15.9% (Model A, 5 conv layers) to 14.9% (Model B, 8 conv layers).
    • Replacing a 5×55\times 5 layer (64 input, 128 output channels) with two 3×33\times 3 layers reduces top-5 error from 15.9% (Model A) to 14.3% (Model C, 6 conv layers) and from 14.9% (Model B) to 13.9% (Model D, 9 conv layers).
    • Further replacing the 5×55\times 5 layer with four 2×22\times 2 layers reduces top-5 error to 13.3% (Model E, 11 conv layers).

    When deploying 2×22\times 2 filters in sequence, spatial feature map size is preserved by applying zero padding in the first 2×22\times 2 layer and 1-pixel padding on each side in the second 2×22\times 2 layer. These results establish that network depth is strictly higher priority than spatial filter size.

  4. Knowl 4 — Priority of Depth over Channel Width under Constant Time Cost

    empirical result

    Under a fixed computational time complexity budget, increasing network depth while proportionally decreasing the number of filters per layer (channel width) provides substantial accuracy improvements on ImageNet 2012 (10-view validation set, 75 epochs):

    • Replacing three 3×33\times 3 layers with 256 filters in the third stage of a 5-layer baseline (Model A, 15.9% top-5 error) with six 3×33\times 3 layers of width 160 (Model F, depth 8) reduces top-5 error to 14.8% at identical theoretical complexity (128⋅32⋅256+2×256⋅32⋅256=128⋅32⋅160+4×160⋅32⋅160+160⋅32⋅256128 \cdot 3^2 \cdot 256 + 2 \times 256 \cdot 3^2 \cdot 256 = 128 \cdot 3^2 \cdot 160 + 4 \times 160 \cdot 3^2 \cdot 160 + 160 \cdot 3^2 \cdot 256).
    • Further increasing depth to nine 3×33\times 3 layers of width 128 (Model G, depth 11) yields 14.7% top-5 error.
    • In stage 2, replacing two 3×33\times 3 layers with 128 filters by four 3×33\times 3 layers of width 64 reduces top-5 error from 14.3% (Model C) to 14.0% (Model H) and from 13.9% (Model D) to 13.5% (Model I).

    While depth delivers significant gains over width, the marginal improvement saturates as the network becomes extremely deep and narrow.

  5. Knowl 5 — Neutral Trade-off between Width and Filter Size at Fixed Depth

    empirical result

    When network depth is held constant, trading spatial filter size against channel width (e.g., using larger filters with fewer channels versus smaller filters with more channels) results in comparable recognition performance on ImageNet 2012:

    • Comparing Model B (six 2×22\times 2 layers with 256 filters in stage 3, depth 8) and Model F (six 3×33\times 3 layers with 160 filters in stage 3, depth 8) yields nearly identical validation top-1/top-5 errors: 35.7%/14.9% (Model B) vs. 35.5%/14.8% (Model F).
    • Comparing Model E (four 2×22\times 2 layers with 128 filters in stage 2, depth 11) and Model I (four 3×33\times 3 layers with 64 filters in stage 2, depth 11) yields top-1/top-5 errors of 33.8%/13.3% (Model E) vs. 33.9%/13.5% (Model I).

    Unlike depth, neither channel width nor spatial filter size (3×33\times 3 vs. 2×22\times 2) demonstrates an intrinsic advantage over the other when computational time cost and depth are matched.

  6. Knowl 6 — Optimization Degradation from Excessive Depth in Plain Feedforward CNNs

    empirical result

    Solely increasing the depth of a feedforward convolutional neural network by adding identical convolutional layers without architectural shortcuts eventually leads to performance saturation and degradation, even when width and filter sizes are not traded off (increasing time cost):

    Model D D+2 D+4 D+6 D+8
    Top-1 Error (%) 34.5 34.0 33.9 34.0 34.2
    Top-5 Error (%) 13.9 13.6 13.4 13.5 13.6

    In this experiment, ii extra convolutional layers (2×22\times 2 filters, 256 channels) are appended to the final stage of Model D (9 convolutional layers). Accuracy peaks at D+4 (depth 13) and degrades at D+6 and D+8. The performance degradation is observed on both the validation set and the training set, demonstrating that the failure is caused by optimization difficulty rather than overfitting.

    Similarly, adding 1×11\times 1 convolutional layers (Network-in-Network style) after each convolutional layer increases depth and degrades performance: Model B top-1/top-5 error degrades from 35.7%/14.9% to 37.8%/16.5%, and Model D degrades from 34.5%/13.9% to 36.9%/15.7%.

  7. Knowl 7 — Channel Expansion via Spatial Downsampling (Stage 4 Addition)

    model/method

    Because the computational cost of a convolutional layer scales with the square of the feature map spatial dimension (ml2m_l^2), introducing a spatial downsampling pooling layer enables a massive increase in channel width within the newly formed stage while maintaining fixed total theoretical complexity.

    Applying a max-pooling layer with spatial filter size 3 and stride 3 after stage 3 reduces the feature map spatial dimension from 18×1818\times 18 to 6×66\times 6. Two 2×22\times 2 convolutional layers are shifted into this new fourth stage. The reduction in spatial dimensions by a factor of 32=93^2 = 9 permits expanding the channel width of the first stage 4 layer from 256 to 2304 filters at identical computation:

    256⋅22⋅256+256⋅22⋅256=256⋅22⋅2304+2304⋅22⋅25632256 \cdot 2^2 \cdot 256 + 256 \cdot 2^2 \cdot 256 = \frac{256 \cdot 2^2 \cdot 2304 + 2304 \cdot 2^2 \cdot 256}{3^2}

    Applying this transformation to Model E yields Model J, reducing ImageNet 2012 validation top-1/top-5 error from 33.8%/13.3% to 32.9%/12.5% under the same computational time complexity.

  8. Knowl 8 — Delayed Subsampling of Pooling Layers

    model/method

    Standard max-pooling layers simultaneously execute two operations: lateral suppression (max-filtering to achieve local spatial translation invariance) and spatial subsampling (dimensionality reduction via stride s>1s > 1). "Delayed subsampling" separates these two operations:

    1. Set the stride of the max-pooling layer to 1 (performing pure sliding-window lateral suppression without reducing spatial size).
    2. Apply the subsampling stride (stride s>1s > 1, equal to the original pooling stride) to the immediately succeeding convolutional layer.

    This modification does not alter the theoretical complexity of any convolutional layers because the input feature map size of the subsequent convolutional layer is reduced according to its stride. Across various architectures on ImageNet 2012 (75 epochs), delayed subsampling consistently lowers top-5 validation error:

    Delayed Subsampling Model B Model D Model E Model J
    No 14.9% 13.9% 13.3% 12.5%
    Yes 13.9% 13.5% 13.0% 12.0%
  9. Knowl 9 — Model J' Architecture Specification and Training Protocol

    model/method

    Model J' is an 11-convolutional-layer architecture optimized for constrained time cost:

    • Input: 224×224×3224\times 224\times 3 RGB image (mean-subtracted).
    • Stage 1: Conv (7×7,64 filters)(7\times 7, 64\text{ filters}), stride 2, no padding →\to MaxPool 3×33\times 3, stride 1 (output size dominantly 36×3636\times 36).
    • Stage 2: Conv (2×2,128 filters)(2\times 2, 128\text{ filters}), stride 3, no padding →\to 3 sequential Conv (2×2,128 filters)(2\times 2, 128\text{ filters}), stride 1, with 1-pixel padding on alternate layers →\to MaxPool 2×22\times 2, stride 1 (output size dominantly 18×1818\times 18).
    • Stage 3: Conv (2×2,256 filters)(2\times 2, 256\text{ filters}), stride 2, no padding →\to 3 sequential Conv (2×2,256 filters)(2\times 2, 256\text{ filters}), stride 1, with 1-pixel padding on alternate layers →\to MaxPool 3×33\times 3, stride 1 (output size dominantly 6×66\times 6).
    • Stage 4: Conv (2×2,2304 filters)(2\times 2, 2304\text{ filters}), stride 3, no padding →\to Conv (2×2,256 filters)(2\times 2, 256\text{ filters}), stride 1, 1-pixel padding.
    • SPP Layer: 4-level Spatial Pyramid Pooling with spatial bins {6×6,3×3,2×2,1×1}\{6\times 6, 3\times 3, 2\times 2, 1\times 1\} (totaling 50 bins ×256=12,800\times 256 = 12{,}800 dimensions).
    • Classifier: Two 4096-dimensional fully connected layers (with ReLU and 50% dropout) →\to 1000-way softmax layer.

    Training Protocol: Mini-batch size 128; SGD with momentum 0.9 and weight decay 0.0005; Gaussian weight initialization; initial learning rate 0.01 (epochs 1–10), 0.001 (epochs 11–70), 0.0001 (remaining epochs up to 75 or 90). Data augmentation uses 224×224224\times 224 crops from images with shorter side 256, random horizontal flipping, and RGB color altering.

  10. Knowl 10 — ImageNet Performance and Speed Benchmarks of Model J'

    data/table

    When trained for 90 epochs on ImageNet 2012 and evaluated using 10-view testing, Model J' outperforms standard fast models in accuracy while maintaining lower complexity and faster single-GPU runtime:

    Model Top-1 Error (%) Top-5 Error (%) Conv Complexity Seconds / Mini-batch (128)
    AlexNet (re-impl.) 37.6 16.0 1.4×1.4\times 0.50 (1.2×1.2\times)
    ZF (fast, re-impl.) 36.0 14.8 1.5×1.5\times 0.54 (1.3×1.3\times)
    CNN-F – 16.7 0.9×0.9\times 0.30 (0.7×0.7\times)
    SPPnet (ZF5) 35.0 14.1 1.5×1.5\times 0.55 (1.3×1.3\times)
    CNN-M – 13.7 2.1×2.1\times 0.71 (1.7×1.7\times)
    CNN-S – 13.1 3.8×3.8\times 1.23 (3.0×3.0\times)
    SPPnet (O5) 32.9 12.8 3.8×3.8\times 1.24 (3.0×3.0\times)
    SPPnet (O7) 30.4 11.1 5.8×5.8\times 1.85 (4.5×4.5\times)
    Model J' (ours) 31.8 11.8 1.0×\mathbf{1.0\times} 0.41 (1.0×\mathbf{1.0\times})

    Compared to AlexNet, Model J' reduces top-5 error by 4.2% (from 16.0% to 11.8%) and top-1 error by 5.8% (from 37.6% to 31.8%), while possessing 40% less theoretical convolutional complexity and running 20% faster on an Nvidia Titan GPU (0.41 seconds per 128-image batch, 1.0 ms test time per view). It also achieves superior accuracy to CNN-M, CNN-S, and SPPnet (O5) at lower computational cost.

Coverage note — None was omitted; all major architectural trade-offs, principles, layer configurations, and benchmark evaluations are represented.

References

  1. 1.K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman. Return of the devil in the details: Delving deep into convolutional nets. In BMVC, 2014.
  2. 2.D. Ciresan, U. Meier, and J. Schmidhuber. Multi-column deep neural networks for image classification. In CVPR, 2012.
  3. 3.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  4. 4.E. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus. Exploiting linear structure within convolutional networks for efficient evaluation. In NIPS, 2014.
  5. 5.D. Eigen, J. Rolfe, R. Fergus, and Y. LeCun. Understanding deep architectures using a recursive convolutional network. arXiv:1312.1847, 2013.
  6. 6.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014.
  7. 7.X. Glorot and Y. Bengio. Understanding the difficulty of training deep feedforward neural networks. In International Conference on Artificial Intelligence and Statistics, pages 249–256, 2010.
  8. 8.B. Hariharan, P. Arbelaez, R. Girshick, and J. Malik. Simultaneous detection and segmentation. In ECCV, pages 297–312, 2014.
  9. 9.K. He, X. Zhang, S. Ren, and J. Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. arXiv:1406.4729v2, 2014.
  10. 10.A. G. Howard. Some improvements on deep convolutional neural network based image classification. arXiv:1312.5402, 2013.
  11. 11.M. Jaderberg, A. Vedaldi, and A. Zisserman. Speeding up convolutional neural networks with low rank expansions. In BMVC, 2014.
  12. 12.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv:1408.5093, 2014.
  13. 13.A. Krizhevsky. One weird trick for parallelizing convolutional neural networks. arXiv:1404.5997, 2014.
  14. 14.A. Krizhevsky, I. Sutskever, and G. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  15. 15.Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural computation, 1989.
  16. 16.M. Lin, Q. Chen, and S. Yan. Network in network. arXiv:1312.4400, 2013.
  17. 17.F. Mamalet and C. Garcia. Simplifying convnets for fast learning. In Artificial Neural Networks and Machine Learning - ICANN, pages 58–65, 2012.
  18. 18.V. Nair and G. E. Hinton. Rectified linear units improve restricted boltzmann machines. In ICML, pages 807–814, 2010.
  19. 19.A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson. Cnn features off-the-shelf: An astounding baseline for recogniton. In CVPR 2014, DeepVision Workshop, 2014.
  20. 20.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Imagenet large scale visual recognition challenge. arXiv:1409.0575, 2014.
  21. 21.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. 2014.
  22. 22.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556, 2014.
  23. 23.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. arXiv:1409.4842, 2014.
  24. 24.M. D. Zeiler and R. Fergus. Visualizing and understanding convolutional neural networks. In ECCV, 2014.

Citation

MLA
He, K., and J. Sun. “Convolutional Neural Networks at Constrained Time Cost”. arXiv, 2014, http://arxiv.org/abs/1412.1710v1.
APA
He, K., & Sun, J. (2014). Convolutional Neural Networks at Constrained Time Cost. arXiv. http://arxiv.org/abs/1412.1710v1
Chicago
He, K., and J. Sun. 2014. “Convolutional Neural Networks at Constrained Time Cost”. arXiv. http://arxiv.org/abs/1412.1710v1.
Harvard
He, K. and Sun, J. (2014) “Convolutional Neural Networks at Constrained Time Cost”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1412.1710v1.
Vancouver
1. He K, Sun J (2014) Convolutional Neural Networks at Constrained Time Cost. arXiv

BibTeX

@article{he2014convolutional,
  title = {Convolutional Neural Networks at Constrained Time Cost},
  author = {He, Kaiming and Sun, Jian},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1412.1710v1},
  eprint = {1412.1710}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE