Speeding up Convolutional Neural Networks with Low Rank Expansions

Max JaderbergAndrea VedaldiAndrew Zisserman

article2014BMVC1,572 citations

Presents low-rank filter decomposition techniques that accelerate convolutional neural network computation by up to 4.5x on standard hardware with negligible loss in accuracy.

Listen

Convolutional neural networks have set high performance benchmarks across computer vision and machine learning tasks. However, their heavy computational requirements pose significant challenges for real-world and real-time deployment, especially in sliding-window object detection systems. Because stacked convolutional layers account for the vast majority of processing time, improving the efficiency of these operations is vital for practical deployment.

The article demonstrates that convolutional layers contain substantial redundancy across spatial dimensions and feature channels. Its primary objective is to evaluate two hardware-agnostic approximation schemes that decompose standard filters into low-rank, separable components, enabling significant inference speedups on pre-trained models with minimal impact on accuracy.

To evaluate these techniques, the authors applied both schemes to a pre-trained network designed for scene text character recognition. The models were evaluated using standard benchmarks comprising over 5,000 cropped character images, alongside training datasets exceeding 160,000 samples. The team tested two optimization approaches: directly minimizing filter reconstruction error and optimizing approximations based on output data reconstruction across layers.

The findings show that substantial computational gains are achievable with negligible accuracy loss. Using the channel-factored approach combined with data reconstruction optimization, the network achieved a 2.5-fold speedup with no loss in classification accuracy, and a 4.5-fold speedup with less than a one percent drop in accuracy (reducing from 91.3% to 90.3%, which remained state-of-the-art). In layer-specific tests, the targeted intermediate layers accounted for roughly 90% of total run time, making them prime candidates for compression. Additionally, practical tests revealed that the second scheme, which factors layers into sequential vertical and horizontal operations, aligned much better with existing computational libraries, delivering faster real-world execution than the per-channel basis alternative.

These results demonstrate that organizations can drastically reduce compute costs, latency, and hardware constraints for computer vision systems without re-architecting baseline models from scratch. Because the approach is flexible and tunable, engineering teams can adjust the trade-off between speed and accuracy to meet specific application latency budgets and deployment environments.

For practical implementation, engineering teams should prioritize the factored sequential scheme (Scheme 2) paired with data reconstruction optimization when accelerating existing pre-trained convolutional layers. Future efforts should explore learning low-rank, separable filter architectures directly during initial training and investigating varied layer configurations to further optimize efficiency.

Confidence in these findings is high for intermediate convolutional layers within standard computer vision architectures. However, decision-makers should note that the first convolutional layer directly processing raw pixels does not lend itself well to separable approximation due to lack of redundancy. Real-world speed gains also depend closely on how underlying software frameworks handle specific matrix operations, meaning realized speedups may vary across different hardware and software environments.

arXiv: 1405.3866
  • Paper: Network In Network, Min Lin et al. (2014). Introduces 1x1 convolutions (cascaded cross-channel pooling) to manipulate channel dimensions and structure cross-channel feature redundancy, a core foundation exploited by low-rank expansion methods.
  • Paper: Return of the Devil in the Details: Delving Deep into Convolutional Nets, Ken Chatfield et al. (2014). Examines standard deep convolutional network architectures and practical implementation trade-offs that motivate acceleration techniques for convolutional layers.
Cover for Speeding up Convolutional Neural Networks with Low Rank Expansions

Abstract

The focus of this paper is speeding up the evaluation of convolutional neural networks. While delivering impressive results across a range of computer vision and machine learning tasks, these networks are computationally demanding, limiting their deployability. Convolutional layers generally consume the bulk of the processing time, and so in this work we present two simple schemes for drastically speeding up these layers. This is achieved by exploiting cross-channel or filter redundancy to construct a low rank basis of filters that are rank-1 in the spatial domain. Our methods are architecture agnostic, and can be easily applied to existing CPU and GPU convolutional frameworks for tuneable speedup performance. We demonstrate this with a real world network designed for scene text character recognition, showing a possible 2.5x speedup with no loss in accuracy, and 4.5x speedup with less than 1% drop in accuracy, still achieving state-of-the-art on standard benchmarks.

Table of Contents

  • 1 Introduction
  • 2 Filter Approximations
  • 2.1 Approximating Convolutional Neural Network Filter Banks
  • 2.2 Optimization
  • 2.2.1 Filter Reconstruction Optimization
  • 2.2.2 Data Reconstruction Optimization
  • 3 Experiments & Results
  • 4 Conclusions
  • References

Knowls

  1. Knowl 1 — Scheme 2: Sequential Vertical and Horizontal 3D Filter Decomposition

    model/method

    Scheme 2 accelerates a convolutional layer by factoring each d×d×Cd \times d \times C convolutional filter into a sequence of two consecutive convolutional layers using 1D rectangular spatial filters across all channels.

    Let z∈RH×W×Cz \in \mathbb{R}^{H \times W \times C} be the input feature map with height HH, width WW, and CC channels. Let Wn∈Rd×d×CW_n \in \mathbb{R}^{d \times d \times C} denote the nn-th full-rank filter kernel of the layer, for n∈{1,…,N}n \in \{1, \dots, N\}. Scheme 2 decomposes the layer into:

    1. A first convolutional layer of KK vertical 3D filters vk∈Rd×1×Cv_k \in \mathbb{R}^{d \times 1 \times C} (k∈{1,…,K}k \in \{1, \dots, K\}), producing an intermediate feature map V∈RH′×W×KV \in \mathbb{R}^{H' \times W \times K} with channels Vk=∑c=1Cvkc∗zcV^k = \sum_{c=1}^C v_k^c * z^c.
    2. A second convolutional layer of NN horizontal 3D filters hn∈R1×d×Kh_n \in \mathbb{R}^{1 \times d \times K} (n∈{1,…,N}n \in \{1, \dots, N\}), producing the output feature map channels yn∈RH′×W′y_n \in \mathbb{R}^{H' \times W'}.

    The full convolution Wn∗z=∑c=1CWnc∗zcW_n * z = \sum_{c=1}^C W_n^c * z^c is approximated as:

    Wn∗z≈hn∗V=∑k=1Khnk∗(vk∗z)=∑k=1Khnk∗(∑c=1Cvkc∗zc)=∑c=1C[∑k=1Khnk∗vkc]∗zcW_n * z \approx h_n * V = \sum_{k=1}^K h_n^k * (v_k * z) = \sum_{k=1}^K h_n^k * \left( \sum_{c=1}^C v_k^c * z^c \right) = \sum_{c=1}^C \left[ \sum_{k=1}^K h_n^k * v_k^c \right] * z^c

    where ∗* denotes 2D spatial convolution and vkc∈Rd×1v_k^c \in \mathbb{R}^{d \times 1}, hnk∈R1×dh_n^k \in \mathbb{R}^{1 \times d}.

    Whereas standard convolution requires O(NCd2H′W′)\mathcal{O}(N C d^2 H' W') operations, Scheme 2 requires O(KCdH′W)\mathcal{O}(K C d H' W) operations for the vertical stage and O(NKdH′W′)\mathcal{O}(N K d H' W') for the horizontal stage. Assuming spatial input and output widths satisfy W≈W′W \approx W' (valid when W≫dW \gg d), the total computational complexity is O(K(N+C)dH′W′)\mathcal{O}(K (N + C) d H' W'). The scheme provides computational acceleration whenever K(N+C)≪NCdK(N + C) \ll N C d. When K,N,CK, N, C are of the same order of magnitude, the theoretical speedup is approximately d×d\times.

  2. Knowl 2 — Scheme 1: Channel-Wise Separable Basis Expansion

    model/method

    Scheme 1 approximates the 2D filter slices within a convolutional layer by projecting them onto a shared basis of spatially rank-1 (separable) 2D filters and linearly recombining the filtered outputs.

    Let an input feature map be z∈RH×W×Cz \in \mathbb{R}^{H \times W \times C} and let the layer have NN filters Wn∈Rd×d×CW_n \in \mathbb{R}^{d \times d \times C}, where Wnc∈Rd×dW_n^c \in \mathbb{R}^{d \times d} is the 2D filter slice for input channel c∈{1,…,C}c \in \{1, \dots, C\}. The set of NN filters acting on channel cc is approximated as a linear combination of a shared basis of M<NM < N separable 2D filters S={sm∈Rd×d:m∈{1,…,M}}S = \{s_m \in \mathbb{R}^{d \times d} : m \in \{1, \dots, M\}\}, where each sm=vm∗hms_m = v_m * h_m is rank-1 with vertical component vm∈Rd×1v_m \in \mathbb{R}^{d \times 1} and horizontal component hm∈R1×dh_m \in \mathbb{R}^{1 \times d}.

    The convolution is approximated as:

    Wn∗z=∑c=1CWnc∗zc≈∑c=1C∑m=1Mancm(sm∗zc)W_n * z = \sum_{c=1}^C W_n^c * z^c \approx \sum_{c=1}^C \sum_{m=1}^M a_n^{cm} (s_m * z^c)

    where ancm∈Ra_n^{cm} \in \mathbb{R} are scalar linear reconstruction coefficients.

    Direct computation costs O(NCd2H′W′)\mathcal{O}(N C d^2 H' W'), where H′,W′H', W' are output spatial dimensions. Computing MM separable convolutions per input channel costs O(MC2dH′W′)\mathcal{O}(M C 2d H' W'), and recombining the responses across the NN output channels costs O(MCNH′W′)\mathcal{O}(M C N H' W'). The total computational complexity is O(MC(d+N)H′W′)\mathcal{O}(M C (d + N) H' W'), yielding a speedup over direct convolution when M≪dmin⁡{d,N}M \ll d \min\{d, N\}.

  3. Knowl 3 — Data Reconstruction Optimization for Low-Rank CNN Approximation

    model/method

    Data reconstruction optimization determines the low-rank filter parameters of an approximated convolutional layer by minimizing the empirical L2L_2 error between the output feature maps produced by the original layer and those produced by the approximated layer over a dataset of training inputs.

    For Scheme 2 at convolutional layer index ll, the objective is:

    min⁡{hnk},{vkc}∑i=1∣X∣∑n=1N∥Wn∗Φl−1(xi)−∑c=1C∑k=1Khnk∗vkc∗Φl−1(xi)∥22\min_{\{h_n^k\}, \{v_k^c\}} \sum_{i=1}^{|X|} \sum_{n=1}^N \left\| W_n * \Phi_{l-1}(x_i) - \sum_{c=1}^C \sum_{k=1}^K h_n^k * v_k^c * \Phi_{l-1}(x_i) \right\|_2^2

    where XX is the set of training examples, xi∈Xx_i \in X, Φl−1(xi)\Phi_{l-1}(x_i) is the activation tensor produced up to layer l−1l-1, WnW_n are the original layer-ll filters, vkc∈Rd×1v_k^c \in \mathbb{R}^{d \times 1} are vertical filter components, and hnk∈R1×dh_n^k \in \mathbb{R}^{1 \times d} are horizontal filter components.

    Optimization is implemented by mirroring the network architecture up to layer ll with the factored approximation layers and back-propagating the L2L_2 output error layer by layer. To prevent cascading approximation errors across stacked layers, the training input for approximating layer ll is fed through the previously approximated layers Φl−1approx(xi)\Phi_{l-1}^{\text{approx}}(x_i) rather than the original network layers. Furthermore, all stacked approximation layers can be fine-tuned jointly using back-propagation.

    This approach restricts filter approximation error to the manifold of real training data, outperforming direct filter weight reconstruction and avoiding the under-fitting or overfitting observed when retraining low-rank layers from scratch using classification loss.

  4. Knowl 4 — Filter Reconstruction Optimization via Nuclear Norm and Alternating Conjugate Gradients

    model/method

    Filter reconstruction optimization finds the parameters of the low-rank approximation schemes by minimizing the L2L_2 difference between the original filter kernels and the reconstructed filter kernels without using training data.

    For Scheme 1, the objective optimizes the shared basis filters sm∈Rd×ds_m \in \mathbb{R}^{d \times d} and mixing coefficients ancm∈Ra_n^{cm} \in \mathbb{R} while penalizing the nuclear norm ∥sm∥∗\|s_m\|_* (the sum of singular values) as a convex proxy for rank-1 spatial separability:

    min⁡{sm},{an}∑n=1N∑c=1C∥Wnc−∑m=1Mancmsm∥22+λ∑m=1M∥sm∥∗\min_{\{s_m\}, \{a_n\}} \sum_{n=1}^N \sum_{c=1}^C \left\| W_n^c - \sum_{m=1}^M a_n^{cm} s_m \right\|_2^2 + \lambda \sum_{m=1}^M \|s_m\|_*

    where λ>0\lambda > 0 is a regularization parameter chosen large enough to enforce rank-1 solutions. This optimization problem is biconvex and is solved by alternating between optimizing sms_m and ancma_n^{cm}.

    For Scheme 2, spatial separability is enforced directly by parameterizing the reconstruction with horizontal filters hnk∈R1×dh_n^k \in \mathbb{R}^{1 \times d} and vertical filters vkc∈Rd×1v_k^c \in \mathbb{R}^{d \times 1}:

    min⁡{hnk},{vkc}∑n=1N∑c=1C∥Wnc−∑k=1Khnk∗vkc∥22\min_{\{h_n^k\}, \{v_k^c\}} \sum_{n=1}^N \sum_{c=1}^C \left\| W_n^c - \sum_{k=1}^K h_n^k * v_k^c \right\|_2^2

    This formulation avoids nuclear norm constraints and is solved by alternating conjugate gradient descent between the horizontal filter set {hnk}\{h_n^k\} and the vertical filter set {vkc}\{v_k^c\} until convergence.

  5. Knowl 5 — Architectural Efficiency Divergence Between Scheme 1 and Scheme 2 in Standard Frameworks

    empirical result

    While Scheme 1 achieves a superior theoretical error-to-speedup ratio compared to Scheme 2 when evaluated by theoretical operation counts, Scheme 2 yields significantly greater practical speedups when executed on standard CPU and GPU convolutional engines (such as Caffe with BLAS).

    In standard implementations, multi-channel convolutions are executed via an image-to-column transformation (im2col) followed by a single Basic Linear Algebra Subprograms (BLAS) general matrix multiplication (GEMM) summing across all input channels.

    Scheme 2 operates as two standard 3D convolutions with rectangular kernels (d×1×Cd \times 1 \times C followed by 1×d×K1 \times d \times K), requiring exactly two im2col operations and two BLAS GEMM calls per forward pass of the layer.

    In contrast, Scheme 1 requires computing separable convolutions per input channel independently before recombining them, requiring CC separate im2col calls without channel summation followed by an additional recombination step. The resulting memory rearrangement overhead negates the theoretical FLOP savings of Scheme 1, making Scheme 2 faster in practice.

  6. Knowl 6 — Layer Configuration and Computational Profile of Scene Text Character Recognition CNN

    data/table

    The base model evaluated for low-rank approximation is a 4-layer convolutional neural network for 37-class case-insensitive scene text character recognition (26 letters, 10 digits, and 1 background class) operating on 24×2424 \times 24 zero-centered and variance-normalized grayscale image patches, utilizing Maxout non-linearities between convolutional layers.

    Layer name Filter size In channels Out channels Filters Maxout groups Time
    Conv1 9×99 \times 9 1 48 96 2 0.473ms (8.3%)
    Conv2 9×99 \times 9 48 64 128 2 3.008ms (52.9%)
    Conv3 8×88 \times 8 64 128 512 4 2.160ms (38.0%)
    Conv4 1×11 \times 1 128 37 148 4 0.041ms (0.7%)
    Softmax - 37 37 - - 0.004ms (0.1%)

    Conv2 and Conv3 collectively account for 90.9% of the total network inference time (5.168 ms5.168\text{ ms} out of 5.686 ms5.686\text{ ms}). Low-rank approximation is applied exclusively to Conv2 and Conv3. Conv4 is omitted because its 1×11 \times 1 filters lack spatial redundancy. Conv1 is omitted because it operates directly on raw pixels where filters cannot be well approximated by separable filters without significant classification degradation.

  7. Knowl 7 — Speedup and Accuracy Trade-Offs on ICDAR 2003 Character Classification

    empirical result

    On the ICDAR 2003 character recognition test set (5,379 cropped alphanumeric characters), the uncompressed 4-layer baseline CNN achieves 91.3% case-insensitive classification accuracy.

    Applying Scheme 2 with layer-wise and joint data reconstruction optimization yields the following speedup-accuracy trade-offs:

    • A 2.5×2.5\times full-network speedup with 0.0% loss in classification accuracy (91.3%).
    • A 4.5×4.5\times full-network speedup with a 1.0% drop in accuracy (90.3%), which still exceeds the prior non-CNN state of the art (89.8%).

    The 4.5×4.5\times speedup network is parameterized by:

    1. Approximating the 128 filters (9×9×489 \times 9 \times 48) of Conv2 using K=31K=31 vertical filters of size 9×1×489 \times 1 \times 48 followed by 128 horizontal filters of size 1×9×311 \times 9 \times 31.
    2. Approximating the 512 filters (8×8×648 \times 8 \times 64) of Conv3 using K=26K=26 vertical filters of size 8×1×648 \times 1 \times 64 followed by 512 horizontal filters of size 1×8×261 \times 8 \times 26.

    In sliding-window scene text character detection maps, the spatial response quality remains sufficient to localize scene text even at a 6.7×6.7\times full-network speedup.

  8. Knowl 8 — Performance Comparison with FFT Convolutions and OverFeat Layer Approximation

    empirical result

    Scheme 2 outperforms Fourier-based convolution acceleration and existing low-rank clustering methods in speedup and accuracy retention:

    1. Compared to FFT-based CNN acceleration on a standard convolutional layer (5×55 \times 5 kernel, 16×16×25616 \times 16 \times 256 input, 384 output filters), Scheme 2 using K=256K=256 basis filters achieves an actual runtime speedup of 2.4×2.4\times on GPU with no expected accuracy loss, compared to 2.2×2.2\times achieved by FFT convolution, while avoiding power-of-two padding and FFT memory overheads.
    2. On layer 2 of the OverFeat ImageNet network, Scheme 2 using filter reconstruction optimization achieves a 2×2\times theoretical speedup with a 0.5% drop in top-5 classification accuracy, compared to a 1.2% drop for the same 2×2\times theoretical speedup using the low-rank approximation and filter clustering method of Denton et al. (2014).

Coverage note — None was omitted; all key theoretical formulations, optimization algorithms, implementation analyses, and empirical benchmark results from the paper are fully covered.

References

  1. 1.http://algoval.essex.ac.uk/icdar/datasets.html.
  2. 2.http://www.iapr-tc11.org/mediawiki/index.php/kaist_scene_text_database.
  3. 3.O. Alsharif and J. Pineau. End-to-End Text Recognition with Hybrid HMM Maxout Models. In International Conference on Learning Representations, 2014.
  4. 4.A. Bissacco, M. Cummins, Y. Netzer, and H. Neven. PhotoOCR: Reading text in uncontrolled conditions. In International Conference of Computer Vision, 2013.
  5. 5.T. de Campos, B. R. Babu, and M. Varma. Character recognition in natural images. 2009.
  6. 6.M. Denil, B. Shakibi, L. Dinh, and N. de Freitas. Predicting parameters in deep learning. In Advances in Neural Information Processing Systems, pages 2148–2156, 2013.
  7. 7.E. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus. Exploiting linear structure within convolutional networks for efficient evaluation. arXiv preprint arXiv:1404.0736, 2014.
  8. 8.C. Farabet, Y. LeCun, K. Kavukcuoglu, E. Culurciello, B. Martini, P. Akselrod, and S. Talay. Large-scale fpga-based convolutional networks. Machine Learning on Very Large Data Sets, 2011.
  9. 9.C. Farabet, C. Couprie, L. Najman, and Y. LeCun. Scene parsing with multiscale feature learning, purity trees, and optimal covers. arXiv preprint arXiv:1202.2160, 2012.
  10. 10.I. J. Goodfellow, Y. Bulatov, J. Ibarz, S. Arnoud, and V. Shet. Multi-digit number recognition from street view imagery using deep convolutional neural networks. In International Conference on Learning Representations, 2013.
  11. 11.I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio. Maxout networks. arXiv preprint arXiv:1302.4389, 2013.
  12. 12.G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov. Improving neural networks by preventing co-adaptation of feature detectors. arXiv preprint arXiv:1207.0580, 2012.
  13. 13.F. Iandola, M. Moskewicz, S. Karayev, R. Girshick, T. Darrell, and K. Keutzer. Densenet: Implementing efficient convnet descriptor pyramids. arXiv preprint arXiv:1404.1869, 2014.
  14. 14.M. Jaderberg, K. Simonyan, A. Vedaldi, and A. Zisserman. Synthetic data and artificial neural networks for natural scene text recognition. arXiv preprint arXiv:1406.2227, 2014.
  15. 15.M. Jaderberg, A Vedaldi, and A. Zisserman. Deep features for text spotting. In European Conference on Computer Vision, 2014.
  16. 16.Y. Jia. Caffe: An open source convolutional architecture for fast feature embedding. http://caffe.berkeleyvision.org/, 2013.
  17. 17.D. Karatzas, F. Shafait, S. Uchida, M. Iwamura, S. R. Mestre, J. Mas, D. F. Mota, J. Almazan, L. P. de las Heras, et al. ICDAR 2013 robust reading competition. In Document Analysis and Recognition (ICDAR), 2013 12th International Conference on, pages 1484–1493. IEEE, 2013.
  18. 18.K. Kavukcuoglu, P. Sermanet, Y. Boureau, K. Gregor, M. Mathieu, and Y. LeCun. Learning convolutional feature hierarchies for visual recognition. In NIPS, volume 1, page 5, 2010.
  19. 19.A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. In NIPS, volume 1, page 4, 2012.
  20. 20.H. Lee, R. Grosse, R. Ranganath, and A. Y. Ng. Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations. In Proceedings of the 26th Annual International Conference on Machine Learning, pages 609–616. ACM, 2009.
  21. 21.S. Lucas. ICDAR 2005 text locating competition results. In Document Analysis and Recognition, 2005. Proceedings. Eighth International Conference on, pages 80–84. IEEE, 2005.
  22. 22.F. Mamalet and C. Garcia. Simplifying convnets for fast learning. In Artificial Neural Networks and Machine Learning–ICANN 2012, pages 58–65. Springer, 2012.
  23. 23.M. Mathieu, M. Henaff, and Y. LeCun. Fast training of convolutional networks through ffts. CoRR, abs/1312.5851, 2013.
  24. 24.L. Neumann and J. Matas. A method for text localization and recognition in real-world images. In Proc. Asian Conf. on Computer Vision, pages 770–783. Springer, 2010.
  25. 25.L. Neumann and J. Matas. Text localization in real-world images using efficiently pruned exhaustive search. In Proc. ICDAR, pages 687–691. IEEE, 2011.
  26. 26.L. Neumann and J. Matas. Real-time scene text localization and recognition. In Proc. CVPR, volume 3, pages 1187–1190. IEEE, 2012.
  27. 27.L. Neumann and J. Matas. Scene text localization and recognition with oriented stroke detection. In 2013 IEEE International Conference on Computer Vision (ICCV 2013), pages 97–104, California, US, December 2013. IEEE. ISBN 978-1-4799-2839-2. doi: 10.1109/ICCV.2013.19.
  28. 28.M. Oquab, L. Bottou, I. Laptev, and J. Sivic. Learning and transferring mid-level image representations using convolutional neural networks. In Computer Vision and Pattern Recognition (CVPR), 2014.
  29. 29.I. Posner, P. Corke, and P. Newman. Using text-spotting to query the world. In Proc. of the IEEE/RSJ Int. Conf. on Intelligent Robots and Systems (IROS), 2010.
  30. 30.T. Quack. Large scale mining and retrieval of visual data in a multimodal context. PhD thesis, ETH Zurich, 2009.
  31. 31.R. Rigamonti, M. A. Brown, and V. Lepetit. Are sparse representations really relevant for image classification? In Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on, pages 1545–1552. IEEE, 2011.
  32. 32.R. Rigamonti, A. Sironi, V. Lepetit, and P. Fua. Learning separable filters. In Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on, pages 2754–2761. IEEE, 2013.
  33. 33.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. arXiv preprint arXiv:1312.6229, 2013.
  34. 34.A. Shahab, F. Shafait, and A. Dengel. ICDAR 2011 robust reading competition challenge 2: Reading text in scene images. In Proc. ICDAR, pages 1491–1496. IEEE, 2011.
  35. 35.H. O. Song, S. Zickler, T. Althoff, R. Girshick, M. Fritz, C. Geyer, P. Felzenszwalb, and T. Darrell. Sparselet models for efficient multiclass object detection. In Computer Vision–ECCV 2012, pages 802–815. Springer, 2012.
  36. 36.H. O. Song, T. Darrell, and R. B. Girshick. Discriminatively activated sparselets. In Proceedings of the 30th International Conference on Machine Learning (ICML-13), pages 196–204, 2013.
  37. 37.Y. Taigman, M. Yang, M. Ranzato, and L. Wolf. Deep-Face: Closing the gap to human-level performance in face verification. In IEEE CVPR, 2014.
  38. 38.A. Toshev and C. Szegedy. DeepPose: Human pose estimation via deep neural networks. arXiv preprint arXiv:1312.4659, 2013.
  39. 39.K. van de Sande, J. Uijlings, T. Gevers, and A. Smeulders. Segmentation as selective search for object recognition. In Computer Vision (ICCV), 2011 IEEE International Conference on, pages 1879–1886. IEEE, 2011.
  40. 40.V. Vanhoucke, A. Senior, and M. Z. Mao. Improving the speed of neural networks on cpus. In Proc. Deep Learning and Unsupervised Feature Learning NIPS Workshop, 2011.
  41. 41.K. Wang, B. Babenko, and S. Belongie. End-to-end scene text recognition. In Proc. ICCV, pages 1457–1464. IEEE, 2011.
  42. 42.T. Wang, D. J. Wu, A. Coates, and A. Y. Ng. End-to-end text recognition with convolutional neural networks. In Pattern Recognition (ICPR), 2012 21st International Conference on, pages 3304–3308. IEEE, 2012.
  43. 43.H. Yang, B. Quehl, and H. Sack. A framework for improved video text detection and recognition. In Int. Journal of Multimedia Tools and Applications (MTAP), 2012.

Citation

MLA
Jaderberg, M., et al. “Speeding up Convolutional Neural Networks with Low Rank Expansions”. arXiv, 2014, http://arxiv.org/abs/1405.3866v1.
APA
Jaderberg, M., Vedaldi, A., & Zisserman, A. (2014). Speeding up Convolutional Neural Networks with Low Rank Expansions. arXiv. http://arxiv.org/abs/1405.3866v1
Chicago
Jaderberg, M., A. Vedaldi, and A. Zisserman. 2014. “Speeding up Convolutional Neural Networks with Low Rank Expansions”. arXiv. http://arxiv.org/abs/1405.3866v1.
Harvard
Jaderberg, M., Vedaldi, A. and Zisserman, A. (2014) “Speeding up Convolutional Neural Networks with Low Rank Expansions”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1405.3866v1.
Vancouver
1. Jaderberg M, Vedaldi A, Zisserman A (2014) Speeding up Convolutional Neural Networks with Low Rank Expansions. arXiv

BibTeX

@article{jaderberg2014speeding,
  title = {Speeding up Convolutional Neural Networks with Low Rank Expansions},
  author = {Jaderberg, Max and Vedaldi, Andrea and Zisserman, Andrew},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1405.3866v1},
  eprint = {1405.3866}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF