Speeding up Convolutional Neural Networks with Low Rank Expansions
Max JaderbergAndrea VedaldiAndrew Zisserman
Presents low-rank filter decomposition techniques that accelerate convolutional neural network computation by up to 4.5x on standard hardware with negligible loss in accuracy.
Convolutional neural networks have set high performance benchmarks across computer vision and machine learning tasks. However, their heavy computational requirements pose significant challenges for real-world and real-time deployment, especially in sliding-window object detection systems. Because stacked convolutional layers account for the vast majority of processing time, improving the efficiency of these operations is vital for practical deployment.
The article demonstrates that convolutional layers contain substantial redundancy across spatial dimensions and feature channels. Its primary objective is to evaluate two hardware-agnostic approximation schemes that decompose standard filters into low-rank, separable components, enabling significant inference speedups on pre-trained models with minimal impact on accuracy.
To evaluate these techniques, the authors applied both schemes to a pre-trained network designed for scene text character recognition. The models were evaluated using standard benchmarks comprising over 5,000 cropped character images, alongside training datasets exceeding 160,000 samples. The team tested two optimization approaches: directly minimizing filter reconstruction error and optimizing approximations based on output data reconstruction across layers.
The findings show that substantial computational gains are achievable with negligible accuracy loss. Using the channel-factored approach combined with data reconstruction optimization, the network achieved a 2.5-fold speedup with no loss in classification accuracy, and a 4.5-fold speedup with less than a one percent drop in accuracy (reducing from 91.3% to 90.3%, which remained state-of-the-art). In layer-specific tests, the targeted intermediate layers accounted for roughly 90% of total run time, making them prime candidates for compression. Additionally, practical tests revealed that the second scheme, which factors layers into sequential vertical and horizontal operations, aligned much better with existing computational libraries, delivering faster real-world execution than the per-channel basis alternative.
These results demonstrate that organizations can drastically reduce compute costs, latency, and hardware constraints for computer vision systems without re-architecting baseline models from scratch. Because the approach is flexible and tunable, engineering teams can adjust the trade-off between speed and accuracy to meet specific application latency budgets and deployment environments.
For practical implementation, engineering teams should prioritize the factored sequential scheme (Scheme 2) paired with data reconstruction optimization when accelerating existing pre-trained convolutional layers. Future efforts should explore learning low-rank, separable filter architectures directly during initial training and investigating varied layer configurations to further optimize efficiency.
Confidence in these findings is high for intermediate convolutional layers within standard computer vision architectures. However, decision-makers should note that the first convolutional layer directly processing raw pixels does not lend itself well to separable approximation due to lack of redundancy. Real-world speed gains also depend closely on how underlying software frameworks handle specific matrix operations, meaning realized speedups may vary across different hardware and software environments.
- Paper: Network In Network, Min Lin et al. (2014). Introduces 1x1 convolutions (cascaded cross-channel pooling) to manipulate channel dimensions and structure cross-channel feature redundancy, a core foundation exploited by low-rank expansion methods.
- Paper: Return of the Devil in the Details: Delving Deep into Convolutional Nets, Ken Chatfield et al. (2014). Examines standard deep convolutional network architectures and practical implementation trade-offs that motivate acceleration techniques for convolutional layers.
- Paper: Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation, Emily L. Denton et al. (2014). Develops complementary low-rank tensor decomposition and biclustering techniques to exploit linear redundancies across convolutional and fully connected layers.
- Paper: Rethinking the Inception Architecture for Computer Vision, Christian Szegedy et al. (2015). Extends the principle of low-rank spatial filter decomposition into native architecture design by factorizing standard 2D convolutions into asymmetric 1D operations.
- Paper: Xception: Deep Learning with Depthwise Separable Convolutions, François Chollet (2016). Takes the decoupling of spatial and cross-channel operations to its extreme by formalizing depthwise separable convolutions into standard efficient architectures.
- Paper: MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications, Andrew G. Howard et al. (2017). Builds complete lightweight mobile architectures by standardizing factorized depthwise separable convolutions to cut spatial and channel redundancy.
- Paper: Learning Structured Sparsity in Deep Neural Networks, Wei Wen et al. (2016). Contrasts low-rank expansion approaches with structured sparsity learning to remove channel and filter redundancies directly during network training.
- Paper: Pruning Filters for Efficient ConvNets, Hao Li et al. (2016). Presents an alternative structured acceleration paradigm by pruning whole convolutional filters to reduce runtime without low-rank filter expansions.
- Paper: Channel Pruning for Accelerating Very Deep Neural Networks, Yihui He et al. (2017). Explores layer-by-layer channel selection and reconstruction to accelerate deep networks as a complementary post-training compression approach.
- Paper: ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices, Xiangyu Zhang et al. (2017). Advances efficient cross-channel computation by combining pointwise group convolutions with channel shuffling to minimize execution latency.
- Paper: MobileNetV2: Inverted Residuals and Linear Bottlenecks, Mark Sandler et al. (2018). Introduces inverted residuals and linear bottlenecks to further optimize the trade-off between intermediate channel expansion and spatial filtering efficiency.
- Paper: GhostNet: More Features From Cheap Operations, Kai Han et al. (2019). Generates redundant feature maps using cheap linear transformations, building on the concept of intrinsic representations and low-cost basis operations.
