EIE: Efficient Inference Engine on Compressed Deep Neural Network
Song HanXingyu LiuHuizi MaoJing PuArdavan PedramMark A. HorowitzWilliam J. Dally
Presents a specialized hardware accelerator designed to execute inference directly on compressed, sparse neural networks in on-chip SRAM, eliminating costly DRAM transfers to achieve thousands-fold gains in energy efficiency over conventional processors.
Large deep neural networks require hundreds of millions of parameters, making them difficult to deploy on embedded systems because fetching weights from external DRAM dominates energy use and exceeds typical power budgets. The article addresses this by evaluating a specialized hardware accelerator that operates directly on compressed neural network models.
The article set out to demonstrate an energy-efficient inference engine capable of accelerating sparse matrix-vector multiplication on networks compressed through pruning and weight sharing, while exploiting both weight and activation sparsity.
The approach involved designing an array of processing elements that store partitions of the compressed model in on-chip SRAM and perform customized sparse computations. The design was evaluated through RTL implementation, synthesis in 45nm CMOS, and cycle-accurate simulation across nine fully-connected layers drawn from AlexNet, VGG-16, and NeuralTalk models.
The analysis shows that moving from DRAM to SRAM yields a 120-fold energy reduction, while sparsity and weight sharing together provide an additional factor of roughly 24. On the benchmarks, EIE delivered 189 times the speed of a high-end CPU and 13 times that of a desktop GPU when both ran uncompressed models; energy efficiency reached 24,000 times and 3,400 times the respective baselines. A 64-element array processed AlexNet fully-connected layers at 18,800 frames per second while dissipating only 590 milliwatts.
These gains matter because they bring state-of-the-art network accuracy within the power and latency constraints of mobile and embedded devices without requiring batching that would increase response time. The architecture scales nearly linearly to at least 256 elements and maintains high efficiency even when activation sparsity reaches 70 percent.
Further work is needed to quantify performance on very small matrices and to integrate support for convolutional layers. The main limitations are modest load imbalance on the smallest layers and reliance on the 45 nm process node for the reported power figures; results remain robust across the evaluated benchmarks.
- Paper: Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding, Song Han et al. (2015). Read Deep Compression first to understand the pruning and weight-sharing pipeline whose compressed models EIE is designed to execute.
- Paper: DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learning, Tianshi Chen et al. (2014). DianNao establishes the small-footprint neural-network accelerator baseline that EIE adapts to handle compressed, sparse models.
- Paper: Learning both Weights and Connections for Efficient Neural Networks, Song Han et al. (2015). This paper introduces the pruning procedure that creates sparse networks, preparing you to follow one of the compression techniques EIE exploits.
- Paper: Efficient Processing of Deep Neural Networks: A Tutorial and Survey, Vivienne Sze et al. (2017). Read this later survey to see how EIE’s focus on data movement and sparse inference fits into the broader development of energy-efficient DNN processing.
