Efficient Processing of Deep Neural Networks: A Tutorial and Survey

Vivienne SzeYu-Hsin ChenTien-Ju YangJoel Emer

article2017Proc. IEEE3,827 citations

Synthesizes hardware architectures and algorithm-hardware co-design strategies for energy-efficient deep neural network acceleration, establishing clear metrics and trade-offs for deploying high-throughput models on specialized hardware.

Listen

Deep neural networks have become the foundation for many artificial intelligence applications such as computer vision, speech recognition, and robotics because they deliver state-of-the-art accuracy by learning high-level features directly from large amounts of data. This capability, however, comes at the cost of high computational complexity that makes deployment on energy-constrained or real-time embedded platforms difficult. The survey therefore set out to evaluate the full range of techniques that can improve energy efficiency and throughput of DNN inference without sacrificing accuracy or increasing hardware cost.

The authors assembled a comprehensive tutorial that begins with background on neural networks and their history, then surveys hardware platforms, specialized accelerator architectures, near-data processing approaches, and joint algorithm-hardware optimizations. They also catalog publicly available development frameworks, pretrained models, and standard datasets, and they define the key benchmarking metrics needed to compare designs. The analysis draws on published results for popular networks such as AlexNet, VGG-16, GoogLeNet, and ResNet, together with concrete measurements of data movement, arithmetic intensity, and memory hierarchy costs.

The survey shows that data movement, rather than arithmetic, dominates both energy and latency; that carefully designed dataflows can exploit convolutional, feature-map, and filter reuse to cut off-chip accesses by orders of magnitude; and that reduced-precision arithmetic, weight pruning, and compact network architectures can reduce model size and operation count by 410× with negligible accuracy loss. It further demonstrates that mixed-signal and emerging memory technologies offer additional opportunities for near-data computation, while softwarehardware co-design yields the largest overall gains.

These findings matter because they directly determine whether DNN-based systems can be deployed at scale in autonomous vehicles, medical devices, and edge sensors where power, latency, and cost budgets are tight. Without such efficiency improvements, the superior accuracy of DNNs remains largely confined to cloud servers.

Designers should therefore adopt a co-design methodology that jointly tunes network architecture, numerical precision, sparsity, and hardware dataflow, using the metrics and resources summarized in the paper to guide iterative refinement. Additional work is still required on training-time methods that preserve accuracy under aggressive quantization or pruning, on robust support for recurrent networks, and on systematic evaluation across the newest large-scale datasets.

The survey is based on literature available through mid-2017 and therefore does not capture subsequent advances; its quantitative claims rest on the accuracy, energy, and throughput numbers reported in the cited studies, which themselves depend on the specific networks and datasets chosen.

arXiv: 1703.09039
Cover for Efficient Processing of Deep Neural Networks: A Tutorial and Survey

Abstract

Deep neural networks (DNNs) are currently widely used for many artificial intelligence (AI) applications including computer vision, speech recognition, and robotics. While DNNs deliver state-of-the-art accuracy on many AI tasks, it comes at the cost of high computational complexity. Accordingly, techniques that enable efficient processing of DNNs to improve energy efficiency and throughput without sacrificing application accuracy or increasing hardware cost are critical to the wide deployment of DNNs in AI systems.

This article aims to provide a comprehensive tutorial and survey about the recent advances towards the goal of enabling efficient processing of DNNs. Specifically, it will provide an overview of DNNs, discuss various hardware platforms and architectures that support DNNs, and highlight key trends in reducing the computation cost of DNNs either solely via hardware design changes or via joint hardware design and DNN algorithm changes. It will also summarize various development resources that enable researchers and practitioners to quickly get started in this field, and highlight important benchmarking metrics and design considerations that should be used for evaluating the rapidly growing number of DNN hardware designs, optionally including algorithmic co-designs, being proposed in academia and industry.

The reader will take away the following concepts from this article: understand the key design considerations for DNNs; be able to evaluate different DNN hardware implementations with benchmarks and comparison metrics; understand the trade-offs between various hardware architectures and platforms; be able to evaluate the utility of various DNN design techniques for efficient processing; and understand recent implementation trends and opportunities.

Table of Contents

  • I. INTRODUCTION
  • II. BACKGROUND ON DEEP NEURAL NETWORKS (DNN)
  • A. Artificial Intelligence and DNNs
  • B. Neural Networks and Deep Neural Networks (DNNs)
  • C. Inference versus Training
  • D. Development History
  • DNN Timeline
  • E. Applications of DNN
  • F. Embedded versus Cloud
  • III. OVERVIEW OF DNNs
  • A. Convolutional Neural Networks (CNNs)
  • B. Popular DNN Models
  • IV. DNN DEVELOPMENT RESOURCES
  • A. Frameworks
  • B. Models
  • C. Popular Datasets for Classification
  • D. Datasets for Other Tasks
  • V. HARDWARE FOR DNN PROCESSING
  • A. Accelerate Kernel Computation on CPU and GPU Platforms
  • B. Energy-Efficient Dataflow for Accelerators
  • VI. NEAR-DATA PROCESSING
  • A. DRAM
  • B. SRAM
  • C. Non-volatile Resistive Memories
  • D. Sensors
  • VII. CO-DESIGN OF DNN MODELS AND HARDWARE
  • A. Reduce Precision
  • B. Reduce Number of Operations and Model Size
  • VIII. BENCHMARKING METRICS FOR DNN EVALUATION AND COMPARISON
  • A. Metrics for DNN Models
  • B. Metrics for DNN Hardware
  • IX. SUMMARY
  • ACKNOWLEDGMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — Taxonomy of Spatial Accelerator Dataflows for Deep Neural Networks

    model/method

    Spatial architectures for deep neural network (DNN) acceleration can be classified into four primary dataflow categories based on how they utilize the local memory hierarchy (register files within processing elements, global buffers, and inter-processing element networks) to optimize data reuse and minimize memory access energy:

    1. Weight Stationary (WS): Minimizes weight access energy by keeping weights stationary in the local register file (RF) of each processing element (PE). The processing element executes multiply-accumulate (MAC) operations that reuse the resident weight across multiple input activations (convolutional and filter reuse). Input activations are broadcast to PEs, and partial sums are accumulated across the spatial array or global buffer.

    2. Output Stationary (OS): Minimizes the read/write energy of partial sums by maintaining the accumulation of partial sums for a given output activation stationary within the PE's RF. Weights are broadcast across the array while input activations stream through PEs. OS has three primary spatial mapping variants:

      • OSA\text{OS}_A: Processes output activations from the same channel simultaneously to optimize convolutional layer reuse.
      • OSB\text{OS}_B: Processes output activations across a mix of multiple channels and multiple spatial positions.
      • OSC\text{OS}_C: Processes output activations from different channels simultaneously, targeting fully-connected (FC) layers where each channel contains only a single output activation.
    3. No Local Reuse (NLR): Eliminates local RF storage inside individual PEs to maximize global on-chip buffer capacity and reduce off-chip memory bandwidth. PEs contain only compute logic (such as custom adder trees); activations are multicasted, weights are single-casted, and partial sums are spatially accumulated across the array before being written directly back to the global buffer.

    4. Row Stationary (RS): Optimizes the combined energy consumption of all data types (weights, input pixels, and partial sums) by assigning 1-D row convolution primitives to each PE. Filter rows remain stationary in the PE's RF while input activation rows are streamed through, accumulating partial sums locally within the PE. Across the 2-D PE array, filter rows are reused horizontally, input activation rows are reused diagonally, and partial sums are accumulated vertically.

  2. Knowl 2 — Mathematical Formulation of Convolutional and Fully-Connected Layers

    equation

    The multidimensional computation of a convolutional (CONV) layer in a deep neural network is defined by:

    O[z][u][x][y]=B[u]+k=0C1i=0S1j=0R1I[z][k][Ux+i][Uy+j]×W[u][k][i][j]O[z][u][x][y] = B[u] + \sum_{k=0}^{C-1} \sum_{i=0}^{S-1} \sum_{j=0}^{R-1} I[z][k][Ux + i][Uy + j] \times W[u][k][i][j]

    subject to the index ranges and output dimension constraints:

    0z<N,0u<M,0x<F,0y<E0 \le z < N, \quad 0 \le u < M, \quad 0 \le x < F, \quad 0 \le y < E

    E=HR+UU,F=WS+UUE = \frac{H - R + U}{U}, \quad F = \frac{W - S + U}{U}

    where:

    • I,W,O,BI, W, O, B represent the input feature map (ifmap), filter weights, output feature map (ofmap), and 1-D bias tensors, respectively.
    • NN is the batch size of 3-D feature maps.
    • MM is the number of 3-D filters (and the number of ofmap channels).
    • CC is the number of ifmap/filter channels.
    • HH and WW are the height and width of the input feature map planes.
    • RR and SS are the height and width of the 2-D filter planes.
    • EE and FF are the height and width of the output feature map planes.
    • UU is the convolution stride.

    A fully-connected (FC) layer is a special case of this formulation without weight sharing, operating under the parameter constraints H=RH = R, W=SW = S, E=1E = 1, F=1F = 1, and U=1U = 1.

  3. Knowl 3 — Energy Efficiency and Memory Hierarchy Breakdown Across Accelerator Dataflows

    empirical result

    Under an identical silicon area constraint with 256 processing elements (PEs), a 0.5–1.0 kB register file (RF) per PE, and a 100–500 kB shared global buffer evaluating AlexNet with a batch size of 16:

    1. Dataflow Energy Ranking: The Row Stationary (RS) dataflow achieves the lowest overall energy consumption, consuming 1.4×1.4\times to 2.5×2.5\times lower total energy than Weight Stationary (WS), Output Stationary (OS), and No Local Reuse (NLR) dataflows for convolutional (CONV) layers, and 1.3×1.3\times lower energy for fully-connected (FC) layers.

    2. Memory Access Cost Hierarchy: Normalized energy cost per access scales significantly across the storage levels:

      • Local Register File (RF, 0.5–1.0 kB): 1×1\times (baseline reference)
      • Inter-PE Network-on-Chip (NoC): 2×2\times
      • Global Buffer (100–500 kB): 6×6\times
      • Off-chip DRAM: 200×200\times
    3. Architectural Trade-offs:

      • WS minimizes weight access energy at the cost of high activation/partial sum movement.
      • OS minimizes partial sum accumulation energy.
      • NLR minimizes off-chip DRAM access due to its large global buffer but expends higher energy accessing the global buffer.
      • RS achieves overall efficiency by shifting the vast majority of accesses for all data types into the lowest-energy RF level.
    4. Layer Energy Distribution: In AlexNet, CONV layers account for approximately 80% of total system energy (dominated by local RF access), whereas FC layers consume the remaining 20% (dominated by off-chip DRAM access).

  4. Knowl 4 — Execution Mechanisms of the Row Stationary Dataflow

    model/method

    The Row Stationary (RS) dataflow maps multidimensional convolutions onto a 2-D processing element (PE) array through three coordinated mechanisms:

    1. 1-D Convolution inside a PE: A 1-D row of filter weights (SS elements) is kept stationary inside the PE register file (RF). The corresponding 1-D row of input activations is streamed into the PE. Sliding-window overlaps are retained in the local RF to maximize convolutional reuse while accumulating partial sums locally within a single memory space.

    2. 2-D Convolution across the Spatial PE Array: For a filter of height RR, a column of RR PEs is used to compute each 1-D convolution concurrently. Partial sums are passed and accumulated vertically across the column to yield a row of the output feature map. Successive columns of PEs receive input activation rows shifted by the vertical stride to generate subsequent output rows. Filter rows are reused horizontally, and input activations are reused diagonally across PEs.

    3. Handling High-Dimensional Mappings:

      • Replication: When the layer's 2-D convolution size requires fewer PEs than available in the physical array, the mapping is replicated across unused PEs to process different channels and filters in parallel.
      • Folding: When the required logical PE array exceeds the physical dimensions (e.g., 27×527 \times 5 logical PEs on a 14×1214 \times 12 physical array), the logical array is divided into smaller tiles (e.g., 14×514 \times 5 and 13×513 \times 5) and executed in sequence, clock-gating unused PEs to save static power.
      • Channel/Filter Interleaving: Multiple channels, filters, or batch feature maps are interleaved or concatenated through each PE to extend reuse.
  5. Knowl 5 — Energy-Aware Network Pruning and Estimation

    model/method

    Conventional magnitude-based pruning removes weights with small absolute values and uses weight counts or multiply-accumulate (MAC) counts as proxies for complexity. However, weight counts do not correlate directly with energy because convolutional (CONV) layers dominate energy consumption despite having far fewer weights than fully-connected (FC) layers.

    Energy-Aware Pruning Methodology:

    1. Direct Energy Estimation: Pruning decisions are guided by a hardware-accurate energy model that computes total energy as: Etotal=Ecomp+EdataE_{\text{total}} = E_{\text{comp}} + E_{\text{data}} where EcompE_{\text{comp}} accounts for active MAC operations, and EdataE_{\text{data}} sums the data movement energy across each level ii of the memory hierarchy (register files, global buffers, DRAM): Edata=i(Number of Accesses at Level i)×(Energy per Access at Level i)E_{\text{data}} = \sum_{i} (\text{Number of Accesses at Level } i) \times (\text{Energy per Access at Level } i) incorporating the exact layer shape configurations, dataflow mapping, and sparsity of activations and weights.
    2. Energy-Guided Saliency: Weights are prioritized for pruning based on their overall contribution to system energy reduction rather than absolute numerical magnitude.
    3. Performance: For AlexNet on ImageNet, energy-aware pruning reduces total network energy by 3.7×3.7\times, outperforming magnitude-based pruning (2.1×2.1\times reduction) by 1.74×1.74\times at equivalent accuracy. When applied to compact models such as GoogLeNet, it achieves an additional 1.6×1.6\times energy reduction.
  6. Knowl 6 — Hardware Impact and Taxonomy of Precision Reduction and Quantization

    model/method

    Reducing numerical precision lowers storage requirements, memory bandwidth, and arithmetic circuit cost in deep neural network processing:

    1. Arithmetic Energy and Area Scaling:

      • Adders: Energy and area scale linearly with bitwidth. An 8-bit fixed-point add uses 3.3×3.3\times less energy (3.8×3.8\times less area) than a 32-bit fixed-point add, and 30×30\times less energy (116×116\times less area) than a 32-bit floating-point add.
      • Multipliers: Energy and area scale quadratically with bitwidth. An 8-bit fixed-point multiply uses 15.5×15.5\times less energy (12.4×12.4\times less area) than a 32-bit fixed-point multiply, and 18.5×18.5\times less energy (27.5×27.5\times less area) than a 32-bit floating-point multiply.
    2. Accumulation Bitwidth Requirements: For NN-bit fixed-point inputs and weights, multiplying produces a 2N2N-bit product. Exact accumulation across a filter plane of size C×R×SC \times R \times S requires 2N+M2N + M bits, where M=log2(C×R×S)M = \lceil \log_2(C \times R \times S) \rceil (typically 1010 to 1616 bits for standard CNNs), before quantizing the final output activation back to NN bits.

    3. Quantization Strategies:

      • Dynamic Fixed Point: Uses an adjustable scaling exponent ff per layer or tensor to optimize dynamic range: (1)s×m×2f(-1)^s \times m \times 2^{-f}.
      • Logarithmic Quantization: Quantizes values to power-of-two levels (2p2^p), converting multiplications into bit-shifts.
      • Weight Sharing (Clustering/Hashing): Compresses weights into UU unique values using kk-means clustering or hash functions. Each weight is stored as a log2U\log_2 U-bit index pointing to a UU-entry codebook of full-precision values, saving storage without modifying the arithmetic precision of MAC hardware.
      • Binary and Ternary Networks: Constrains weights to {w,w}\{-w, w\} (binary) or {w1,0,w2}\{-w_1, 0, w_2\} (ternary), replacing multiplications with additions, subtractions, or XNOR operations.
  7. Knowl 7 — Numerical Precision Reduction Methods on AlexNet ImageNet Classification

    data/table

    The following table details the impact of linear, non-linear, binary, and ternary quantization methods on the top-5 classification error of AlexNet on the ImageNet dataset relative to a 32-bit floating-point baseline:

    Quantization Method Weight Bits Activation Bits Accuracy Loss vs. FP32 (%)
    Dynamic Fixed Point (w/o fine-tuning) 8 10 0.4
    Dynamic Fixed Point (w/ fine-tuning) 8 8 0.6
    BinaryConnect 1 32 (float) 19.2
    Binary Weight Network (BWN)^* 1 32 (float) 0.8
    Ternary Weight Networks (TWN)^* 2 32 (float) 3.7
    Trained Ternary Quantization (TTQ)^* 2 32 (float) 0.6
    XNOR-Net^* 1 1 11.0
    Binarized Neural Networks (BNN) 1 1 29.8
    DoReFa-Net^* 1 2 7.63
    Quantized Neural Networks (QNN) 1 2 6.5
    HWGQ-Net^* 1 2 5.2
    LogNet 5 (conv), 4 (fc) 4 3.2
    Incremental Network Quantization (INQ) 5 32 (float) -0.2
    Deep Compression (Weight Sharing) 8 (conv), 4 (fc) 16 0.0
    Deep Compression (Aggressive) 4 (conv), 2 (fc) 16 2.6

    ^*Method does not quantize the first and/or last layers (they remain in higher precision).

    Key takeaways from the data:

    • Quantizing weights and activations to 8-bit dynamic fixed point incurs negligible accuracy degradation (<0.6%<0.6\%).
    • Binarizing both weights and activations to 1-bit causes substantial accuracy drops (11.0%11.0\% to 29.8%29.8\%), but increasing activation precision to 2 bits (e.g., HWGQ-Net, QNN) recovers accuracy significantly.
    • Codebook weight sharing (Deep Compression) preserves baseline accuracy (0.0%0.0\% loss) at 8-bit convolutional and 4-bit fully connected weight indexes.
  8. Knowl 8 — Near-Data and In-Memory Processing Architectures for DNN Acceleration

    model/method

    To counter the high energy cost of off-chip memory access, near-data and processing-in-memory (PIM) architectures place computation directly within or adjacent to memory technologies:

    1. Embedded DRAM (eDRAM): Integrates dense memory on-chip, offering 2.85×2.85\times higher density than SRAM and 321×321\times lower access energy than DDR3 DRAM (as implemented in DaDianNao to store tens of megabytes of weights and activations on-chip).

    2. 3-D Stacked Memory (HMC / HBM): Connects DRAM dies to logic dies using Through-Silicon Vias (TSVs), providing an order of magnitude higher memory bandwidth and up to 5×5\times lower access energy than 2-D DRAM. Accelerator designs (e.g., TETRIS, Neurocube) exploit this bandwidth by placing SIMD logic or spatial arrays adjacent to the memory layers, reallocating on-chip area toward PEs rather than large global buffers.

    3. Compute-in-SRAM: Merges analog MAC computation into standard 6T SRAM bit-cells by driving wordlines with digital-to-analog converter (DAC) voltages representing inputs and discharging bitlines with column-summed cell currents (IBC=Vinput×WI_{\text{BC}} = V_{\text{input}} \times W), achieving up to 12×12\times energy reduction compared to reading SRAM data for external digital computation.

    4. Non-Volatile Resistive Memory (Memristors / RRAM / PCM / STT-MRAM): Implements in-situ matrix-vector multiplication in crossbar arrays where conductance represents weight (GG), input voltage represents activation (VV), and output current represents the dot-product sum (I=ViGiI = \sum V_i G_i) via Ohm's and Kirchhoff's laws. Trade-offs include high density and zero standby power, offset by analog ADC/DAC conversion overhead, IR drop over large crossbar wires, and device variation non-idealities.

    5. Near-Sensor Analog Processing: Executes early convolutional layers or gradient extractions within the sensor using switched-capacitor ADCs or Angle Sensitive Pixels before digitalization, reducing sensor-to-processor data transfer bandwidth by up to 10×10\times to 21×21\times.

  9. Knowl 9 — Compact Convolutional Filter Decompositions and Topologies

    model/method

    Network architecture co-design reduces the parameter count and computational complexity of standard convolutions (C×R×S×MC \times R \times S \times M) through structural decompositions:

    1. Spatial Filter Factorization: Replaces large 2-D filters with a sequence of smaller filters having an identical effective receptive field. For example, a 5×55 \times 5 convolution (2525 parameters/channel) is replaced by two stacked 3×33 \times 3 convolutions (1818 parameters/channel), reducing weights and MACs by 28%28\%.

    2. Separable 1-D Convolutions: Decomposes an N×NN \times N spatial filter into a 1-D horizontal convolution (1×N1 \times N) followed by a 1-D vertical convolution (N×1N \times 1), reducing computational operations from O(N2)O(N^2) to O(2N)O(2N).

    3. Depthwise Separable Convolutions: Decouples spatial filtering from channel mixing by replacing a standard 3-D convolution with a 2-D depthwise convolution (applied independently to each channel) followed by a 1×11 \times 1 pointwise 3-D convolution across channels (used in MobileNets and Xception).

    4. Bottleneck 1×11 \times 1 Layers: Uses 1×11 \times 1 convolutional layers to project high-dimensional channel spaces down to a lower-dimensional intermediate subspace before executing expensive spatial convolutions (e.g., Inception modules, ResNet bottleneck blocks, and SqueezeNet Fire modules).

    5. Post-Training Tensor Decompositions: Decomposes trained 4-D weight tensors into lower-rank factor tensors using Canonical Polyadic (CP) or Tucker decompositions. Low-rank approximations combined with fine-tuning achieve speedups (e.g., 4.5×4.5\times on CPUs) without requiring network retraining from scratch.

  10. Knowl 10 — Standardized Benchmarking and Metric Reporting Framework for DNN Hardware

    model/method

    To enable fair, reproducible, and holistic evaluation across diverse DNN hardware accelerators and algorithm-hardware co-designs, performance must be assessed across four interdependent metric axes:

    1. Accuracy and Workload Specification:

      • Target AI task and dataset difficulty (e.g., ImageNet top-5 error rather than simple tasks like MNIST).
      • Evaluation protocol: Single-crop single-model inference versus multi-crop ensemble testing.
      • Precision: Specific bitwidth formats for weights, activations, and accumulation across every layer.
    2. Compute Complexity and Sparsity:

      • Total weights and multiply-accumulate (MAC) counts.
      • Explicit count of non-zero (NZ) weights and non-zero (NZ) MAC operations evaluated on standard validation sets (e.g., ImageNet 50,000 validation images), quantifying the actual exploited sparsity in both activations and weights.
    3. Throughput, Latency, and Memory Bandwidth:

      • Latency and throughput measured under specific batch sizes on actual target networks, accounting for PE utilization, mapping overhead, and memory bottlenecks (rather than theoretical peak throughput).
      • Off-chip memory traffic: Total DRAM data volume read and written per image inference (in MBytes), capturing system-level power impact.
    4. Hardware Cost and Energy Efficiency:

      • Silicon area and density: Core area (mm2\text{mm}^2), process technology node (nm), supply voltage, core area per multiplier (mm2/multiplier\text{mm}^2/\text{multiplier}), and on-chip memory capacity per multiplier (kB/multiplier\text{kB}/\text{multiplier}).
      • Operating power (mW) and system-level energy per inference measured on real hardware test setups (distinguishing physical chip measurements from post-synthesis or post-place-and-route simulations).

Coverage note — None omitted. The extracted knowls cover all core contributions of the paper: the dataflow taxonomy (WS, OS, NLR, RS), mathematical layer formulations, comparative energy analysis across memory hierarchies, row stationary mechanics, energy-aware pruning and estimation, quantization methods and data, near-data/in-memory processing architectures, compact convolutional topologies, and the hardware evaluation/benchmarking framework.

References

  1. 1.[1] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436–444, May 2015.
  2. 2.[2] L. Deng, J. Li, J.-T. Huang, K. Yao, D. Yu, F. Seide, M. Seltzer, G. Zweig, X. He, J. Williams et al., “Recent advances in deep learning for speech research at Microsoft,” in ICASSP, 2013.
  3. 3.[3] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in NIPS, 2012.
  4. 4.[4] C. Chen, A. Seff, A. Kornhauser, and J. Xiao, “Deepdriving: Learning affordance for direct perception in autonomous driving,” in ICCV, 2015.
  5. 5.[5] A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, and S. Thrun, “Dermatologist-level classification of skin cancer with deep neural networks,” Nature, vol. 542, no. 7639, pp. 115–118, 2017.
  6. 6.[6] D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis, “Mastering the game of Go with deep neural networks and tree search,” Nature, vol. 529, no. 7587, pp. 484–489, Jan. 2016.
  7. 7.[7] F.-F. Li, A. Karpathy, and J. Johnson, “Stanford CS class CS231n: Convolutional Neural Networks for Visual Recogni­tion,” http://cs231n.stanford.edu/.
  8. 8.[8] P. A. Merolla, J. V. Arthur, R. Alvarez-Icaza, A. S. Cassidy, J. Sawada, F. Akopyan, B. L. Jackson, N. Imam, C. Guo, Y. Nakamura et al., “A million spiking-neuron integrated circuit with a scalable communication network and interface,” Science, vol. 345, no. 6197, pp. 668–673, 2014.
  9. 9.[9] S. K. Esser, P. A. Merolla, J. V. Arthur, A. S. Cassidy, R. Appuswamy, A. Andreopoulos, D. J. Berg, J. L. McKinstry, T. Melano, D. R. Barch et al., “Convolutional networks for fast, energy-efficient neuromorphic computing,” Proceedings of the National Academy of Sciences, 2016.
  10. 10.[10] M. Mathieu, M. Henaff, and Y. LeCun, “Fast training of convolutional networks through FFTs,” in ICLR, 2014.
  11. 11.[11] Y. LeCun, L. D. Jackel, B. Boser, J. S. Denker, H. P. Graf, I. Guyon, D. Henderson, R. E. Howard, and W. Hubbard, “Handwritten digit recognition: applications of neural network chips and automatic learning,” IEEE Commun. Mag., vol. 27, no. 11, pp. 41–46, Nov 1989.
  12. 12.[12] B. Widrow and M. E. Hoff, “Adaptive switching circuits,” in 1960 IRE WESCON Convention Record, 1960.
  13. 13.[13] B. Widrow, “Thinking about thinking: the discovery of the LMS algorithm,” IEEE Signal Process. Mag., 2005.
  14. 14.[14] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision (IJCV), vol. 115, no. 3, pp. 211–252, 2015.
  15. 15.[15] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in CVPR, 2016.
  16. 16.[16] “Complete Visual Networking Index (VNI) Forecast,” Cisco, June 2016.
  17. 17.[17] J. Woodhouse, “Big, big, big data: higher and higher resolution video surveillance,” technology.ihs.com, January 2016.
  18. 18.[18] R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation,” in CVPR, 2014.
  19. 19.[19] J. Long, E. Shelhamer, and T. Darrell, “Fully Convolutional Networks for Semantic Segmentation,” in CVPR, 2015.
  20. 20.[20] K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in NIPS, 2014.
  21. 21.[21] G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al., “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Process. Mag., vol. 29, no. 6, pp. 82–97, 2012.
  22. 22.[22] R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa, “Natural language processing (almost) from scratch,” Journal of Machine Learning Research, vol. 12, no. Aug, pp. 2493–2537, 2011.
  23. 23.[23] A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” CoRR abs/1609.03499, 2016.
  24. 24.[24] H. Y. Xiong, B. Alipanahi, L. J. Lee, H. Bretschneider, D. Merico, R. K. Yuen, Y. Hua, S. Gueroussov, H. S. Najafabadi, T. R. Hughes et al., “The human splicing code reveals new insights into the genetic determinants of disease,” Science, vol. 347, no. 6218, p. 1254806, 2015.
  25. 25.[25] J. Zhou and O. G. Troyanskaya, “Predicting effects of noncod­ing variants with deep learning-based sequence model,” Nature methods, vol. 12, no. 10, pp. 931–934, 2015.
  26. 26.[26] B. Alipanahi, A. Delong, M. T. Weirauch, and B. J. Frey, “Predicting the sequence specificities of dna-and rna-binding proteins by deep learning,” Nature biotechnology, vol. 33, no. 8, pp. 831–838, 2015.
  27. 27.[27] H. Zeng, M. D. Edwards, G. Liu, and D. K. Gifford, “Convolu­tional neural network architectures for predicting dna–protein binding,” Bioinformatics, vol. 32, no. 12, pp. i121–i127, 2016.
  28. 28.[28] M. Jermyn, J. Desroches, J. Mercier, M.-A. Tremblay, K. St­Arnaud, M.-C. Guiot, K. Petrecca, and F. Leblond, “Neural net­works improve brain cancer detection with raman spectroscopy in the presence of operating room light artifacts,” Journal of Biomedical Optics, vol. 21, no. 9, pp. 094 002–094 002, 2016.
  29. 29.[29] D. Wang, A. Khosla, R. Gargeya, H. Irshad, and A. H. Beck, “Deep learning for identifying metastatic breast cancer,” arXiv preprint arXiv:1606.05718, 2016.
  30. 30.[30] L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Rein­forcement learning: A survey,” Journal of artificial intelligence research, vol. 4, pp. 237–285, 1996.
  31. 31.[31] V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing Atari with Deep Reinforcement Learning,” in NIPS Deep Learning Workshop, 2013.
  32. 32.[32] S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” Journal of Machine Learning Research, vol. 17, no. 39, pp. 1–40, 2016.
  33. 33.[33] M. Pfeiffer, M. Schaeuble, J. Nieto, R. Siegwart, and C. Cadena, “From Perception to Decision: A Data-driven Approach to End­to-end Motion Planning for Autonomous Ground Robots,” in ICRA, 2017.
  34. 34.[34] S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik, “Cognitive mapping and planning for visual navigation,” in CVPR, 2017.
  35. 35.[35] T. Zhang, G. Kahn, S. Levine, and P. Abbeel, “Learning deep control policies for autonomous aerial vehicles with mpc-guided policy search,” in ICRA, 2016.
  36. 36.[36] S. Shalev-Shwartz, S. Shammah, and A. Shashua, “Safe, multi­agent, reinforcement learning for autonomous driving,” in NIPS Workshop on Learning, Inference and Control of Multi-Agent Systems, 2016.
  37. 37.[37] N. Hemsoth, “The Next Wave of Deep Learning Applications,” Next Platform, September 2016.
  38. 38.[38] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997.
  39. 39.[39] T. N. Sainath, A.-r. Mohamed, B. Kingsbury, and B. Ramab­hadran, “Deep convolutional neural networks for LVCSR,” in ICASSP, 2013.
  40. 40.[40] V. Nair and G. E. Hinton, “Rectified Linear Units Improve Restricted Boltzmann Machines,” in ICML, 2010.
  41. 41.[41] A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlin­earities improve neural network acoustic models,” in ICML, 2013.
  42. 42.[42] K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in ICCV, 2015.
  43. 43.[43] D.-A. Clevert, T. Unterthiner, and S. Hochreiter, “Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs),” ICLR, 2016.
  44. 44.[44] X. Zhang, J. Trmal, D. Povey, and S. Khudanpur, “Improving deep neural network acoustic models using generalized maxout networks,” in ICASSP, 2014.
  45. 45.[45] Y. Zhang, M. Pezeshki, P. Brakel, S. Zhang, , C. Laurent, Y. Bengio, and A. Courville, “Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks,” in Interspeech, 2016.
  46. 46.[46] Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Gir­shick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in ACM International Conference on Multimedia, 2014.
  47. 47.[47] S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML, 2015.
  48. 48.[48] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient­based learning applied to document recognition,” Proc. IEEE, vol. 86, no. 11, pp. 2278–2324, Nov 1998.
  49. 49.[49] P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, “OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks,” in ICLR, 2014.
  50. 50.[50] K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in ICLR, 2015.
  51. 51.[51] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going Deeper With Convolutions,” in CVPR, 2015.
  52. 52.[52] M. Lin, Q. Chen, and S. Yan, “Network in Network,” in ICLR, 2014.
  53. 53.[53] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in CVPR, 2016.
  54. 54.[54] C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi, “Inception­v4, Inception-ResNet and the Impact of Residual Connections on Learning,” in AAAI, 2017.
  55. 55.[55] G. Urban, K. J. Geras, S. E. Kahou, O. Aslan, S. Wang, R. Caruana, A. Mohamed, M. Philipose, and M. Richardson, “Do Deep Convolutional Nets Really Need to be Deep and Convolutional?” ICLR, 2017.
  56. 56.[56] “Caffe LeNet MNIST,” http://caffe.berkeleyvision.org/gathered/ examples/mnist.html.
  57. 57.[57] “Caffe Model Zoo,” http://caffe.berkeleyvision.org/model zoo. html.
  58. 58.[58] “Matconvnet Pretrained Models,” http://www.vlfeat.org/ matconvnet/pretrained/.
  59. 59.[59] “TensorFlow-Slim image classification library,” https://github. com/tensorflow/models/tree/master/slim.
  60. 60.[60] “Deep Learning Frameworks,” https://developer.nvidia.com/ deep-learning-frameworks.
  61. 61.[61] Y.-H. Chen, T. Krishna, J. Emer, and V. Sze, “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolu­tional Neural Networks,” IEEE J. Solid-State Circuits, vol. 51, no. 1, 2017.
  62. 62.[62] C. J. B. Yann LeCun, Corinna Cortes, “THE MNIST DATABASE of handwritten digits,” http://yann.lecun.com/exdb/ mnist/.
  63. 63.[63] L. Wan, M. Zeiler, S. Zhang, Y. L. Cun, and R. Fergus, “Regularization of neural networks using dropconnect,” in ICML, 2013.
  64. 64.[64] A. Krizhevsky, V. Nair, and G. Hinton, “The CIFAR-10 dataset,” https://www.cs.toronto.edu/∼kriz/cifar.html.
  65. 65.[65] A. Torralba, R. Fergus, and W. T. Freeman, “80 million tiny images: A large data set for nonparametric object and scene recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 30, no. 11, pp. 1958–1970, 2008.
  66. 66.[66] A. Krizhevsky and G. Hinton, “Convolutional deep belief networks on cifar-10,” Unpublished manuscript, vol. 40, 2010.
  67. 67.[67] B. Graham, “Fractional max-pooling,” arXiv preprint arXiv:1412.6071, 2014.
  68. 68.[68] “Pascal VOC data sets,” http://host.robots.ox.ac.uk/pascal/ VOC/.
  69. 69.[69] “Microsoft Common Objects in Context (COCO) dataset,” http: //mscoco.org/.
  70. 70.[70] “Google Open Images,” https://github.com/openimages/dataset.
  71. 71.[71] “YouTube-8M,” https://research.google.com/youtube8m/.
  72. 72.[72] “AudioSet,” https://research.google.com/audioset/index.html.
  73. 73.[73] S. Condon, “Facebook unveils Big Basin, new server geared for deep learning,” ZDNet, March 2017.
  74. 74.[74] C. Dubout and F. Fleuret, “Exact acceleration of linear object detectors,” in ECCV, 2012.
  75. 75.[75] J. Cong and B. Xiao, “Minimizing computation in convolutional neural networks,” in ICANN, 2014.
  76. 76.[76] A. Lavin and S. Gray, “Fast algorithms for convolutional neural networks,” in CVPR, 2016.
  77. 77.[77] “Intel Math Kernel Library,” https://software.intel.com/en-us/ mkl.
  78. 78.[78] S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer, “cuDNN: Efficient Primitives for Deep Learning,” arXiv preprint arXiv:1410.0759, 2014.
  79. 79.[79] M. Horowitz, “Computing’s energy problem (and what we can do about it),” in ISSCC, 2014.
  80. 80.[80] Y.-H. Chen, J. Emer, and V. Sze, “Eyeriss: A Spatial Archi­tecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” in ISCA, 2016.
  81. 81.[81] ——, “Using Dataflow to Optimize Energy Efficiency of Deep Neural Network Accelerators,” IEEE Micro’s Top Picks from the Computer Architecture Conferences, vol. 37, no. 3, May-June 2017.
  82. 82.[82] M. Sankaradas, V. Jakkula, S. Cadambi, S. Chakradhar, I. Dur­danovic, E. Cosatto, and H. P. Graf, “A Massively Parallel Coprocessor for Convolutional Neural Networks,” in ASAP, 2009.
  83. 83.[83] V. Sriram, D. Cox, K. H. Tsoi, and W. Luk, “Towards an embedded biologically-inspired machine vision processor,” in FPT, 2010.
  84. 84.[84] S. Chakradhar, M. Sankaradas, V. Jakkula, and S. Cadambi, “A Dynamically Configurable Coprocessor for Convolutional Neural Networks,” in ISCA, 2010.
  85. 85.[85] V. Gokhale, J. Jin, A. Dundar, B. Martini, and E. Culurciello, “A 240 G-ops/s Mobile Coprocessor for Deep Neural Networks,” in CVPR Workshop, 2014.
  86. 86.[86] S. Park, K. Bong, D. Shin, J. Lee, S. Choi, and H.-J. Yoo, “A 1.93TOPS/W scalable deep learning/inference processor with tetra-parallel MIMD architecture for big-data applications,” in ISSCC, 2015.
  87. 87.[87] L. Cavigelli, D. Gschwend, C. Mayer, S. Willi, B. Muheim, and L. Benini, “Origami: A Convolutional Network Accelerator,” in GLVLSI, 2015.
  88. 88.[88] S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep Learning with Limited Numerical Precision,” in ICML, 2015.
  89. 89.[89] Z. Du, R. Fasthuber, T. Chen, P. Ienne, L. Li, T. Luo, X. Feng, Y. Chen, and O. Temam, “ShiDianNao: Shifting Vision Processing Closer to the Sensor,” in ISCA, 2015.
  90. 90.[90] M. Peemen, A. A. A. Setio, B. Mesman, and H. Corporaal, “Memory-centric accelerator design for Convolutional Neural Networks,” in ICCD, 2013.
  91. 91.[91] C. Zhang, P. Li, G. Sun, Y. Guan, B. Xiao, and J. Cong, “Opti­mizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks,” in FPGA, 2015.
  92. 92.[92] T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “DianNao: A Small-footprint High-throughput Accelerator for Ubiquitous Machine-learning,” in ASPLOS, 2014.
  93. 93.[93] Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “DaDianNao: A Machine-Learning Supercomputer,” in MICRO, 2014.
  94. 94.[94] Y.-H. Chen, T. Krishna, J. Emer, and V. Sze, “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convo­lutional Neural Networks,” in ISSCC, 2016.
  95. 95.[95] V. Sze, M. Budagavi, and G. J. Sullivan, “High Efficiency Video Coding (HEVC): Algorithms and Architectures,” in Integrated Circuit and Systems. Springer, 2014, pp. 1–375.
  96. 96.[96] M. Alwani, H. Chen, M. Ferdman, and P. Milder, “Fused-layer CNN accelerators,” in MICRO, 2016.
  97. 97.[97] D. Keitel-Schulz and N. Wehn, “Embedded DRAM develop­ment: Technology, physical design, and application issues,” IEEE Des. Test. Comput., vol. 18, no. 3, pp. 7–15, 2001.
  98. 98.[98] J. Jeddeloh and B. Keeth, “Hybrid memory cube new DRAM architecture increases density and performance,” in Symp. on VLSI, 2012.
  99. 99.[99] J. Standard, “High bandwidth memory (HBM) DRAM,” JESD235, 2013.
  100. 100.[100] D. Kim, J. Kung, S. Chai, S. Yalamanchili, and S. Mukhopad­hyay, “Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory,” in ISCA, 2016.
  101. 101.[101] M. Gao, J. Pu, X. Yang, M. Horowitz, and C. Kozyrakis, “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,” in ASPLOS, 2017.
  102. 102.[102] J. Zhang, Z. Wang, and N. Verma, “A machine-learning classifier implemented in a standard 6T SRAM array,” in Symp. on VLSI, 2016.
  103. 103.[103] Z. Wang, R. Schapire, and N. Verma, “Error-adaptive classifier boosting (EACB): Exploiting data-driven training for highly fault-tolerant hardware,” in ICASSP, 2014.
  104. 104.[104] A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P. Strachan, M. Hu, R. S. Williams, and V. Srikumar, “ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,” in ISCA, 2016.
  105. 105.[105] L. Chua, “Memristor-the missing circuit element,” IEEE Trans. Circuit Theory, vol. 18, no. 5, pp. 507–519, 1971.
  106. 106.[106] L. Wilson, “International technology roadmap for semiconduc­tors (ITRS),” Semiconductor Industry Association, 2013.
  107. 107.[107] Lu, Darsen, “Tutorial on Emerging Memory Devices,” 2016.
  108. 108.[108] S. B. Eryilmaz, S. Joshi, E. Neftci, W. Wan, G. Cauwenberghs, and H.-S. P. Wong, “Neuromorphic architectures with electronic synapses,” in ISQED, 2016.
  109. 109.[109] P. Chi, S. Li, Z. Qi, P. Gu, C. Xu, T. Zhang, J. Zhao, Y. Liu, Y. Wang, and Y. Xie, “PRIME: A Novel Processing-In-Memory Architecture for Neural Network Computation in ReRAM-based Main Memory,” in ISCA, 2016.
  110. 110.[110] M. Prezioso, F. Merrikh-Bayat, B. Hoskins, G. Adam, K. K. Likharev, and D. B. Strukov, “Training and operation of an integrated neuromorphic network based on metal-oxide memristors,” Nature, vol. 521, no. 7550, pp. 61–64, 2015.
  111. 111.[111] J. Zhang, Z. Wang, and N. Verma, “A matrix-multiplying ADC implementing a machine-learning classifier directly with data conversion,” in ISSCC, 2015.
  112. 112.[112] E. H. Lee and S. S. Wong, “A 2.5 GHz 7.7 TOPS/W switched­capacitor matrix multiplier with co-designed local memory in 40nm,” in ISSCC, 2016.
  113. 113.[113] R. LiKamWa, Y. Hou, J. Gao, M. Polansky, and L. Zhong, “RedEye: analog ConvNet image sensor architecture for contin­uous mobile vision,” in ISCA, 2016.
  114. 114.[114] A. Wang, S. Sivaramakrishnan, and A. Molnar, “A 180nm CMOS image sensor with on-chip optoelectronic image com­pression,” in CICC, 2012.
  115. 115.[115] H. Chen, S. Jayasuriya, J. Yang, J. Stephen, S. Sivaramakrish­nan, A. Veeraraghavan, and A. Molnar, “ASP Vision: Optically Computing the First Layer of Convolutional Neural Networks using Angle Sensitive Pixels,” in CVPR, 2016.
  116. 116.[116] A. Suleiman and V. Sze, “Energy-efficient HOG-based object detection at 1080HD 60 fps with multi-scale support,” in SiPS, 2014.
  117. 117.[117] E. H. Lee, D. Miyashita, E. Chai, B. Murmann, and S. S. Wong, “Lognet: Energy-Efficient Neural Networks Using Logrithmic Computations,” in ICASSP, 2017.
  118. 118.[118] S. Han, H. Mao, and W. J. Dally, “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” in ICLR, 2016.
  119. 119.[119] I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Ben­gio, “Quantized neural networks: Training neural networks with low precision weights and activations,” arXiv preprint arXiv:1609.07061, 2016.
  120. 120.[120] S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou, “DoReFa­Net: Training low bitwidth convolutional neural networks with low bitwidth gradients,” arXiv preprint arXiv:1606.06160, 2016.
  121. 121.[121] Y. Ma, N. Suda, Y. Cao, J.-S. Seo, and S. Vrudhula, “Scalable and modularized RTL compilation of Convolutional Neural Networks onto FPGA,” in FPL, 2016.
  122. 122.[122] P. Gysel, M. Motamedi, and S. Ghiasi, “Hardware-oriented Approximation of Convolutional Neural Networks,” in ICLR, 2016.
  123. 123.[123] S. Higginbotham, “Google Takes Unconventional Route with Homegrown Machine Learning Chips,” Next Platform, May 2016.
  124. 124.[124] T. P. Morgan, “Nvidia Pushes Deep Learning Inference With New Pascal GPUs,” Next Platform, September 2016.
  125. 125.[125] P. Judd, J. Albericio, T. Hetherington, T. M. Aamodt, and A. Moshovos, “Stripes: Bit-serial deep neural network comput­ing,” in MICRO, 2016.
  126. 126.[126] B. Moons and M. Verhelst, “A 0.3–2.6 TOPS/W precision­scalable processor for real-time large-scale ConvNets,” in Symp. on VLSI, 2016.
  127. 127.[127] M. Courbariaux, Y. Bengio, and J.-P. David, “Binaryconnect: Training deep neural networks with binary weights during propagations,” in NIPS, 2015.
  128. 128.[128] M. Courbariaux and Y. Bengio, “Binarynet: Training deep neural networks with weights and activations constrained to+ 1 or-1,” arXiv preprint arXiv:1602.02830, 2016.
  129. 129.[129] M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “XNOR­Net: ImageNet Classification Using Binary Convolutional Neural Networks,” in ECCV, 2016.
  130. 130.[130] Z. Cai, X. He, J. Sun, and N. Vasconcelos, “Deep learning with low precision by half-wave gaussian quantization,” in CVPR, 2017.
  131. 131.[131] F. Li and B. Liu, “Ternary weight networks,” in NIPS Workshop on Efficient Methods for Deep Neural Networks, 2016.
  132. 132.[132] C. Zhu, S. Han, H. Mao, and W. J. Dally, “Trained Ternary Quantization,” ICLR, 2017.
  133. 133.[133] R. Andri, L. Cavigelli, D. Rossi, and L. Benini, “YodaNN: An Ultra-Low Power Convolutional Neural Network Accelerator Based on Binary Weights,” in ISVLSI, 2016.
  134. 134.[134] K. Ando, K. Ueyoshi, K. Orimo, H. Yonekawa, S. Sato, H. Nakahara, M. Ikebe, T. Asai, S. Takamaeda-Yamazaki, and M. Kuroda, T.and Motomura, “BRein Memory: A 13-Layer 4.2 K Neuron/0.8 M Synapse Binary/Ternary Reconfigurable In-Memory Deep Neural Network Accelerator in 65nm CMOS,” in Symp. on VLSI, 2017.
  135. 135.[135] D. Miyashita, E. H. Lee, and B. Murmann, “Convolutional Neural Networks using Logarithmic Data Representation,” arXiv preprint arXiv:1603.01025, 2016.
  136. 136.[136] A. Zhou, A. Yao, Y. Guo, L. Xu, and Y. Chen, “Incremental Network Quantization: Towards Lossless CNNs with Low­precision Weights,” in ICLR, 2017.
  137. 137.[137] W. Chen, J. T. Wilson, S. Tyree, K. Q. Weinberger, and Y. Chen, “Compressing Neural Networks with the Hashing Trick,” in ICML, 2015.
  138. 138.[138] J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: ineffectual-neuron-free deep neural network computing,” in ISCA, 2016.
  139. 139.[139] B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. K. Lee, J. M. Hernandez-Lobato, G.-Y. Wei, and D. Brooks, ´ “Minerva: Enabling low-power, highly-accurate deep neural network accelerators,” in ISCA, 2016.
  140. 140.[140] Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal Brain Damage,” in NIPS, 1990.
  141. 141.[141] S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both weights and connections for efficient neural networks,” in NIPS, 2015.
  142. 142.[142] T.-J. Yang, Y.-H. Chen, and V. Sze, “Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning,” in CVPR, 2017.
  143. 143.[143] “DNN Energy Estimation,” http://eyeriss.mit.edu/energy.html.
  144. 144.[144] R. Dorrance, F. Ren, and D. Markovic, “A scalable sparse ´ matrix-vector multiplication kernel for energy-efficient sparse­blas on FPGAs,” in ISFPGA, 2014.
  145. 145.[145] S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “EIE: efficient inference engine on compressed deep neural network,” in ISCA, 2016.
  146. 146.[146] A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “Scnn: An accelerator for compressed-sparse convolutional neural networks,” in ISCA, 2017.
  147. 147.[147] W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” in NIPS, 2016.
  148. 148.[148] S. Anwar, K. Hwang, and W. Sung, “Structured pruning of deep convolutional neural networks,” ACM Journal of Emerging Technologies in Computing Systems, vol. 13, no. 3, p. 32, 2017.
  149. 149.[149] J. Yu, A. Lukefahr, D. Palframan, G. Dasika, R. Das, and S. Mahlke, “Scalpel: Customizing dnn pruning to the underlying hardware parallelism,” in ISCA, 2017.
  150. 150.[150] H. Mao, S. Han, J. Pool, W. Li, X. Liu, Y. Wang, and W. J. Dally, “Exploring the regularity of sparse structure in convolutional neural networks,” in CVPR Workshop on Tensor Methods In Computer Vision, 2017.
  151. 151.[151] J. S. Lim, “Two-dimensional signal and image processing,” Englewood Cliffs, NJ, Prentice Hall, 1990, 710 p., 1990.
  152. 152.[152] F. Chollet, “Xception: Deep Learning With Depthwise Separa­ble Convolutions,” CVPR, 2017.
  153. 153.[153] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017.
  154. 154.[154] F. N. Iandola, M. W. Moskewicz, K. Ashraf, S. Han, W. J. Dally, and K. Keutzer, “SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <1MB model size,” ICLR , 2017.
  155. 155.[155] E. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus, “Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation,” in NIPS, 2014.
  156. 156.[156] V. Lebedev, Y. Ganin, M. Rakhuba1, I. Oseledets, and V. Lem­pitsky, “Speeding-Up Convolutional Neural Networks Using Fine-tuned CP-Decomposition,” ICLR, 2015.
  157. 157.[157] Y.-D. Kim, E. Park, S. Yoo, T. Choi, L. Yang, and D. Shin, “Compression of Deep Convolutional Neural Networks for Fast and Low Power Mobile Applications,” in ICLR, 2016.
  158. 158.[158] C. Bucilu, R. Caruana, and A. Niculescu-Mizil, “Model Compression,” in SIGKDD, 2006.
  159. 159.[159] L. Ba and R. Caurana, “Do Deep Nets Really Need to be Deep?” NIPS, 2014.
  160. 160.[160] G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” in NIPS Deep Learning Workshop, 2014.
  161. 161.[161] A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “Fitnets: Hints for Thin Deep Nets,” ICLR, 2015.
  162. 162.[162] “Benchmarking DNN Processors,” http://eyeriss.mit.edu/ benchmarking.html.

Citation

MLA
Sze, V., et al. “Efficient Processing of Deep Neural Networks: A Tutorial and Survey”. arXiv, 2017, http://arxiv.org/abs/1703.09039v2.
APA
Sze, V., Chen, Y.-H., Yang, T.-J., & Emer, J. (2017). Efficient Processing of Deep Neural Networks: A Tutorial and Survey. arXiv. http://arxiv.org/abs/1703.09039v2
Chicago
Sze, V., Y.-H. Chen, T.-J. Yang, and J. Emer. 2017. “Efficient Processing of Deep Neural Networks: A Tutorial and Survey”. arXiv. http://arxiv.org/abs/1703.09039v2.
Harvard
Sze, V. et al. (2017) “Efficient Processing of Deep Neural Networks: A Tutorial and Survey”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1703.09039v2.
Vancouver
1. Sze V, Chen Y-H, Yang T-J, Emer J (2017) Efficient Processing of Deep Neural Networks: A Tutorial and Survey. arXiv

BibTeX

@article{sze2017efficient,
  title = {Efficient Processing of Deep Neural Networks: A Tutorial and Survey},
  author = {Sze, Vivienne and Chen, Yu-Hsin and Yang, Tien-Ju and Emer, Joel},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1703.09039v2},
  eprint = {1703.09039}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF