Resource-Efficient Neural Networks for Embedded Systems

Wolfgang RothGünther SchindlerBernhard KleinRobert PeharzSebastian TschiatschekHolger FröningFranz PernkopfZoubin Ghahramani

article2024JMLR84 citations

Presents a systematic review and empirical analysis of deep neural network optimization techniques, including quantization, pruning, and structural efficiency, to guide the trade-offs between prediction accuracy, latency, and energy consumption across embedded CPUs, GPUs, and FPGAs.

Listen

Modern deep learning achieves outstanding performance across computer vision, speech, and natural language processing, but its high memory and computational demands prevent practical deployment on resource-constrained embedded systems, autonomous platforms, and edge devices. As machine learning transitions into real-world applications with strict limits on power, latency, and hardware capacity, finding effective trade-offs between predictive accuracy and operational resource efficiency has become essential.

The article systematically reviews algorithmic techniques for resource-efficient deep neural network inference and evaluates how these compression methods interact with embedded hardware platforms. The authors categorize optimization methods into three primary domains—quantization, network pruning, and structural efficiency—and analyze their real-world impact through benchmarking on embedded central processing units, graphics processing units, and field-programmable gate arrays using standardized image classification datasets.

The evaluation reveals several crucial operational findings. First, prediction accuracy is significantly more sensitive to the quantization of activations than to the quantization of weights. While reducing weight precision down to one or two bits causes moderate accuracy degradation, preserving two to four bits for activations is essential to avoid severe performance drops. Second, structured channel pruning consistently outperforms kernel and grouped convolution pruning when evaluating total memory footprint; although grouped convolutions drastically reduce theoretical floating-point operations and parameter counts, they fail to reduce activation memory and do not improve inference speed on target devices. Third, hardware architecture heavily dictates compression benefits. On general-purpose embedded central processing units, low-bit quantization fails to improve throughput due to instruction-level overhead and efficient native floating-point units, whereas structured pruning preserves speed. Conversely, custom data-flow architectures on field-programmable gate arrays require low-bit quantization to fit entire models into limited on-chip memory, enabling extremely high throughput, while embedded graphics processing units achieve the best balance of programmability, high throughput, and accuracy when paired with structured channel pruning.

These findings demonstrate that theoretical efficiency metrics, such as parameter counts and operation counts, do not accurately predict real-world execution latency, memory footprint, or energy consumption. Algorithmic compression strategies cannot be designed in isolation from the deployment hardware. Misaligning compression techniques with processor architectures can degrade predictive performance without delivering any throughput or energy advantages.

Organizations deploying deep learning models on edge devices should tailor compression techniques directly to the target hardware platform: structured channel pruning is best suited for embedded central processing units and graphics processing units, whereas aggressive mixed-precision quantization is necessary for data-flow architectures such as field-programmable gate arrays. Furthermore, engineering teams should prioritize the reduction of intermediate activation memory over theoretical floating-point operations to achieve meaningful latency improvements.

Decision-makers should note that the empirical evaluations in the article focus on vision classification benchmarks and fixed 5-Watt power envelopes, and specialized domain-specific accelerators were omitted due to toolchain and flexibility constraints. Despite these boundaries, the evidence provides high confidence that joint algorithmic-hardware optimization is required for effective edge AI deployment.

arXiv: 2001.03048
Cover for Resource-Efficient Neural Networks for Embedded Systems

Abstract

While machine learning is traditionally a resource intensive task, embedded systems, autonomous navigation, and the vision of the Internet of Things fuel the interest in resource-efficient approaches. These approaches aim for a carefully chosen trade-off between performance and resource consumption in terms of computation and energy. The development of such approaches is among the major challenges in current machine learning research and key to ensure a smooth transition of machine learning technology from a scientific environment with virtually unlimited computing resources into everyday’s applications. In this article, we provide an overview of the current state of the art of machine learning techniques facilitating these real-world requirements. In particular, we focus on resource-efficient inference based on deep neural networks (DNNs), the predominant machine learning models of the past decade. We give a comprehensive overview of the vast literature that can be mainly split into three non-mutually exclusive categories: (i) quantized neural networks, (ii) network pruning, and (iii) structural efficiency. These techniques can be applied dur-

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1 Feed-forward Deep Neural Networks
  • 2.2 Training of Deep Neural Networks
  • 2.3 Batch Normalization
  • 2.4 Dropout
  • 2.5 Common Neural Architectures
  • 2.5.1 AlexNet
  • 2.5.2 VGGNet
  • 2.5.3 InceptionNet
  • 2.5.4 ResNet
  • 2.5.5 DenseNet
  • 2.5.6 MobileNet
  • 2.5.7 EfficientNet
  • 2.5.8 Transformers
  • 2.6 The Straight-Through Gradient Estimator
  • 2.7 Bayesian Neural Networks
  • 3. Resource Efficiency in Deep Neural Networks
  • 3.1 Quantized Neural Networks
  • 3.1.1 Early Quantization Approaches
  • 3.1.2 Quantization-aware Training
  • 3.1.3 Bayesian Approaches for Quantization
  • 3.2 Network Pruning
  • 3.2.1 Unstructured Pruning
  • 3.2.2 Structured Pruning
  • 3.2.3 Bayesian Pruning
  • 3.2.4 Dynamic Network Pruning
  • 3.3 Structural Efficiency in DNNs
  • 3.3.1 Weight Sharing
  • 3.3.2 Knowledge Distillation
  • 3.3.3 Special Matrix Structures
  • 3.3.4 Manual Architecture Design
  • 3.3.5 Neural Architecture Search
  • 4. Embedded Hardware for Deep Neural Networks
  • 4.1 CPUs
  • 4.2 GPUs
  • 4.3 FPGAs
  • 4.4 Domain-Specific Accelerators
  • 4.5 Loop-Back vs. Data-Flow Architectures
  • 5. Experimental Results
  • 5.1 Prediction Quality of Compressed DNNs
  • 5.1.1 Prediction Quality using Different Quantization Approaches
  • 5.1.2 Prediction Quality using Different Pruning Structures
  • 5.2 Evaluating Compressed DNNs on Embedded Hardware
  • 5.2.1 Evaluating Compressed DNNs on CPU
  • 5.2.2 Evaluating Quantized DNNs on FPGAs
  • 5.2.3 Evaluating Pruned DNNs on GPU
  • 5.2.4 Overall Comparison
  • 6. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Three dimensions of resource-efficient neural networks

    definition

    Resource efficiency in a deployed deep neural network (DNN) is the joint trade-off among prediction quality, representational efficiency, and computational efficiency. Representational efficiency concerns the model’s memory footprint, which is governed by its architecture, parameter count, numerical representation, and sparsity. Computational efficiency concerns actual inference time, throughput, and energy consumption on target hardware; theoretical operation counts alone may not predict these quantities because memory movement and hardware mapping are also important. The practical deployment objective is therefore to satisfy constraints on accuracy, latency or throughput, chip area, power, and memory simultaneously.

    The survey organizes DNN resource-efficiency methods into three overlapping categories: quantized neural networks reduce the numerical precision of weights and activations; network pruning removes parameters or larger architectural structures; and structural-efficiency methods redesign or train the network to use efficient parameter sharing, matrix structures, building blocks, or automatically discovered architectures. These categories can be combined, such as by applying pruning and quantization to the same network.

  2. Knowl 2 — Quantization and quantization-aware training

    model/method

    DNN quantization reduces the number of bits used to store or compute with weights, activations, gradients, or intermediate values. Let WW denote real-valued weights and XX denote real-valued activations; a bb-bit quantizer maps each value to a finite set of at most 2b2^b representable values. Lower precision reduces storage and can replace expensive arithmetic with cheaper operations. With binary weights and activations in {−1,1}\{-1,1\}, dot products can be implemented using XNOR and bit-count operations.

    Quantization-aware training maintains full-precision auxiliary weights during optimization but applies the quantizer to weights and, where applicable, activations in the forward pass. During backpropagation, the straight-through gradient estimator treats the non-differentiable quantizer as having an approximate nonzero derivative, commonly the identity derivative, and updates the auxiliary full-precision weights. At deployment, only the quantized weights are retained. The surveyed methods include deterministic and stochastic rounding, binary and ternary weights, power-of-two representations, integer-only inference, learned step sizes, and mixed-precision quantization in which different layers receive different bit widths.

    For a linear quantizer with step size Qd>0Q_d>0, dynamic range Qmax⁡>0Q_{\max}>0, and integer bit width Qb≥1Q_b\geq 1, the characteristic quantities satisfy

    Qmax⁡=(2Qb−1−1)Qd.Q_{\max}=(2^{Q_b-1}-1)Q_d.

    The survey emphasizes that lower precision generally improves memory and arithmetic efficiency but makes training harder and can reduce prediction quality, with activation precision often being more consequential than weight precision.

  3. Knowl 3 — Unstructured, structured, and dynamic network pruning

    model/method

    Network pruning creates resource-efficient DNNs by setting parameters or larger network structures to zero. Unstructured pruning removes individual weights wherever they occur. It is usually less damaging to prediction quality, but practical speedups require sparse tensor representations and sufficiently high sparsity. Structured pruning removes complete dimensions or groups, such as neurons, filters, channels, kernels, or layers, leaving dense tensors that can use optimized dense matrix and convolution kernels. It is more hardware-compatible but can be more sensitive to accuracy loss.

    The surveyed pruning criteria include second-order estimates of loss increase, magnitude thresholds followed by retraining, data-driven estimates of the effect of removing channels, group-lasso or batch-normalization scale regularization, stochastic gates with sparsity penalties, and Bayesian dropout-based signal-to-noise criteria. Some approaches allow previously removed weights to return during training rather than making pruning decisions irreversible. Other methods prune conditionally during inference: a lightweight decision network or gating mechanism examines the input or intermediate features and selects which channels, spatial locations, or convolutional outputs to compute. Dynamic pruning can therefore spend more computation on difficult inputs and less on easy inputs, but its gating overhead must be smaller than the computation it saves.

  4. Knowl 4 — Structural efficiency through distillation, sharing, efficient operators, and architecture search

    model/method

    Structural-efficiency methods reduce resource use by changing how a DNN represents or computes functions rather than only reducing numerical precision or zeroing parameters. Knowledge distillation trains a small student network to reproduce the outputs of a larger teacher, usually using soft class probabilities in addition to hard labels. If aia_i is the teacher logit for class ii, CC is the number of classes, and τ>0\tau>0 is a temperature, the teacher soft target is

    y^i=exp⁡(ai/τ)∑j=1Cexp⁡(aj/τ).\hat y_i=\frac{\exp(a_i/\tau)}{\sum_{j=1}^{C}\exp(a_j/\tau)}.

    Temperatures greater than one soften the targets. Distillation can also transfer intermediate features, compress ensembles, improve quantized students, or provide accurate early exits in multi-exit networks.

    Weight sharing assigns many connections to a smaller set of learned values, while hashing can store the assignments implicitly. Low-rank and tensor decompositions replace a matrix W∈Rm×nW\in\mathbb{R}^{m\times n} with factors such as UVUV, where U∈Rm×kU\in\mathbb{R}^{m\times k}, V∈Rk×nV\in\mathbb{R}^{k\times n}, and k<min⁡(m,n)k<\min(m,n); structured matrices such as circulant or Fastfood matrices further reduce storage and accelerate multiplication. Manually designed architectures use global average pooling, 1×11\times1 convolutions, spatially separable convolutions, depthwise-separable convolutions, bottlenecks, grouped convolutions, channel shuffling, residual connections, and squeeze-and-excitation modules to reduce computation while retaining expressive power.

    Neural architecture search (NAS) automates the selection of architectures from a predefined search space. Resource-aware NAS incorporates measured or modeled latency, energy, memory, bit width, or sparsity into the objective. The surveyed approaches use reinforcement-learning controllers, differentiable gates, shared-weight supernets, or hardware measurements. Their effectiveness depends strongly on the search space and target device; the survey notes that current NAS methods are unlikely to discover fundamental design principles absent from that predefined space.

  5. Knowl 5 — Hardware determines which compression is computationally useful

    model/method

    The practical benefit of a compressed DNN depends on the target processor. CPUs provide relatively high clock frequency, short vector units, caches, multithreading, and support for low-precision integer arithmetic. These properties make CPUs comparatively suitable for irregular sparsity and quantization, although their limited parallelism restricts peak throughput. GPUs provide massive parallelism, high memory bandwidth, and high throughput for regular dense operations, but their execution model makes fine-grained sparse computation difficult; common embedded GPUs typically support formats such as 8-bit integers and half-precision floating point rather than arbitrary very-low-bit formats.

    FPGAs contain configurable logic and memory and can tailor compute units, data formats, parallelism, and sparse logic to a particular DNN. Their lower frequency and limited on-chip memory are important constraints. Domain-specific accelerators commonly use dense systolic arrays optimized for regular matrix operations; they reduce data movement but generally support only selected formats such as 8-bit integers and half precision and do not efficiently exploit fine-grained sparsity.

    The survey distinguishes loop-back and data-flow inference architectures. Loop-back systems reuse a fixed processor and memory system across layers, allowing arbitrary operations but potentially incurring costly off-chip transfers and poor utilization. Data-flow systems assign dedicated compute and memory resources to layers and forward intermediate results between them, enabling pipelining and high utilization, but they require longer hardware-development effort and generally must keep the complete network, including weights and activations, on chip. Consequently, operation-count reductions do not necessarily translate into latency or energy reductions unless the compressed representation matches the hardware and software stack.

  6. Knowl 6 — Quantization accuracy trade-offs on CIFAR-100

    empirical result

    The quantization comparison used a 100-layer DenseNet-BC-100 with bottleneck and compression layers and growth rate k=12k=12. The network classified 32×3232\times32 RGB images from CIFAR-100, using 50,000 training images and 10,000 test images. The evaluated methods were Binary Weight Networks (BWN), Binarized Neural Networks (BNN), DoReFa-Net, Trained Ternary Quantization (TTQ), and LQ-Net. Quantization was tested in weight-only, activation-only, and combined weight-and-activation modes; BWN and TTQ were evaluated as weight methods, whereas BNN was designed for combined quantization.

    Across methods and quantization modes, test error generally decreased as the bit width increased. Activation quantization caused a larger prediction-quality penalty than weight quantization. LQ-Net consistently outperformed the simpler linear quantization of DoReFa-Net and the specialized binary or ternary methods, but its training cost was higher: relative to unquantized training, the reported per-iteration time increased by a factor of about 1.51.5 for DoReFa-Net and by up to 4.64.6 for LQ-Net, depending on bit width. The experiment demonstrates that the best accuracy-efficient quantization choice is not necessarily the cheapest method to train.

  7. Knowl 7 — Structured pruning favors channels over kernels and fixed groups

    empirical result

    The structured-pruning comparison used 28-layer Wide Residual Networks (WRNs) and a 28-layer DenseNet variant on CIFAR-10. The DenseNet width was adjusted so that its parameter count and computation approximately matched the WRN, enabling comparison across architectures. Parameterized structured pruning (PSP) learned which tensor structures were unimportant by parameterizing candidate structures and using weight decay to drive removable structures toward zero. The experiment separately evaluated channel pruning, kernel-size pruning, and group pruning, and also evaluated channel pruning combined with fixed group size G=64G=64.

    Channel pruning and group pruning achieved substantially lower floating-point operation counts and parameter counts than kernel pruning at comparable test accuracy, indicating that convolution-kernel size is particularly sensitive to accuracy. Channel and group pruning were similar in FLOPs and parameters for highly compressed models, but group pruning retained the number of input and output channels and therefore produced many activations. It consequently performed much worse in activation count and total memory, making channel pruning the strongest isolated compression structure.

    Combining fixed grouping with channel pruning performed worse than pure channel pruning across the reported metrics and required many activations. The DenseNet variant was more efficient than residual blocks in FLOPs and parameters, whereas the residual architecture used fewer activations and less total memory. Thus, DenseNets were more parameter- and computation-efficient in this comparison, while ResNets were more memory-efficient.

  8. Knowl 8 — Low-bit quantization did not accelerate the ARM CPU experiment

    empirical result

    Inference throughput was evaluated on an ARM Cortex-A53 using 28-layer WRNs trained on CIFAR-10. The compared models varied in width, used LQ-Net for weight and activation quantization at different bit widths, or used PSP channel pruning. The evaluation measured both test accuracy and throughput.

    On this processor, quantization did not improve throughput. Efficient floating-point units, fast on-chip memory, overhead from low-bit computation, and layer dimensions that did not fit the processor’s bit-level vector instructions outweighed the potential arithmetic savings. One-bit weights and activations also produced the largest accuracy degradation. Increasing activation precision to two or three bits while retaining low-precision weights substantially improved accuracy, but the resulting models still did not provide a convincing overall trade-off. Channel-pruned models maintained accuracy and throughput comparable to the uncompressed baseline, making structured channel pruning more suitable for this CPU than the tested low-bit quantization schemes.

  9. Knowl 9 — FPGA throughput increases as precision decreases, with activation bits as the main accuracy bottleneck

    empirical result

    Quantized inference was evaluated with the FINN data-flow framework on a Xilinx Ultra96 FPGA. Because the evaluated FINN configuration did not support residual connections, the experiment used a VGG-style network on CIFAR-10. The hardware configuration targeted the highest throughput permitted by the device resources, including block RAM and lookup tables, and varied weight and activation bit widths.

    Higher bit widths increased test accuracy but reduced throughput. Along the accuracy-throughput Pareto frontier, the most favorable configurations first used one-bit weights while increasing activation precision up to three bits. After that point, increasing weight precision to two bits and activation precision to four bits gave the more favorable trade-off. The result supports the broader finding that activation precision is more important to prediction quality than weight precision, while also showing that an FPGA can turn aggressive quantization into actual throughput gains when the data-flow hardware is designed around the low-bit representation.

  10. Knowl 10 — Embedded GPU and cross-platform results expose the gap between FLOPs and latency

    empirical result

    Pruned WRNs and DenseNets were evaluated with TensorRT on an embedded NVIDIA Jetson Nano GPU using half-precision weights and activations. Kernel pruning, group pruning, and fixed grouping combined with channel pruning gave the poorest throughput behavior. Group pruning was especially revealing: despite sharply reducing FLOPs and parameter counts, it did not translate those reductions into faster inference. Pure channel pruning applied to residual and dense architectures produced the best throughput behavior. These results indicate that reducing activation memory and data movement can matter more for latency than reducing FLOPs alone.

    A cross-platform comparison considered an ARM CPU, an NVIDIA Nano GPU, and a Xilinx Ultra96 FPGA, all with power consumption of roughly five watts. The CPU could deploy relatively large and accurate models because of its comparatively large memory, but its limited parallelism reduced throughput. The FPGA achieved extremely high throughput and hardware utilization, but its data-flow design required the complete model, including activations, to fit on chip, restricting model size. The GPU offered the most balanced compromise between programmability, memory capacity, throughput, and accuracy, although its optimized libraries and software stack limited the compression formats and sparse structures it could exploit.

Coverage note — Detailed Bayesian formulations, straight-through-estimator background, dynamic-pruning variants, named-architecture chronology, and the large method-by-method citation survey were compressed into the broader quantization, pruning, structural-efficiency, and hardware knowls because they support rather than replace the paper’s central taxonomy and experiments.

References

  1. 1.Jan Achterhold, Jan M. K¨ohler, Anke Schmeink, and Tim Genewein. Variational network quantization. In International Conference on Learning Representations (ICLR), 2018.
  2. 2.Alexander G. Anderson and Cory P. Berg. The high-dimensional geometry of binary neural networks. In International Conference on Learning Representations (ICLR), 2018.
  3. 3.Jimmy Ba and Rich Caruana. Do deep nets really need to be deep? In Advances in Neural Information Processing Systems (NIPS), pages 2654–2662, 2014.
  4. 4.Yoshua Bengio, Nicholas L´eonard, and Aaron C. Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. CoRR, abs/1308.3432, 2013.
  5. 5.Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. Weight uncertainty in neural networks. In International Conference on Machine Learning (ICML), pages 1613–1622, 2015.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. CoRR, abs/2005.14165, 2020. URL https://arxiv.org/abs/2005.14165.
  7. 7.Cristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil. Model compression. In ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pages 535–541, 2006.
  8. 8.Han Cai, Ligeng Zhu, and Song Han. ProxylessNAS: Direct neural architecture search on target task and hardware. In International Conference on Learning Representations (ICLR), 2019.
  9. 9.Zhaowei Cai, Xiaodong He, Jian Sun, and Nuno Vasconcelos. Deep learning with low precision by half-wave Gaussian quantization. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 5406–5414, 2017.
  10. 10.Wenlin Chen, James T. Wilson, Stephen Tyree, Kilian Q. Weinberger, and Yixin Chen. Compressing neural networks with the hashing trick. In International Conference on Machine Learning (ICML), pages 2285–2294, 2015.
  11. 11.Yu Cheng, Felix X. Yu, Rog´erio Schmidt Feris, Sanjiv Kumar, Alok N. Choudhary, and Shih-Fu Chang. An exploration of parameter redundancy in deep networks with circulant projections. In International Conference on Computer Vision (ICCV), pages 2857–2865, 2015.
  12. 12.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research (JMLR), 24(240):1–113, 2023.
  13. 13.Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Training deep neural networks with low precision multiplications. In International Conference on Learning Representations (ICLR) Workshop, volume abs/1412.7024, 2015a.
  14. 14.Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. BinaryConnect: Training deep neural networks with binary weights during propagations. In Advances in Neural Information Processing Systems (NIPS), pages 3123–3131, 2015b.
  15. 15.Misha Denil, Babak Shakibi, Laurent Dinh, Marc’Aurelio Ranzato, and Nando de Freitas. Predicting parameters in deep learning. In Advances in Neural Information Processing Systems (NIPS), pages 2148–2156, 2013.
  16. 16.R. H. Dennard, F. H. Gaensslen, H. Yu, V. L. Rideout, E. Bassous, and A. R. LeBlanc. Design of ion-implanted MOSFET’s with very small physical dimensions. IEEE Journal of Solid-State Circuits, 1974.
  17. 17.Emily L. Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus. Exploiting linear structure within convolutional networks for efficient evaluation. In Advances in Neural Information Processing Systems (NIPS), pages 1269–1277, 2014.
  18. 18.Xuanyi Dong, Junshi Huang, Yi Yang, and Shuicheng Yan. More is less: A more complicated network with less inference complexity. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 1895–1903, 2017.
  19. 19.Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. HAWQ: Hessian aware quantization of neural networks with mixed-precision. In International Conference on Computer Vision (ICCV), pages 293–302, 2019.
  20. 20.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. CoRR, abs/2010.11929, Jun 2021.
  21. 21.Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Neural architecture search: A survey. The Journal of Machine Learning Research (JMLR), 20(1):1997–2017, 2019.
  22. 22.Steven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S. Modha. Learned step size quantization. In International Conference on Learning Representations (ICLR), 2020.
  23. 23.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. OPTQ: Accurate quantization for generative pre-trained transformers. In International Conference on Learning Representations (ICLR), 2023.
  24. 24.Advait Harshal Gadhikar, Sohom Mukherjee, and Rebekka Burkholz. Why random pruning is all we need to start sparse. In Proceedings of Machine Learning Research (PMLR), volume 202, 2023.
  25. 25.Xitong Gao, Yiren Zhao, Lukasz Dudziak, Robert Mullins, and Cheng-Zhong Xu. Dynamic channel pruning: Feature boosting and suppression. In International Conference on Learning Representations (ICLR), 2019.
  26. 26.Alex Graves. Practical variational inference for neural networks. In Advances in Neural Information Processing Systems (NIPS), pages 2348–2356, 2011.
  27. 27.Peter D. Gr¨unwald. The minimum description length principle. MIT press, 2007.
  28. 28.Yiwen Guo, Anbang Yao, and Yurong Chen. Dynamic network surgery for efficient DNNs. In Advances in Neural Information Processing Systems (NIPS), pages 1379–1387, 2016.
  29. 29.Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. Deep learning with limited numerical precision. In International Conference on Machine Learning (ICML), pages 1737–1746, 2015.
  30. 30.Song Han, Jeff Pool, John Tran, and William J. Dally. Learning both weights and connections for efficient neural networks. In Advances in Neural Information Processing Systems (NIPS), pages 1135–1143, 2015.
  31. 31.Song Han, Huizi Mao, and William J. Dally. Deep compression: Compressing deep neural network with pruning, trained quantization and Huffman coding. In International Conference on Learning Representations (ICLR), 2016.
  32. 32.Babak Hassibi and David G. Stork. Second order derivatives for network pruning: Optimal brain surgeon. In Advances in Neural Information Processing Systems (NIPS), pages 164–171, 1992.
  33. 33.Marton Havasi, Robert Peharz, and Jos´e Miguel Hern´andez-Lobato. Minimal random code learning: Getting bits back from compressed model parameters. In International Conference on Learning Representations (ICLR), 2019.
  34. 34.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016.
  35. 35.Geoffrey E. Hinton and Drew van Camp. Keeping the neural networks simple by minimizing the description length of the weights. In Conference on Computational Learning Theory (COLT), pages 5–13, 1993.
  36. 36.Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. Distilling the knowledge in a neural network. In Deep Learning and Representation Learning Workshop @ NIPS, 2015.
  37. 37.Markus H¨ohfeld and Scott E. Fahlman. Learning with limited numerical precision using the cascade-correlation algorithm. IEEE Transactions on Neural Networks, 3(4):602–611, 1992a.
  38. 38.Markus H¨ohfeld and Scott E. Fahlman. Probabilistic rounding in neural network learning with limited precision. Neurocomputing, 4(6):291–299, 1992b.
  39. 39.Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V. Le, and Hartwig Adam. Searching for mobilenetv3. CoRR, abs/1905.02244, May 2019. URL https://arxiv.org/abs/1905.02244.
  40. 40.Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. MobileNets: Efficient convolutional neural networks for mobile vision applications. CoRR, abs/1704.04861, 2017a.
  41. 41.Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. CoRR, abs/1704.04861, Apr 2017b. URL https://arxiv.org/abs/1704.04861.
  42. 42.Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 2261–2269, 2017.
  43. 43.Zehao Huang and Naiyan Wang. Data-driven sparse structure selection for deep neural networks. In European Conference on Computer Vision (ECCV), pages 317–334, 2018.
  44. 44.Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Binarized neural networks. In Advances in Neural Information Processing Systems (NIPS), pages 4107–4115, 2016.
  45. 45.Forrest N. Iandola, Matthew W. Moskewicz, Khalid Ashraf, Song Han, William J. Dally, and Kurt Keutzer. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <1mb model size. CoRR, abs/1602.07360, 2016.
  46. 46.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning (ICML), pages 448–456, 2015.
  47. 47.Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, and Dmitry Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 2704–2713, 2018.
  48. 48.Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman. Speeding up convolutional neural networks with low rank expansions. In British Machine Vision Conference (BMVC), 2014.
  49. 49.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. In International Conference on Learning Representations (ICLR), 2017.
  50. 50.Prabhu Kaliamoorthi, Aditya Siddhant, Edward Li, and Melvin Johnson. Distilling large language models into tiny and effective students using pQRNN. CoRR, abs/2101.08890, 2021.
  51. 51.Jangho Kim, Seonguk Park, and Nojun Kwak. Paraphrasing complex network: Network compression via factor transfer. In Advances in Neural Information Processing Systems (NeurIPS), pages 2765–2774, 2018.
  52. 52.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  53. 53.Diederik P Kingma, Tim Salimans, and Max Welling. Variational dropout and the local reparameterization trick. In Advances in Neural Information Processing Systems (NIPS), pages 2575–2583, 2015.
  54. 54.Anoop Korattikara, Vivek Rathod, Kevin P. Murphy, and Max Welling. Bayesian dark knowledge. In Advances in Neural Information Processing Systems (NIPS), pages 3438–3446, 2015.
  55. 55.Torben Krieger, Bernhard Klein, and Holger Fr¨oning. Towards hardware-specific automatic compression of neural networks. CoRR, abs/2212.07818, Dec 2022. URL https://arxiv.org/abs/2212.07818.
  56. 56.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems (NIPS), pages 1106–1114, 2012.
  57. 57.Vadim Lebedev, Yaroslav Ganin, Maksim Rakhuba, Ivan V. Oseledets, and Victor S. Lempitsky. Speeding-up convolutional neural networks using fine-tuned CP-decomposition. In International Conference on Learning Representations (ICLR), 2015.
  58. 58.Yann LeCun, John S. Denker, and Sara A. Solla. Optimal brain damage. In Advances in Neural Information Processing Systems (NIPS), pages 598–605, 1989.
  59. 59.Fengfu Li, Bo Zhang, and Bin Liu. Ternary weight networks. CoRR, abs/1605.04711, 2016.
  60. 60.Hao Li, Soham De, Zheng Xu, Christoph Studer, Hanan Samet, and Tom Goldstein. Training quantized nets: A deeper understanding. In Advances in Neural Information Processing Systems (NIPS), pages 5811–5821, 2017.
  61. 61.Jinyu Li, Rui Zhao, Jui-Ting Huang, and Yifan Gong. Learning small-size DNN with output-distribution-based criteria. In INTERSPEECH: Conference of the International Speech Communication Association, pages 1910–1914, 2014.
  62. 62.Yang Li and Shihao Ji. L0-ARM: Network sparsification via stochastic binary optimization. In European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD), 2019.
  63. 63.Darryl Dexu Lin, Sachin S. Talathi, and V. Sreekanth Annapureddy. Fixed point quantization of deep convolutional networks. In International Conference on Machine Learning (ICML), pages 2849–2858, 2016.
  64. 64.Ji Lin, Yongming Rao, Jiwen Lu, and Jie Zhou. Runtime neural pruning. In Advances in Neural Information Processing Systems (NIPS), pages 2181–2191, 2017a.
  65. 65.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, Chuang Gan, and Song Han. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. CoRR, abs/2306.00978, 2023.
  66. 66.Min Lin, Qiang Chen, and Shuicheng Yan. Network in network. In International Conference on Learning Representations (ICLR), 2014a.
  67. 67.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zitnick. Microsoft COCO: Common objects in context. In European Conference on Computer Vision (ECCV), pages 740–755, 2014b.
  68. 68.Xiaofan Lin, Cong Zhao, and Wei Pan. Towards accurate binary convolutional neural network. In Neural Information Processing Systems (NIPS), pages 345–353, 2017b.
  69. 69.Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, and Yoshua Bengio. Neural networks with few multiplications. In International Conference on Learning Representations (ICLR), 2015.
  70. 70.Hanxiao Liu, Karen Simonyan, and Yiming Yang. DARTS: Differentiable architecture search. In International Conference on Learning Representations (ICLR), 2019a.
  71. 71.Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang, Wei Liu, and Kwang-Ting Cheng. Bi-Real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm. In European Conference on Computer Vision (ECCV), pages 747–763, 2018.
  72. 72.Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. Learning efficient convolutional networks through network slimming. In International Conference on Computer Vision (ICCV), pages 2755–2763, 2017.
  73. 73.Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. Rethinking the value of network pruning. In International Conference on Learning Representations (ICLR), 2019b.
  74. 74.Christos Louizos, Karen Ullrich, and Max Welling. Bayesian compression for deep learning. In Advances in Neural Information Processing Systems (NIPS), pages 3288–3298, 2017.
  75. 75.Christos Louizos, Max Welling, and Diederik P. Kingma. Learning sparse neural networks through L0 regularization. In International Conference on Learning Representations (ICLR), 2018.
  76. 76.Christos Louizos, Matthias Reisser, Tijmen Blankevoort, Efstratios Gavves, and Max Welling. Relaxed quantization for discretized neural networks. In International Conference on Learning Representations (ICLR), 2019.
  77. 77.Jian-Hao Luo, Jianxin Wu, and Weiyao Lin. ThiNet: A filter level pruning method for deep neural network compression. In International Conference on Computer Vision (ICCV), pages 5068–5076, 2017.
  78. 78.Zelda Mariet and Suvrit Sra. Diversity networks: Neural network compression using determinantal point processes. In International Conference on Learning Representations (ICLR), 2016.
  79. 79.Gaurav Menghani. Efficient deep learning: A survey on making deep learning models smaller, faster, and better. ACM Computing Surveys, 55(12), 2023.
  80. 80.Thomas P. Minka. Expectation propagation for approximate Bayesian inference. In Uncertainty in Artificial Intelligence (UAI), pages 362–369, 2001.
  81. 81.Asit K. Mishra and Debbie Marr. Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy. In International Conference on Learning Representations (ICLR), 2018.
  82. 82.Daisuke Miyashita, Edward H. Lee, and Boris Murmann. Convolutional neural networks using logarithmic data representation. CoRR, abs/1603.01025, 2016.
  83. 83.Dmitry Molchanov, Arsenii Ashukha, and Dmitry P. Vetrov. Variational dropout sparsifies deep neural networks. In International Conference on Machine Learning (ICML), pages 2498–2507, 2017.
  84. 84.Radford M. Neal. Bayesian training of backpropagation networks by the hybrid Monte Carlo method. Technical report, Dept. of Computer Science, University of Toronto, 1992.
  85. 85.Jorge Nocedal and Stephen Wright. Numerical Optimization. Springer New York, 2 edition, 2006.
  86. 86.Alexander Novikov, Dmitry Podoprikhin, Anton Osokin, and Dmitry P. Vetrov. Tensorizing neural networks. In Advances in Neural Information Processing Systems (NIPS), pages 442–450, 2015.
  87. 87.Steven J. Nowlan and Geoffrey E. Hinton. Simplifying neural networks by soft weight-sharing. Neural Computation, 4(4):473–493, 1992.
  88. 88.Jorn W. T. Peters and Max Welling. Probabilistic binary neural networks. CoRR, abs/1809.03368, 2018.
  89. 89.Mary Phuong and Christoph Lampert. Distillation-based training for multi-exit architectures. In IEEE International Conference on Computer Vision (ICCV), pages 1355–1364, 2019.
  90. 90.Antonio Polino, Razvan Pascanu, and Dan Alistarh. Model compression via distillation and quantization. In International Conference on Learning Representations (ICLR), 2018.
  91. 91.Rajesh Ranganath, Sean Gerrish, and David M. Blei. Black box variational inference. In International Conference on Artificial Intelligence and Statistics (AISTATS), pages 814–822, 2014.
  92. 92.Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. XNOR-Net: ImageNet classification using binary convolutional neural networks. In European Conference on Computer Vision (ECCV), pages 525–542, 2016.
  93. 93.Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. FitNets: Hints for thin deep nets. In International Conference on Learning Representations (ICLR), 2015.
  94. 94.Wolfgang Roth and Franz Pernkopf. Bayesian neural networks with weight sharing using Dirichlet processes. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2018.
  95. 95.Wolfgang Roth, G¨unther Schindler, Holger Fr¨oning, and Franz Pernkopf. Training discrete-valued neural networks with sign activations using weight distributions. In European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD), 2019.
  96. 96.David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning representations by back-propagating errors. Nature, 323:533–536, 1986.
  97. 97.Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Inverted residuals and linear bottlenecks: Mobile networks for classification, detection and segmentation. CoRR, abs/1801.04381, Jan 2018a. URL https://arxiv.org/abs/1801.04381.
  98. 98.Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. MobileNetV2: Inverted residuals and linear bottlenecks. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 4510–4520, 2018b.
  99. 99.G¨unther Schindler, Wolfgang Roth, Franz Pernkopf, and Holger Fr¨oning. Parameterized structured pruning for deep neural networks. In 6th International Conference on Machine Learning, Optimization, and Data Science (LOD), 2020.
  100. 100.Oran Shayer, Dan Levi, and Ethan Fetaya. Learning discrete weights using the local reparameterization trick. In International Conference on Learning Representations (ICLR), 2018.
  101. 101.Alexander Shekhovtsov, Viktor Yanush, and Boris Flach. Path sample-analytic gradient estimators for stochastic binary networks. In Advances in Neural Information Processing Systems (NeurIPS), pages 12884–12894, 2020.
  102. 102.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. CoRR, abs/1909.08053, 2019. URL https://arxiv.org/abs/1909.08053.
  103. 103.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations (ICLR), 2015.
  104. 104.Daniel Soudry, Itay Hubara, and Ron Meir. Expectation backpropagation: Parameter-free training of multilayer neural networks with continuous or discrete weights. In Advances in Neural Information Processing Systems (NIPS), pages 963–971, 2014.
  105. 105.Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):1929–1958, 2014.
  106. 106.Dimitrios Stamoulis, Ruizhou Ding, Di Wang, Dimitrios Lymberopoulos, Bodhi Priyantha, Jie Liu, and Diana Marculescu. Single-path NAS: Designing hardware-efficient convnets in less than 4 hours. In European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD), 2019.
  107. 107.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 1–9, 2015.
  108. 108.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 2818–2826, 2016.
  109. 109.Mingxing Tan and Quoc V. Le. Efficientnet: Rethinking model scaling for convolutional neural networks. CoRR, abs/1905.11946, May 2019. URL https://arxiv.org/abs/1905.11946.
  110. 110.Mingxing Tan and Quoc V. Le. Efficientnetv2: Smaller models and faster training. CoRR, abs/2104.00298, Apr 2021.
  111. 111.Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, and Quoc V. Le. MnasNet: Platform-aware neural architecture search for mobile. CoRR, abs/1807.11626, 2018.
  112. 112.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey. ACM Computing Surveys, 55(6), 2022.
  113. 113.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Thibaut Lavril, Timoth´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. CoRR, 2023. URL https://arxiv.org/abs/2302.13971.
  114. 114.Stefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama, Javier Alonso Garc´ıa, Stephen Tiedemann, Thomas Kemp, and Akira Nakamura. Mixed precision DNNs: All you need is a good parametrization. In International Conference on Learning Representations (ICLR), 2020.
  115. 115.Karen Ullrich, Edward Meeds, and Max Welling. Soft weight-sharing for neural network compression. In International Conference on Learning Representations (ICLR), 2017.
  116. 116.Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Heng Wai Leong, Magnus Jahre, and Kees A. Vissers. FINN: A framework for fast, scalable binarized neural network inference. In ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (ISFPGA), pages 65–74, 2017.
  117. 117.Mart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad, Ying Wang, Tijmen Blankevoort, and Max Welling. Bayesian bits: Unifying quantization and pruning. In Advances in Neural Information Processing Systems (NeurIPS), pages 5741–5752, 2020.
  118. 118.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. CoRR, abs/1706.03762, Jun 2017. URL https://arxiv.org/abs/1706.03762.
  119. 119.Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. HAQ: Hardware-aware automated quantization with mixed precision. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 8612–8620, 2019.
  120. 120.Max Welling and Yee Whye Teh. Bayesian learning via stochastic gradient Langevin dynamics. In International Conference on Machine Learning (ICML), pages 681–688, 2011.
  121. 121.Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. Learning structured sparsity in deep neural networks. In Advances in Neural Information Processing Systems (NIPS), pages 2074–2082, 2016.
  122. 122.Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda, and Kurt Keutzer. Mixed precision quantization of convnets via differentiable neural architecture search. CoRR, abs/1812.00090, 2018a. URL http://arxiv.org/abs/1812.00090.
  123. 123.Shuang Wu, Guoqi Li, Feng Chen, and Luping Shi. Training and inference with integers in deep neural networks. In International Conference on Learning Representations (ICLR), 2018b.
  124. 124.Saining Xie, Ross B. Girshick, Piotr Doll´ar, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 5987–5995, 2017.
  125. 125.Zichao Yang, Marcin Moczulski, Misha Denil, Nando de Freitas, Alexander J. Smola, Le Song, and Ziyu Wang. Deep fried convnets. In International Conference on Computer Vision (ICCV), pages 1476–1483, 2015.
  126. 126.Mingzhang Yin and Mingyuan Zhou. ARM: augment-REINFORCE-merge gradient for stochastic binary networks. In International Conference on Learning Representations (ICLR), 2019.
  127. 127.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. In Proceedings of the British Machine Vision Conference (BMVC), 2016.
  128. 128.Xinchuan Zeng and Tony R. Martinez. Using a neural network to approximate an ensemble of classifiers. Neural Processing Letters, 12(3):225–237, 2000.
  129. 129.Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua. LQ-Nets: Learned quantization for highly accurate and compact deep neural networks. In European Conference on Computer Vision (ECCV), pages 373–390, 2018a.
  130. 130.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. ShuffleNet: An extremely efficient convolutional neural network for mobile devices. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 6848–6856, 2018b.
  131. 131.Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu, and Yurong Chen. Incremental network quantization: Towards lossless CNNs with low-precision weights. In International Conference on Learning Representations (ICLR), 2017.
  132. 132.Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou. DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients. CoRR, abs/1606.06160, 2016.
  133. 133.Chenzhuo Zhu, Song Han, Huizi Mao, and William J. Dally. Trained ternary quantization. In International Conference on Learning Representations (ICLR), 2017.
  134. 134.Barret Zoph and Quoc V. Le. Neural architecture search with reinforcement learning. In International Conference on Learning Representations (ICLR), 2017.
  135. 135.Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le. Learning transferable architectures for scalable image recognition. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 8697–8710, 2018.

Citation

MLA
Roth, W., et al. “Resource-Efficient Neural Networks for Embedded Systems”. Journal of Machine Learning Research, vol. 25, no. 50, 2024, pp. 1–1, https://www.jmlr.org/papers/v25/18-566.html.
APA
Roth, W., Schindler, G., Klein, B., Peharz, R., Tschiatschek, S., Fröning, H., Pernkopf, F., & Ghahramani, Z. (2024). Resource-Efficient Neural Networks for Embedded Systems. Journal of Machine Learning Research, 25(50), 1–51. https://www.jmlr.org/papers/v25/18-566.html
Chicago
Roth, W., G. Schindler, B. Klein, et al. 2024. “Resource-Efficient Neural Networks for Embedded Systems”. Journal of Machine Learning Research 25 (50): 1–51. https://www.jmlr.org/papers/v25/18-566.html.
Harvard
Roth, W. et al. (2024) “Resource-Efficient Neural Networks for Embedded Systems”, Journal of Machine Learning Research, 25(50), pp. 1–51. Available at: https://www.jmlr.org/papers/v25/18-566.html.
Vancouver
1. Roth W, Schindler G, Klein B, Peharz R, Tschiatschek S, Fröning H, Pernkopf F, Ghahramani Z (2024) Resource-Efficient Neural Networks for Embedded Systems. Journal of Machine Learning Research 25:1–51

BibTeX

@article{JMLR:v25:18-566,
  author  = {Wolfgang Roth and G{{\"u}}nther Schindler and Bernhard Klein and Robert Peharz and Sebastian Tschiatschek and Holger Fr{{\"o}}ning and Franz Pernkopf and Zoubin Ghahramani},
  title   = {Resource-Efficient Neural Networks for Embedded Systems},
  journal = {Journal of Machine Learning Research},
  year    = {2024},
  volume  = {25},
  number  = {50},
  pages   = {1--51},
  url     = {http://jmlr.org/papers/v25/18-566.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/