Resource-Efficient Neural Networks for Embedded Systems
Wolfgang RothGünther SchindlerBernhard KleinRobert PeharzSebastian TschiatschekHolger FröningFranz PernkopfZoubin Ghahramani
Presents a systematic review and empirical analysis of deep neural network optimization techniques, including quantization, pruning, and structural efficiency, to guide the trade-offs between prediction accuracy, latency, and energy consumption across embedded CPUs, GPUs, and FPGAs.
Modern deep learning achieves outstanding performance across computer vision, speech, and natural language processing, but its high memory and computational demands prevent practical deployment on resource-constrained embedded systems, autonomous platforms, and edge devices. As machine learning transitions into real-world applications with strict limits on power, latency, and hardware capacity, finding effective trade-offs between predictive accuracy and operational resource efficiency has become essential.
The article systematically reviews algorithmic techniques for resource-efficient deep neural network inference and evaluates how these compression methods interact with embedded hardware platforms. The authors categorize optimization methods into three primary domains—quantization, network pruning, and structural efficiency—and analyze their real-world impact through benchmarking on embedded central processing units, graphics processing units, and field-programmable gate arrays using standardized image classification datasets.
The evaluation reveals several crucial operational findings. First, prediction accuracy is significantly more sensitive to the quantization of activations than to the quantization of weights. While reducing weight precision down to one or two bits causes moderate accuracy degradation, preserving two to four bits for activations is essential to avoid severe performance drops. Second, structured channel pruning consistently outperforms kernel and grouped convolution pruning when evaluating total memory footprint; although grouped convolutions drastically reduce theoretical floating-point operations and parameter counts, they fail to reduce activation memory and do not improve inference speed on target devices. Third, hardware architecture heavily dictates compression benefits. On general-purpose embedded central processing units, low-bit quantization fails to improve throughput due to instruction-level overhead and efficient native floating-point units, whereas structured pruning preserves speed. Conversely, custom data-flow architectures on field-programmable gate arrays require low-bit quantization to fit entire models into limited on-chip memory, enabling extremely high throughput, while embedded graphics processing units achieve the best balance of programmability, high throughput, and accuracy when paired with structured channel pruning.
These findings demonstrate that theoretical efficiency metrics, such as parameter counts and operation counts, do not accurately predict real-world execution latency, memory footprint, or energy consumption. Algorithmic compression strategies cannot be designed in isolation from the deployment hardware. Misaligning compression techniques with processor architectures can degrade predictive performance without delivering any throughput or energy advantages.
Organizations deploying deep learning models on edge devices should tailor compression techniques directly to the target hardware platform: structured channel pruning is best suited for embedded central processing units and graphics processing units, whereas aggressive mixed-precision quantization is necessary for data-flow architectures such as field-programmable gate arrays. Furthermore, engineering teams should prioritize the reduction of intermediate activation memory over theoretical floating-point operations to achieve meaningful latency improvements.
Decision-makers should note that the empirical evaluations in the article focus on vision classification benchmarks and fixed 5-Watt power envelopes, and specialized domain-specific accelerators were omitted due to toolchain and flexibility constraints. Despite these boundaries, the evidence provides high confidence that joint algorithmic-hardware optimization is required for effective edge AI deployment.
- Paper: Efficient Processing of Deep Neural Networks: A Tutorial and Survey, Vivienne Sze et al. (2017). This comprehensive tutorial establishes the foundational hardware and algorithmic trade-offs of deep neural network efficiency that underpin modern embedded inference techniques.
- Paper: Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding, Song Han et al. (2015). This seminal paper introduces the core three-stage compression pipeline combining pruning, trained quantization, and encoding upon which subsequent resource-efficient deep learning builds.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). It provides a rigorous taxonomy and theoretical foundation for low-precision integer quantization methods essential for the survey's quantization section.
- Paper: MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications, Andrew G. Howard et al. (2017). This paper establishes depthwise separable convolutions as the foundational design paradigm for structural efficiency in mobile and embedded neural networks.
- Paper: Learning both Weights and Connections for Efficient Neural Networks, Song Han et al. (2015). It introduces iterative magnitude-based connection pruning, providing the foundational technique for network sparsification discussed throughout the survey.
- Paper: Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, Benoit Jacob et al. (2018). This work formulates integer-only quantization and quantization-aware training for practical deployment on embedded and mobile processors.
- Paper: Pruning Filters for Efficient ConvNets, Hao Li et al. (2016). It details structured filter-level pruning to accelerate convolutional neural networks on standard hardware without requiring custom sparse libraries.
- Paper: MobileNetV2: Inverted Residuals and Linear Bottlenecks, Mark Sandler et al. (2018). This architecture introduces inverted residual blocks and linear bottlenecks, establishing standard structural efficiency concepts for on-device inference.
- Paper: Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations, Itay Hubara et al. (2016). It explores extreme low-precision arithmetic by training networks with binary weights and activations using bitwise operations.
- Paper: Learning Structured Sparsity in Deep Neural Networks, Wei Wen et al. (2016). This paper develops structured sparsity learning via group Lasso to regularize hardware-friendly architectural dimensions during training.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). This survey extends resource efficiency principles and structural acceleration paradigms from classic deep neural networks to modern large language model architectures.
- Paper: Sparser, Faster, Lighter Transformer Language Models, Edoardo Cetin et al. (2026). It applies activation sparsity and hardware-optimized execution to transformer language models, building directly upon core neural network compression concepts.
- Paper: Scaling Laws for Fine-Grained Mixture of Experts, Jan Ludziejewski et al. (2024). This work advances structural efficiency and conditional computation by formulating scaling laws for fine-grained mixture-of-experts models.
