AMC: AutoML for Model Compression and Acceleration on Mobile Devices
Yihui HeJi LinZhijian LiuHanrui WangLi-Jia LiSong Han
Introduces a reinforcement learning-based compression framework that automates the exploration of neural network design spaces, outperforming handcrafted heuristics to deliver higher inference speeds on mobile hardware with minimal accuracy loss.
Deploying deep neural networks to resource-constrained edge devices like smartphones, autonomous vehicles, and robotics requires reducing model size and computational demand without sacrificing accuracy. Historically, this model compression relies on manual, rule-based heuristics developed by domain experts. However, because modern deep networks contain complex inter-layer dependencies, human-designed pruning rules are labor-intensive, difficult to transfer across different architectures, and frequently suboptimal.
The article aims to introduce and evaluate AutoML for Model Compression, an automated framework that uses reinforcement learning to determine the optimal compression policy for neural networks across both mobile and server platforms.
The framework automates the compression pipeline by evaluating a network layer by layer using a continuous-control reinforcement learning agent. Rather than requiring expensive model retraining at every search step, the agent evaluates intermediate accuracy on a validation set without fine-tuning, drastically accelerating policy search to about one hour on a single graphics processing unit. The authors evaluated the approach across multiple established computer vision models (including VGG-16, ResNet-50, MobileNet, and MobileNet-V2) using benchmark datasets (CIFAR-10, ImageNet, and PASCAL VOC) and measured latency directly on an Android smartphone and a desktop graphics processor.
The evaluation yielded several key findings. First, the automated approach consistently outperformed human-designed heuristics; for example, it pushed the compression ratio of ResNet-50 on ImageNet from 3.4-fold to 5-fold with zero loss in classification accuracy. Second, on compact architectures like MobileNet, the automated system achieved a 1.81-fold to 1.95-fold real-world speedup on an Android mobile device with only a 0.1% to 0.4% drop in top-1 accuracy, doubling processing speed from 8.1 to 16.0 frames per second. Third, under a four-fold computational reduction on VGG-16, the system delivered 2.7% higher accuracy than rule-based alternatives and generalized effectively to object detection tasks, improving detection accuracy over the uncompressed baseline.
These results demonstrate that automated, learning-based policies can replace labor-intensive manual tuning while unlocking superior performance, lower power consumption, and smaller memory footprints on edge hardware. By supporting direct optimization for specific hardware constraints (such as measured mobile latency rather than theoretical computation counts), organizations can significantly shorten engineering cycles and accelerate the deployment of advanced artificial intelligence capabilities to mobile applications.
Engineering and product teams deploying computer vision models on edge hardware should adopt automated, continuous reinforcement learning pipelines in place of manual pruning rules. Depending on deployment priorities, teams can configure the pipeline for resource-constrained targets to maximize speed within strict hardware budgets, or accuracy-guaranteed targets to shrink models without compromising performance.
While the findings show strong transferability across vision tasks and standard hardware, confidence in deployment depends on target-specific profiling. The framework was evaluated primarily on standard convolutional architectures and computer vision benchmarks; organizations deploying non-vision architectures, proprietary edge accelerators, or novel hardware environments should run pilot validations to verify real-world latency and accuracy before broad rollout.
- Paper: Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding, Song Han et al. (2015). This foundational paper establishes the standard multi-stage neural network compression pipeline—combining pruning, quantization, and coding—that AMC automates and optimizes via reinforcement learning.
- Paper: Learning both Weights and Connections for Efficient Neural Networks, Song Han et al. (2015). It introduces iterative threshold-based weight pruning, providing the foundational compression primitives and motivation for automated policy search in AMC.
- Paper: MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications, Andrew G. Howard et al. (2017). It introduces MobileNets and depthwise separable convolutions, which serve as a primary target architecture and baseline for AMC's automated compression experiments.
- Paper: Channel Pruning for Accelerating Very Deep Neural Networks, Yihui He et al. (2017). It details channel pruning techniques for convolutional layers that AMC replaces with reinforcement learning-driven per-layer compression policies.
- Paper: Pruning Filters for Efficient ConvNets, Hao Li et al. (2016). It introduces structured filter pruning for convolutional networks based on weight norms, offering the baseline structured pruning paradigm automated by AMC.
- Paper: Learning Efficient Convolutional Networks through Network Slimming, Zhuang Liu et al. (2017). It provides a rule-based channel pruning method using batch normalization scaling factors, serving as an important heuristic baseline that AMC seeks to outperform.
- Paper: Designing Neural Network Architectures using Reinforcement Learning, Bowen Baker et al. (2016). It demonstrates using reinforcement learning to design neural architectures sequentially, providing the algorithmic foundation for applying RL agents to model design decisions.
- Paper: ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression, Jian-Hao Luo et al. (2017). It outlines a greedy filter-level pruning framework for VGG-16, representing the hand-crafted, layer-by-layer compression heuristics AMC replaces with an automated agent.
- Paper: ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware, Han Cai et al. (2018). It extends hardware-aware automated optimization by performing direct neural architecture search and latency optimization on target devices using path binarization.
- Paper: Once for All: Train One Network and Specialize it for Efficient Deployment, Han Cai et al. (2019). It generalizes hardware-specialized model optimization by training a single once-for-all supernet that yields diverse sub-networks tailored to multiple device constraints without separate retraining.
- Paper: Searching for MobileNetV3, Andrew Howard et al. (2019). It combines platform-aware neural architecture search with per-layer automated network adaptation algorithms to build optimized mobile vision models.
- Paper: Rethinking the Value of Network Pruning, Zhuang Liu et al. (2019). It critically re-examines the mechanisms underlying structured pruning pipelines, investigating whether automated pruning algorithms primarily function as architecture search.
- Paper: DepGraph: Towards Any Structural Pruning, Gongfan Fang et al. (2023). It generalizes automated structural pruning to arbitrary neural network topologies by automatically building dependency graphs across diverse architectures.
- Paper: EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks, Mingxing Tan et al. (2019). It builds upon automated mobile architecture design principles to systematically scale depth, width, and resolution under fixed computational budgets.
- Paper: AutoML: A Survey of the State-of-the-Art, Xin He et al. (2019). It provides a broad retrospective survey of automated machine learning techniques, contextualizing reinforcement learning-based model search and compression within the wider AutoML landscape.
