AMC: AutoML for Model Compression and Acceleration on Mobile Devices

Yihui HeJi LinZhijian LiuHanrui WangLi-Jia LiSong Han

article2018ECCV1,518 citations

Introduces a reinforcement learning-based compression framework that automates the exploration of neural network design spaces, outperforming handcrafted heuristics to deliver higher inference speeds on mobile hardware with minimal accuracy loss.

Listen

Deploying deep neural networks to resource-constrained edge devices like smartphones, autonomous vehicles, and robotics requires reducing model size and computational demand without sacrificing accuracy. Historically, this model compression relies on manual, rule-based heuristics developed by domain experts. However, because modern deep networks contain complex inter-layer dependencies, human-designed pruning rules are labor-intensive, difficult to transfer across different architectures, and frequently suboptimal.

The article aims to introduce and evaluate AutoML for Model Compression, an automated framework that uses reinforcement learning to determine the optimal compression policy for neural networks across both mobile and server platforms.

The framework automates the compression pipeline by evaluating a network layer by layer using a continuous-control reinforcement learning agent. Rather than requiring expensive model retraining at every search step, the agent evaluates intermediate accuracy on a validation set without fine-tuning, drastically accelerating policy search to about one hour on a single graphics processing unit. The authors evaluated the approach across multiple established computer vision models (including VGG-16, ResNet-50, MobileNet, and MobileNet-V2) using benchmark datasets (CIFAR-10, ImageNet, and PASCAL VOC) and measured latency directly on an Android smartphone and a desktop graphics processor.

The evaluation yielded several key findings. First, the automated approach consistently outperformed human-designed heuristics; for example, it pushed the compression ratio of ResNet-50 on ImageNet from 3.4-fold to 5-fold with zero loss in classification accuracy. Second, on compact architectures like MobileNet, the automated system achieved a 1.81-fold to 1.95-fold real-world speedup on an Android mobile device with only a 0.1% to 0.4% drop in top-1 accuracy, doubling processing speed from 8.1 to 16.0 frames per second. Third, under a four-fold computational reduction on VGG-16, the system delivered 2.7% higher accuracy than rule-based alternatives and generalized effectively to object detection tasks, improving detection accuracy over the uncompressed baseline.

These results demonstrate that automated, learning-based policies can replace labor-intensive manual tuning while unlocking superior performance, lower power consumption, and smaller memory footprints on edge hardware. By supporting direct optimization for specific hardware constraints (such as measured mobile latency rather than theoretical computation counts), organizations can significantly shorten engineering cycles and accelerate the deployment of advanced artificial intelligence capabilities to mobile applications.

Engineering and product teams deploying computer vision models on edge hardware should adopt automated, continuous reinforcement learning pipelines in place of manual pruning rules. Depending on deployment priorities, teams can configure the pipeline for resource-constrained targets to maximize speed within strict hardware budgets, or accuracy-guaranteed targets to shrink models without compromising performance.

While the findings show strong transferability across vision tasks and standard hardware, confidence in deployment depends on target-specific profiling. The framework was evaluated primarily on standard convolutional architectures and computer vision benchmarks; organizations deploying non-vision architectures, proprietary edge accelerators, or novel hardware environments should run pilot validations to verify real-world latency and accuracy before broad rollout.

Cover for AMC: AutoML for Model Compression and Acceleration on Mobile Devices

Abstract

Model compression is a critical technique to efficiently deploy neural network models on mobile devices which have limited computation resources and tight power budgets. Conventional model compression techniques rely on hand-crafted heuristics and rule-based policies that require domain experts to explore the large design space trading off among model size, speed, and accuracy, which is usually sub-optimal and time-consuming. In this paper, we propose AutoML for Model Compression (AMC) which leverage reinforcement learning to provide the model compression policy. This learning-based compression policy outperforms conventional rule-based compression policy by having higher compression ratio, better preserving the accuracy and freeing human labor. Under 4x FLOPs reduction, we achieved 2.7% better accuracy than the handcrafted model compression policy for VGG-16 on ImageNet. We applied this automated, push-the-button compression pipeline to MobileNet and achieved 1.81x speedup of measured inference latency on an Android phone and 1.43x speedup on the Titan XP GPU, with only 0.1% loss of ImageNet Top-1 accuracy.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Problem Definition
  • 3.2 Automated Compression with Reinforcement Learning
  • The State Space
  • The Action Space
  • DDPG Agent
  • 3.3 Search Protocols
  • Resource-Constrained Compression
  • Accuracy-Guaranteed Compression
  • 4 Experimental Results
  • 4.1 CIFAR-10 and Analysis
  • FLOPs-Constrained Compression.
  • Accuracy-Guaranteed Compression.
  • Speedup policy exploration.
  • 4.2 ImageNet
  • Push the Limit of Fine-grained Pruning.
  • Comparison with Heuristic Channel Reduction.
  • Speedup Mobile Inference.
  • Generalization Ability.
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — AutoML for Model Compression (AMC) Framework

    model/method

    AutoML for Model Compression (AMC) is an automated framework that uses reinforcement learning (RL) to determine layer-wise compression ratios for deep neural networks. Rather than relying on rule-based heuristics or manual hyperparameter tuning across layers, AMC processes a pre-trained convolutional neural network sequentially from the first layer L1L_1 to the final layer LTL_T as a Markov Decision Process (MDP).

    At each step tt, a continuous-control Deep Deterministic Policy Gradient (DDPG) agent observes a layer feature embedding state sts_t, outputs a continuous action at∈(0,1]a_t \in (0, 1] corresponding to the sparsity ratio for layer LtL_t, compresses the layer, and transitions to state st+1s_{t+1}. After all TT layers are compressed, the full compressed network is evaluated on a validation set without fine-tuning to produce a scalar reward RR that reflects both accuracy and hardware efficiency. The agent's policy is updated using this terminal reward. Once policy search finishes, the single best-performing pruned architecture is fine-tuned to restore accuracy.

  2. Knowl 2 — Layer Embedding State Space Representation in AMC

    equation

    For each convolutional or fully connected layer tt in a network with TT layers, the state sts_t provided to the reinforcement learning agent is an 11-dimensional feature vector:

    st=(t,n,c,h,w,stride,k,FLOPs[t],reduced,rest,at−1)s_t = (t, n, c, h, w, \text{stride}, k, \text{FLOPs}[t], \text{reduced}, \text{rest}, a_{t-1})

    where:

    • tt is the layer index (1≤t≤T1 \le t \le T).
    • nn is the number of output channels.
    • cc is the number of input channels.
    • hh and ww are the spatial height and width of the input feature map.
    • stride\text{stride} is the convolution stride.
    • kk is the kernel size (for an n×c×k×kn \times c \times k \times k weight tensor).
    • FLOPs[t]\text{FLOPs}[t] is the computational cost of layer LtL_t in floating-point operations.
    • reduced\text{reduced} is the cumulative number of FLOPs reduced in layers 1,…,t−11, \dots, t-1.
    • rest\text{rest} is the total baseline FLOPs of the remaining layers t+1,…,Tt+1, \dots, T.
    • at−1a_{t-1} is the compression action taken for the previous layer Lt−1L_{t-1}.

    All state features are normalized to [0,1][0, 1] before being fed to the agent network.

  3. Knowl 3 — Continuous Action Policy and Critic Optimization with Baseline Subtraction

    equation

    AMC employs the continuous Deep Deterministic Policy Gradient (DDPG) actor-critic algorithm to output a continuous layer compression ratio at∈(0,1]a_t \in (0, 1]. To balance exploration and exploitation, exploration noise is drawn from a truncated normal distribution:

    μ′(st)∼TN(μ(st∣θμ),σ2,0,1)\mu'(s_t) \sim \text{TN}\left(\mu(s_t \mid \theta^\mu), \sigma^2, 0, 1\right)

    where μ(st∣θμ)\mu(s_t \mid \theta^\mu) is the deterministic actor network parameterized by θμ\theta^\mu, TN(⋅,σ2,0,1)\text{TN}(\cdot, \sigma^2, 0, 1) denotes a normal distribution truncated within [0,1][0, 1], and the noise standard deviation σ\sigma is initialized to 0.50.5 and decayed exponentially across episodes.

    Each transition in an episode is represented as (st,at,R,st+1)(s_t, a_t, R, s_{t+1}), where RR is the global reward returned after compressing the entire network. The critic network Q(s,a∣θQ)Q(s, a \mid \theta^Q) is trained by minimizing the mean squared Bellman error:

    Loss=1N∑i(yi−Q(si,ai∣θQ))2\text{Loss} = \frac{1}{N} \sum_{i} \left( y_i - Q(s_i, a_i \mid \theta^Q) \right)^2

    yi=ri−b+γQ(si+1,μ(si+1)∣θQ)y_i = r_i - b + \gamma Q\left(s_{i+1}, \mu(s_{i+1}) \mid \theta^Q\right)

    where NN is the mini-batch size, the discount factor is set to γ=1\gamma = 1 to avoid over-prioritizing short-term layer rewards, and bb is a baseline reward subtracted to reduce gradient variance, computed as an exponential moving average of previous episode rewards.

  4. Knowl 4 — Resource-Constrained Compression Protocol and Reward

    model/method

    The resource-constrained compression protocol is designed for latency- and resource-critical applications where a model must fit within a fixed budget of FLOPs, parameters, or on-device inference latency. Under this protocol, the RL agent receives the negative validation classification error as its reward:

    Rerr=−ErrorR_{\text{err}} = -\text{Error}

    Because RerrR_{\text{err}} provides no intrinsic gradient for reducing resources beyond the target threshold, resource satisfaction is enforced by dynamically bounding the allowable action space (at)(a_t) during sequential generation. The agent can select arbitrary actions in early layers, but once the cumulative budget requires maximal compression on all subsequent layers to meet the global target, the action range is clamped. This guarantees that every generated model strictly satisfies the resource budget while the policy learns layer-wise allocations that minimize error.

  5. Knowl 5 — Action Bounding for Resource-Constrained Compression

    algorithm

    Algorithm for bounding the sparsity action ata_t for layer LtL_t to guarantee that the final pruned model meets a target global sparsity ratio α\alpha:

    Input: Layer index tt, state sts_t, layer parameter counts W1,…,WTW_1, \dots, W_T, target global sparsity α∈[0,1]\alpha \in [0, 1], maximum allowed layer sparsity amax⁡a_{\max}
    Output: Bounded sparsity action actiont\text{action}_t for layer LtL_t
    if t=1t = 1 then
        Wreduced←0W_{\text{reduced}} \leftarrow 0
    end if
    actiont←μ′(st)\text{action}_t \leftarrow \mu'(s_t)
    actiont←min⁡(actiont,amax⁡)\text{action}_t \leftarrow \min(\text{action}_t, a_{\max})
    Wall←∑k=1TWkW_{\text{all}} \leftarrow \sum_{k=1}^T W_k
    Wrest←∑k=t+1TWkW_{\text{rest}} \leftarrow \sum_{k=t+1}^T W_k
    Wduty←α⋅Wall−amax⁡⋅Wrest−WreducedW_{\text{duty}} \leftarrow \alpha \cdot W_{\text{all}} - a_{\max} \cdot W_{\text{rest}} - W_{\text{reduced}}
    actiont←max⁡(actiont,WdutyWt)\text{action}_t \leftarrow \max\left(\text{action}_t, \frac{W_{\text{duty}}}{W_t}\right)
    Wreduced←Wreduced+actiont⋅WtW_{\text{reduced}} \leftarrow W_{\text{reduced}} + \text{action}_t \cdot W_t
    return actiont\text{action}_t

    The algorithm computes WdutyW_{\text{duty}}, which is the parameter reduction that layer LtL_t must supply if all future layers k>tk > t are pruned at maximum aggressiveness amax⁡a_{\max}. If the proposed action actiont\text{action}_t provides less reduction than Wduty/WtW_{\text{duty}} / W_t, it is bounded upward to ensure the global reduction budget αWall\alpha W_{\text{all}} is mathematically reachable.

  6. Knowl 6 — Accuracy-Guaranteed Compression Protocol and Reward

    equation

    The accuracy-guaranteed compression protocol is designed to find the maximum possible compression ratio without degrading baseline model accuracy. Under this protocol, the action space is unconstrained, and the RL agent is driven by a composite reward function sensitive to validation classification error while offering a continuous incentive for resource reduction:

    RFLOPs=−Error⋅log⁡(FLOPs)R_{\text{FLOPs}} = -\text{Error} \cdot \log(\text{FLOPs})

    RParam=−Error⋅log⁡(#Param)R_{\text{Param}} = -\text{Error} \cdot \log(\#\text{Param})

    where Error∈[0,1]\text{Error} \in [0, 1] is the top-1 validation error, FLOPs\text{FLOPs} is the total floating-point operation count of the compressed network, and #Param\#\text{Param} is the total count of remaining non-zero parameters. Because classification error is empirically inversely proportional to the logarithm of FLOPs or parameter count, this reward enables the policy to automatically discover the compression boundary where model size is minimized with negligible accuracy loss.

  7. Knowl 7 — Validation Accuracy without Fine-Tuning as an Exploration Proxy

    empirical result

    To eliminate the substantial computational cost of fine-tuning candidate models during RL search, AMC uses validation accuracy measured immediately after pruning (pre-fine-tuning) as the reward signal. Experiments on CIFAR-10 and ImageNet demonstrate a strong monotonic correlation between the pre-fine-tuning validation accuracy and post-fine-tuning accuracy across varied compression policies.

    Evaluating pre-fine-tuning accuracy on a small validation subset (5,000 images on CIFAR-10, 3,000 images from the training split on ImageNet) provides an effective surrogate for the final converged model quality. This proxy enables the reinforcement learning search to complete within 1 hour on a single NVIDIA GeForce GTX TITAN Xp GPU for CIFAR-10 networks.

  8. Knowl 8 — CIFAR-10 Pruning Policy Benchmarks

    data/table

    Comparison of AMC against rule-based heuristic pruning policies (uniform, shallow, deep) on CIFAR-10 for Plain-20, ResNet-56, and ResNet-50. RErrR_{\text{Err}} denotes FLOPs-constrained channel pruning, and RParamR_{\text{Param}} denotes accuracy-guaranteed fine-grained pruning.

    Model Policy Ratio Val Acc. (%) Test Acc. (%) Acc. after FT (%)
    Plain-20 (90.5%) deep (handcraft) 50% FLOPs 79.6 79.2 88.3
    Plain-20 (90.5%) shallow (handcraft) 50% FLOPs 83.2 82.9 89.2
    Plain-20 (90.5%) uniform (handcraft) 50% FLOPs 84.0 83.9 89.7
    Plain-20 (90.5%) AMC (RErrR_{\text{Err}}) 50% FLOPs 86.4 86.0 90.2
    ResNet-56 (92.8%) uniform (handcraft) 50% FLOPs 87.5 87.4 89.8
    ResNet-56 (92.8%) deep (handcraft) 50% FLOPs 88.4 88.4 91.5
    ResNet-56 (92.8%) AMC (RErrR_{\text{Err}}) 50% FLOPs 90.2 90.1 91.9
    ResNet-50 (93.53%) AMC (RParamR_{\text{Param}}) 60% Params 93.64 93.55 -

    AMC achieves 90.2% post-fine-tuning test accuracy on Plain-20 and 91.9% on ResNet-56 under a 50% FLOPs budget, outperforming handcrafted heuristic policies by 0.5%–1.9%. The learned pruning policy produces a sawtooth pattern across layers that mirrors a bottleneck architecture. On ResNet-50, accuracy-guaranteed AMC achieves a 60% parameter reduction while improving test accuracy from 93.53% to 93.55% without fine-tuning.

  9. Knowl 9 — ImageNet Channel Pruning on VGG-16 and MobileNets

    data/table

    Comparison of AMC with handcrafted and rule-based channel pruning methods on ImageNet across VGG-16, MobileNet, and MobileNet-V2. Baseline Top-1 accuracies: VGG-16 = 70.5%, MobileNet = 70.6%, MobileNet-V2 = 71.8%.

    Model Policy FLOPs Remaining (%) ΔAcc\Delta\text{Acc} Top-1 (%)
    VGG-16 FP (handcraft) 20% -14.6
    VGG-16 RNP (handcraft) 20% -3.58
    VGG-16 SPP (handcraft) 20% -2.3
    VGG-16 CP (handcraft) 20% -1.7
    VGG-16 AMC (ours) 20% -1.4
    MobileNet uniform (0.75-224) 56% -2.5
    MobileNet AMC (ours) 50% -0.4
    MobileNet uniform (0.75-192) 41% -3.7
    MobileNet AMC (ours) 40% -1.7
    MobileNet-V2 uniform (0.75-224) 70% -2.0
    MobileNet-V2 AMC (ours) 70% -1.0

    On VGG-16 at 5×5\times FLOPs reduction (20% FLOPs remaining), AMC yields a top-1 accuracy loss of −1.4%-1.4\%, outperforming expert-designed channel pruning (CP, −1.7%-1.7\%) and other heuristics (FP, RNP, SPP). On MobileNet at 50% FLOPs, AMC loses only 0.4% top-1 accuracy, outperforming the 0.75 width multiplier baseline (−2.5%-2.5\% at 56% FLOPs). On MobileNet-V2, AMC achieves a 1.0% accuracy improvement over the uniform multiplier baseline at 70% FLOPs.

  10. Knowl 10 — On-Device Mobile Latency and Speedup on MobileNet

    data/table

    Performance of AMC applied to MobileNet with FLOPs-constrained and latency-constrained search protocols. Latency is measured on a Google Pixel 1 phone (Qualcomm Snapdragon 821 SoC, TensorFlow Lite, batch size 1) and an NVIDIA Titan XP GPU (batch size 50) using 224×224224 \times 224 input resolution.

    Model MACs
    (×106\times 10^6)
    Top-1
    Acc. (%)
    Top-5
    Acc. (%)
    GPU Latency
    / Speed
    Android Latency
    / Speed
    Model
    Memory
    100% MobileNet 569 70.6% 89.5% 0.46 ms / 2191 fps 123.3 ms / 8.1 fps 20.1 MB
    75% MobileNet 325 68.4% 88.2% 0.34 ms / 2944 fps 72.3 ms / 13.8 fps 14.8 MB
    NetAdapt - 69.8% - - / - 70.0 ms / 14.3 fps -
    AMC (50% FLOPs) 285 70.5% 89.3% 0.32 ms / 3127 fps (1.43×1.43\times) 68.3 ms / 14.6 fps (1.81×1.81\times) 14.3 MB
    AMC (50% Latency) 272 70.2% 89.2% 0.30 ms / 3350 fps (1.53×1.53\times) 63.3 ms / 16.0 fps (1.95×1.95\times) 13.2 MB

    Targeting 50% latency directly allows AMC to attain a 1.95×1.95\times speedup on the Android phone (reducing latency from 123.3 ms to 63.3 ms; throughput from 8.1 fps to 16.0 fps) and a 1.53×1.53\times speedup on GPU, with only a 0.4% drop in Top-1 accuracy (70.2% vs. 70.6%). AMC achieves a 2.01×2.01\times speedup specifically on 1×11 \times 1 convolutions, while depth-wise convolutions exhibit less speedup due to their lower computation-to-communication ratio.

  11. Knowl 11 — Layer Sparsity Allocation in ResNet-50 Fine-Grained Pruning

    empirical result

    In a 4-iteration fine-grained pruning schedule on ResNet-50 (targeting overall model densities of 50%, 35%, 25%, and 20%), the AMC reinforcement learning policy discovers a distinctive structural allocation:

    1. 3×33 \times 3 convolutional layers receive higher sparsity (lower density), indicating greater redundancy.
    2. 1×11 \times 1 convolutional layers receive lower sparsity (higher density), indicating lower parameter redundancy.

    By following this learned policy, AMC pushes the compression ratio of ResNet-50 on ImageNet to 5×5\times (20% non-zero parameters) without loss of accuracy (Top-1: 76.11%, Top-5: 92.89%, compared to baseline Top-1: 76.13%, Top-5: 92.86%). In contrast, handcrafted human-expert pruning achieved a 3.4×3.4\times compression ratio (29% density) on the same architecture.

  12. Knowl 12 — Transferability of AMC-Compressed Models to Object Detection

    data/table

    Evaluation of Faster R-CNN on the PASCAL VOC 2007 object detection task using VGG-16 backbones compressed via AMC versus handcrafted pruning baselines:

    Backbone Model mAP (%) mAP [0.5, 0.95] (%)
    Baseline VGG-16 68.7 36.7
    2×2\times handcrafted (He et al., 2017) 68.3 (-0.4) 36.7 (-0.0)
    4×4\times handcrafted (He et al., 2017) 66.9 (-1.8) 35.1 (-1.6)
    4×4\times handcrafted (Zhang et al., 2016) 67.8 (-0.9) 36.5 (-0.2)
    4×4\times AMC (ours) 68.8 (+0.1) 37.2 (+0.5)

    Using a 4×4\times compressed VGG-16 backbone pruned with AMC, Faster R-CNN achieves 68.8% mAP and 37.2% mAP@[0.5, 0.95] on PASCAL VOC 2007. This outperforms 4×4\times handcrafted baselines by 1.0%–1.9% mAP and surpasses the unpruned baseline by 0.1% mAP (and +0.5% mAP@[0.5, 0.95]), demonstrating that the sparsity distribution discovered on classification tasks generalizes effectively to downstream object detection.

Coverage note — None was omitted; all key components of AMC, including MDP formulations, state representations, policy updates, action bounding algorithms, proxy evaluations, and empirical benchmarks on CIFAR-10, ImageNet, mobile latency, and object detection transfer, are captured.

References

  1. 1.Anwar, S., Sung, W.: Compact deep convolutional neural networks with coarse pruning. arXiv preprint arXiv:1610.09639 (2016)
  2. 2.Ashok, A., Rhinehart, N., Beainy, F., Kitani, K.M.: N2n learning: Network to network compression via policy gradient reinforcement learning. arXiv preprint arXiv:1709.06030 (2017)
  3. 3.Bagherinezhad, H., Rastegari, M., Farhadi, A.: Lcnn: Lookup-based convolutional neural network. arXiv preprint arXiv:1611.06473 (2016)
  4. 4.Baker, B., Gupta, O., Naik, N., Raskar, R.: Designing neural network architectures using reinforcement learning. arXiv preprint arXiv:1611.02167 (2016)
  5. 5.Brock, A., Lim, T., Ritchie, J.M., Weston, N.: Smash: one-shot model architecture search through hypernetworks. arXiv preprint arXiv:1708.05344 (2017)
  6. 6.Cai, H., Chen, T., Zhang, W., Yu, Y., Wang, J.: Reinforcement learning for architecture search by network transformation. arXiv preprint arXiv:1707.04873 (2017)
  7. 7.Canziani, A., Paszke, A., Culurciello, E.: An analysis of deep neural network models for practical applications. arXiv preprint arXiv:1605.07678 (2016)
  8. 8.Chen, T., Goodfellow, I., Shlens, J.: Net2net: Accelerating learning via knowledge transfer. arXiv preprint arXiv:1511.05641 (2015)
  9. 9.Chollet, F.: Xception: Deep learning with depthwise separable convolutions. arXiv preprint arXiv:1610.02357 (2016)
  10. 10.Courbariaux, M., Bengio, Y.: Binarynet: Training deep neural networks with weights and activations constrained to+ 1 or-1. arXiv preprint arXiv:1602.02830 (2016)
  11. 11.Denton, E.L., Zaremba, W., Bruna, J., LeCun, Y., Fergus, R.: Exploiting linear structure within convolutional networks for efficient evaluation. In: Advances in Neural Information Processing Systems. pp. 1269–1277 (2014)
  12. 12.Dong, X., Huang, J., Yang, Y., Yan, S.: More is less: A more complicated network with less inference complexity. arXiv preprint arXiv:1703.08651 (2017)
  13. 13.Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A.: The PAS-CAL Visual Object Classes Challenge 2007 (VOC2007) Results. http://www.pascal-network.org/challenges/VOC/voc2007/workshop/index.html
  14. 14.Girshick, R.: Fast r-cnn. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 1440–1448 (2015)
  15. 15.Gong, Y., Liu, L., Yang, M., Bourdev, L.: Compressing deep convolutional networks using vector quantization. arXiv preprint arXiv:1412.6115 (2014)
  16. 16.Han, S.: Efficient methods and hardware for deep learning, https://stacks.stanford.edu/file/druid:qf934gh3708/EFFICIENT%20METHODS%20AND%20HARDWARE%20FOR%20DEEP%20LEARNING-augmented.pdf
  17. 17.Han, S., Kang, J., Mao, H., Hu, Y., Li, X., Li, Y., Xie, D., Luo, H., Yao, S., Wang, Y., et al.: Ese: Efficient speech recognition engine with sparse lstm on fpga. In: Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays. pp. 75–84. ACM (2017)
  18. 18.Han, S., Liu, X., Mao, H., Pu, J., Pedram, A., Horowitz, M.A., Dally, W.J.: Eie: efficient inference engine on compressed deep neural network. In: Proceedings of the 43rd International Symposium on Computer Architecture. pp. 243–254. IEEE Press (2016)
  19. 19.Han, S., Mao, H., Dally, W.J.: Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149 (2015)
  20. 20.Han, S., Pool, J., Tran, J., Dally, W.: Learning both weights and connections for efficient neural network. In: Advances in Neural Information Processing Systems. pp. 1135–1143 (2015)
  21. 21.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 770–778 (2016)
  22. 22.He, Y., Zhang, X., Sun, J.: Channel pruning for accelerating very deep neural networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 1389–1397 (2017)
  23. 23.Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., Adam, H.: Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017)
  24. 24.Hu, H., Peng, R., Tai, Y.W., Tang, C.K.: Network trimming: A data-driven neuron pruning approach towards efficient deep architectures. arXiv preprint arXiv:1607.03250 (2016)
  25. 25.Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167 (2015)
  26. 26.Jaderberg, M., Vedaldi, A., Zisserman, A.: Speeding up convolutional neural networks with low rank expansions. arXiv preprint arXiv:1405.3866 (2014)
  27. 27.Kim, Y.D., Park, E., Yoo, S., Choi, T., Yang, L., Shin, D.: Compression of deep convolutional neural networks for fast and low power mobile applications. arXiv preprint arXiv:1511.06530 (2015)
  28. 28.Krizhevsky, A., Hinton, G.: Learning multiple layers of features from tiny images (2009)
  29. 29.Lavin, A.: Fast algorithms for convolutional neural networks. arXiv preprint arXiv:1509.09308 (2015)
  30. 30.Lebedev, V., Ganin, Y., Rakhuba, M., Oseledets, I., Lempitsky, V.: Speeding-up convolutional neural networks using fine-tuned cp-decomposition. arXiv preprint arXiv:1412.6553 (2014)
  31. 31.Li, H., Kadav, A., Durdanovic, I., Samet, H., Graf, H.P.: Pruning filters for efficient convnets. arXiv preprint arXiv:1608.08710 (2016)
  32. 32.Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., Wierstra, D.: Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 (2015)
  33. 33.Lin, J., Rao, Y., Lu, J.: Runtime neural pruning. In: Advances in Neural Information Processing Systems. pp. 2178–2188 (2017)
  34. 34.Luo, J.H., Wu, J., Lin, W.: Thinet: A filter level pruning method for deep neural network compression. arXiv preprint arXiv:1707.06342 (2017)
  35. 35.Masana, M., van de Weijer, J., Herranz, L., Bagdanov, A.D., Alvarez, J.M.: Domain-adaptive deep network compression. In: The IEEE International Conference on Computer Vision (ICCV) (Oct 2017)
  36. 36.Mathieu, M., Henaff, M., LeCun, Y.: Fast training of convolutional networks through ffts. arXiv preprint arXiv:1312.5851 (2013)
  37. 37.Miikkulainen, R., Liang, J., Meyerson, E., Rawal, A., Fink, D., Francon, O., Raju, B., Navruzyan, A., Duffy, N., Hodjat, B.: Evolving deep neural networks. arXiv preprint arXiv:1703.00548 (2017)
  38. 38.Molchanov, P., Tyree, S., Karras, T., Aila, T., Kautz, J.: Pruning convolutional neural networks for resource efficient transfer learning. CoRR, abs/1611.06440 (2016)
  39. 39.Parashar, A., Rhu, M., Mukkara, A., Puglielli, A., Venkatesan, R., Khailany, B., Emer, J., Keckler, S., Dally, W.J.: Scnn: An accelerator for compressed-sparse convolutional neural networks. In: 44th International Symposium on Computer Architecture (2017)
  40. 40.Polyak, A., Wolf, L.: Channel-level acceleration of deep face representations. IEEE Access 3, 2163–2175 (2015)
  41. 41.Rastegari, M., Ordonez, V., Redmon, J., Farhadi, A.: Xnor-net: Imagenet classification using binary convolutional neural networks. In: European Conference on Computer Vision. pp. 525–542. Springer (2016)
  42. 42.Real, E., Moore, S., Selle, A., Saxena, S., Suematsu, Y.L., Le, Q., Kurakin, A.: Large-scale evolution of image classifiers. arXiv preprint arXiv:1703.01041 (2017)
  43. 43.Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. In: Advances in neural information processing systems. pp. 91–99 (2015)
  44. 44.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., Chen, L.C.: Inverted residuals and linear bottlenecks: Mobile networks for classification, detection and segmentation. arXiv preprint arXiv:1801.04381 (2018)
  45. 45.Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)
  46. 46.Stanley, K.O., Miikkulainen, R.: Evolving neural networks through augmenting topologies. Evolutionary computation 10(2), 99–127 (2002)
  47. 47.Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: Going deeper with convolutions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 1–9 (2015)
  48. 48.Vasilache, N., Johnson, J., Mathieu, M., Chintala, S., Piantino, S., LeCun, Y.: Fast convolutional nets with fbfft: A gpu performance evaluation. arXiv preprint arXiv:1412.7580 (2014)
  49. 49.Wang, H., Zhang, Q., Wang, Y., Hu, R.: Structured probabilistic pruning for deep convolutional neural network acceleration. arXiv preprint arXiv:1709.06994 (2017)
  50. 50.Watkins, C.J.C.H.: Learning from delayed rewards. Ph.D. thesis, King’s College, Cambridge (1989)
  51. 51.Xue, J., Li, J., Gong, Y.: Restructuring of deep neural network acoustic models with singular value decomposition. In: INTERSPEECH. pp. 2365–2369 (2013)
  52. 52.Yang, T.J., Howard, A., Chen, B., Zhang, X., Go, A., Sze, V., Adam, H.: Netadapt: Platform-aware neural network adaptation for mobile applications. arXiv preprint arXiv:1804.03230 (2018)
  53. 53.Zhang, X., Zou, J., He, K., Sun, J.: Accelerating very deep convolutional networks for classification and detection. IEEE transactions on pattern analysis and machine intelligence 38(10), 1943–1955 (2016)
  54. 54.Zhong, Z., Yan, J., Liu, C.L.: Practical network blocks design with q-learning. arXiv preprint arXiv:1708.05552 (2017)
  55. 55.Zhu, C., Han, S., Mao, H., Dally, W.J.: Trained ternary quantization. arXiv preprint arXiv:1612.01064 (2016)
  56. 56.Zoph, B., Le, Q.V.: Neural architecture search with reinforcement learning. arXiv preprint arXiv:1611.01578 (2016)
  57. 57.Zoph, B., Vasudevan, V., Shlens, J., Le, Q.V.: Learning transferable architectures for scalable image recognition. arXiv preprint arXiv:1707.07012 (2017)

Citation

MLA
He, Y., et al. “AMC: AutoML for Model Compression and Acceleration on Mobile Devices”. Lecture Notes in Computer Science, Springer International Publishing, 2018, pp. 815–32, https://doi.org/10.1007/978-3-030-01234-2_48.
APA
He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., & Han, S. (2018). AMC: AutoML for Model Compression and Acceleration on Mobile Devices. In Lecture Notes in Computer Science (pp. 815–832). Springer International Publishing. https://doi.org/10.1007/978-3-030-01234-2_48
Chicago
He, Y., J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han. 2018. “AMC: AutoML for Model Compression and Acceleration on Mobile Devices”. In Lecture Notes in Computer Science. Springer International Publishing. https://doi.org/10.1007/978-3-030-01234-2_48.
Harvard
He, Y. et al. (2018) “AMC: AutoML for Model Compression and Acceleration on Mobile Devices”, Lecture Notes in Computer Science. Springer International Publishing, pp. 815–832. Available at: https://doi.org/10.1007/978-3-030-01234-2_48.
Vancouver
1. He Y, Lin J, Liu Z, Wang H, Li L-J, Han S (2018) AMC: AutoML for Model Compression and Acceleration on Mobile Devices. In: Lecture Notes in Computer Science. Springer International Publishing, pp 815–832

BibTeX

@inbook{He_2018, title={AMC: AutoML for Model Compression and Acceleration on Mobile Devices}, ISBN={9783030012342}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/978-3-030-01234-2_48}, DOI={10.1007/978-3-030-01234-2_48}, booktitle={Computer Vision – ECCV 2018}, publisher={Springer International Publishing}, author={He, Yihui and Lin, Ji and Liu, Zhijian and Wang, Hanrui and Li, Li-Jia and Han, Song}, year={2018}, pages={815–832} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF