Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge

Yiping KangJohann HauswaldCao GaoA. RovinskiT. MudgeJason MarsLingjia Tang

article2017ASPLOS1,598 citationsBest Paper Award

Proposes an adaptive runtime system that automatically partitions deep neural network inference layer by layer between mobile devices and datacenters to significantly reduce query latency, lower mobile energy consumption, and increase cloud throughput.

Listen

Modern intelligent applications, such as virtual assistants handling speech, natural language, and computer vision queries, rely heavily on deep neural networks. Traditionally, service providers execute these computationally heavy models entirely within cloud datacenters. However, uploading rich input data such as images and audio over wireless networks creates severe latency bottlenecks, drains mobile battery life, and places heavy computational demands on cloud infrastructure. As mobile hardware becomes significantly more capable, continuing with a cloud-only processing model becomes inefficient.

The article aims to evaluate the viability of splitting neural network computation across mobile devices and cloud servers, and it demonstrates an automated system to identify optimal, layer-by-layer partitioning strategies for intelligent applications.

To evaluate this approach, the researchers analyzed the data and compute profiles of eight production-grade neural networks covering vision, speech, and natural language tasks. Using an experimental setup consisting of a mobile development platform and a GPU-accelerated server, the team developed lightweight, platform-specific regression models to predict execution latency and power consumption for individual network layers. These models informed the creation of Neurosurgeon, an automated runtime scheduler that dynamically decides whether to execute each layer locally or in the cloud based on real-time network bandwidth and datacenter load.

The findings show that data transmission over wireless connections often accounts for over 90% of end-to-end response times in cloud-only setups. Furthermore, neural network layers exhibit distinct structural traits: computer vision models generally see data size decrease after front-end layers while computational intensity rises in back-end layers, creating ideal partition points in the middle of the network. Overall, deploying Neurosurgeon improved end-to-end response latency by an average of 3.1 times (and up to 40.7 times), reduced mobile device energy consumption by an average of 59.5% (and up to 94.7%), and increased cloud datacenter query throughput by an average of 1.5 times (and up to 6.7 times) compared to standard cloud-only execution. Neurosurgeon also outperformed existing code-offloading frameworks by an average of 1.9 times by basing decisions on network layer structures rather than code regions.

These results demonstrate that a dynamic, collaborative computing architecture provides substantial commercial and operational benefits. By offloading selected computation to user devices, organizations can drastically cut cloud hosting requirements, enhance user experience through lower latency, and preserve device battery life. Moreover, dynamic partitioning shields users from network volatility and datacenter traffic spikes without requiring model-specific profiling or developer annotations.

Organizations operating large-scale intelligent services should evaluate shifting from pure cloud processing to layer-aware collaborative intelligence systems. Adopting dynamic partitioning allows engineering teams to maximize the utilization of emerging edge hardware while optimizing datacenter capacity. For systems where full deployment is pending, teams should implement real-time network and server load monitoring to identify candidate applications—particularly vision models—that will benefit most from intermediate offloading.

The analysis carries high confidence for deep neural networks with linear topologies across common mobile and server GPU platforms. However, the evaluation focused on a specific mobile system-on-chip and eight benchmark architectures under consistent experimental conditions. Decision-makers should validate these prediction models across wider fleets of heterogeneous edge devices, varying wireless environments, and non-standard network topologies before enterprise-wide rollout.

Cover for Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge

Abstract

The computation for today's intelligent personal assistants such as Apple Siri, Google Now, and Microsoft Cortana, is performed in the cloud. This cloud-only approach requires significant amounts of data to be sent to the cloud over the wireless network and puts significant computational pressure on the datacenter. However, as the computational resources in mobile devices become more powerful and energy efficient, questions arise as to whether this cloud-only processing is desirable moving forward, and what are the implications of pushing some or all of this compute to the mobile devices on the edge.

In this paper, we examine the status quo approach of cloud-only processing and investigate computation partitioning strategies that effectively leverage both the cycles in the cloud and on the mobile device to achieve low latency, low energy consumption, and high datacenter throughput for this class of intelligent applications. Our study uses 8 intelligent applications spanning computer vision, speech, and natural language domains, all employing state-of-the-art Deep Neural Networks (DNNs) as the core machine learning technique. We find that given the characteristics of DNN algorithms, a fine-grained, layer-level computation partitioning strategy based on the data and computation variations of each layer within a DNN has significant latency and energy advantages over the status quo approach.

Using this insight, we design Neurosurgeon, a lightweight scheduler to automatically partition DNN computation between mobile devices and datacenters at the granularity of neural network layers. Neurosurgeon does not require per-application profiling. It adapts to various DNN architectures, hardware platforms, wireless networks, and server load levels, intelligently partitioning computation for best latency or best mobile energy. We evaluate Neurosurgeon on a state-of-the-art mobile development platform and show that it improves end-to-end latency by 3.1× on average and up to 40.7×, reduces mobile energy consumption by 59.5% on average and up to 94.7%, and improves datacenter throughput by 1.5× on average and up to 6.7×.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 3. Cloud-only Processing: The Status Quo
  • 3.1 Experimental setup
  • 3.2 Examining the Mobile Edge
  • 4. Fine-grained Computation Partitioning
  • 4.1 Layer Taxonomy
  • 4.2 Characterizing Layers in AlexNet
  • 4.3 Layer-granularity Computation Partitioning
  • 4.4 Generalizing to More DNNs
  • 5. Neurosurgeon
  • 5.1 Performance Prediction Model
  • 5.2 Dynamic DNN Partitioning
  • 5.3 Partitioned Execution
  • 6. Evaluation
  • 6.1 Latency Improvement
  • 6.2 Energy Improvement
  • 6.3 Comparing Neurosurgeon to MAUI
  • 6.4 Network Variation
  • 6.5 Server Load Variation
  • 6.6 Datacenter Throughput Improvement
  • 7. Related Work
  • 8. Conclusion
  • 9. Acknowledgment
  • References

Knowls

  1. Knowl 1 — Neurosurgeon Dynamic Layer Partitioning Algorithm

    algorithm

    Neurosurgeon selects an optimal partition boundary within a deep neural network (DNN) by evaluating estimated execution time and energy across candidate cut points. Every boundary after layer j∈{1,…,N}j \in \{1, \dots, N\} represents executing layers 11 through jj on the mobile device, transferring the intermediate tensor output of layer jj across the wireless network to the server, and executing layers j+1j+1 through NN in the cloud. A cut point before the first layer (j=0j=0, denoted as input) corresponds to cloud-only execution, while a cut point after the final layer (j=Nj=N) represents mobile-only execution.

    Input: Number of layers NN
    Input: Layer sequence {Li∣i=1,…,N}\{L_i \mid i = 1, \dots, N\}
    Input: Intermediate data sizes {Di∣i=1,…,N}\{D_i \mid i = 1, \dots, N\}
    Input: Regression latency predictors fmobile,fcloudf_{\text{mobile}}, f_{\text{cloud}}
    Input: Regression power predictor gmobileg_{\text{mobile}}
    Input: Server load level KK
    Input: Wireless uplink bandwidth BB
    Input: Wireless uplink power consumption PUPU
    Input: Optimization target OptTarget∈{latency,energy}\text{OptTarget} \in \{\text{latency}, \text{energy}\}
    Output: Optimal partition index j∗j^*
    procedure PARTITIONDECISION
        for i=1i = 1 to NN do
            TMi←fmobile(Li)TM_i \leftarrow f_{\text{mobile}}(L_i)
            TCi←fcloud(Li,K)TC_i \leftarrow f_{\text{cloud}}(L_i, K)
            PMi←gmobile(Li)PM_i \leftarrow g_{\text{mobile}}(L_i)
            TUi←Di/BTU_i \leftarrow D_i / B
        if OptTarget==latency\text{OptTarget} == \text{latency} then
            return arg⁡min⁡j∈{1,…,N}(∑i=1jTMi+∑k=j+1NTCk+TUj)\arg\min_{j \in \{1, \dots, N\}} \left( \sum_{i=1}^{j} TM_i + \sum_{k=j+1}^{N} TC_k + TU_j \right)
        else if OptTarget==energy\text{OptTarget} == \text{energy} then
            return arg⁡min⁡j∈{1,…,N}(∑i=1jTMi×PMi+TUj×PU)\arg\min_{j \in \{1, \dots, N\}} \left( \sum_{i=1}^{j} TM_i \times PM_i + TU_j \times PU \right)

    TMiTM_i and TCiTC_i denote the predicted execution times of layer LiL_i on the mobile device and server, respectively. PMiPM_i represents the predicted mobile power consumption during the execution of layer LiL_i. TUjTU_j is the data transmission latency for sending the output of layer jj (of size DjD_j) to the cloud over uplink bandwidth BB. Because the regression models are evaluated via closed-form arithmetic, partition selection executes with negligible runtime overhead.

  2. Knowl 2 — Layer-Level Performance and Power Prediction Models in Neurosurgeon

    model/method

    Neurosurgeon models the execution time and power consumption of arbitrary DNN layers without requiring per-application profiling. During a one-time deployment phase on a target mobile and server platform, regression models are constructed for each fundamental layer type by sweeping configurable hyperparameter spaces and measuring throughput (in Giga Floating Point Operations per Second, GFLOPS) and power dissipation.

    The regression models map specific layer configuration parameters to execution metrics using linear or logarithmic functions. Logarithmic functions capture the hardware saturation plateau when computational demand exceeds available compute units:

    • Convolution layers (conv): Characterized using two input variables: the number of features in the input feature maps, and the computational density per pixel given by (filter sizestride)2×(# of filters)\left(\frac{\text{filter size}}{\text{stride}}\right)^2 \times (\#\text{ of filters}).
    • Local and Pooling layers (local, pool): Modeled as a function of the input feature map dimensions and output feature map dimensions.
    • Fully-connected layers (fc), Softmax (softmax), and Argmax (argmax): Modeled as a function of the number of input neurons and output neurons.
    • Activation layers (relu, sig, htanh) and Normalization layers (norm): Modeled based on the total number of constituent neurons due to their element-wise one-to-one mapping.

    These models allow Neurosurgeon to predict per-layer mobile latency TMiTM_i, server latency TCiTC_i, and mobile power PMiPM_i for unseen network topologies at runtime.

  3. Knowl 3 — Layer-Wise Compute and Data Volume Dynamics in DNN Topologies

    empirical result

    Analysis of deep neural network architectures across vision, speech, and natural language domains reveals distinct layer-level computation and data size characteristics:

    1. Computer Vision (CV) Networks (e.g., AlexNet, VGG, DeepFace, MNIST): Early convolutional layers significantly expand data volume relative to the raw input image due to large numbers of learned feature filters. Subsequent pooling layers sharply compress tensor dimensions (reducing intermediate data volume by up to 4.7×4.7\times). Computation latency is heavily skewed toward deeper layers: fully-connected layers (such as fc6 in AlexNet, which accounts for 45%45\% of total network execution time) are up to an order of magnitude slower on mobile platforms than front-end convolutional layers. This creates optimal partition opportunities in the intermediate layers where data transfer volume is minimal and heavy back-end compute can be offloaded to the cloud.
    2. Speech and NLP Networks (e.g., Kaldi ASR, SENNA POS/NER/CHK): These architectures consist almost entirely of fully-connected and element-wise activation layers without data-expanding convolutions or data-reducing pooling layers. Intermediate tensor dimensions remain relatively constant throughout execution. Consequently, per-layer execution times are uniform, and optimal partition points occur only at the network extremities (cloud-only or mobile-only execution).
  4. Knowl 4 — Neurosurgeon System Architecture and Partitioned Runtime Framework

    model/method

    Neurosurgeon operates in two distinct phases:

    1. Deployment Phase: Hardware platforms (mobile device and server) are profiled once across standard neural network layer types to train regression predictors for latency and power. These lightweight models are stored locally on the mobile device.
    2. Runtime Phase: The mobile runtime system (NSmobile) inspects the computational graph of an incoming DNN task, extracts layer configurations, and computes optimal cut points using dynamic inputs: real-time wireless uplink bandwidth (BB), measured uplink transmission power (PUPU), and datacenter load level (KK) obtained via lightweight periodic server pings.

    Execution is orchestrated across modified deep learning framework instances (NSmobile on the client and NSserver on the cloud) communicating via Apache Thrift RPC. Both instances store the full DNN model parameters. NSmobile executes layers from the input up to layer j∗j^*, serializes and transmits the intermediate tensor output across the network to NSserver, which resumes execution from layer j∗+1j^*+1 to the final layer NN. The final inference result is returned to NSmobile. Exactly one network transfer occurs per query.

  5. Knowl 5 — End-to-End Latency and Energy Improvements of Neurosurgeon

    empirical result

    Across an evaluation suite of 8 DNN workloads spanning computer vision, speech recognition, and natural language processing evaluated on 3G, LTE, and Wi-Fi networks using mobile CPU and GPU platforms:

    • End-to-End Latency: Neurosurgeon achieves an average latency speedup of 3.1×3.1\times (geometric mean) and up to 40.7×40.7\times compared to the cloud-only baseline. For computer vision workloads on mobile GPUs, intermediate layer partitioning avoids both high wireless transmission overhead and heavy mobile fully-connected execution. For speech recognition, Neurosurgeon retains cloud-only execution, matching baseline performance.
    • Mobile Energy Consumption: Neurosurgeon reduces mobile device energy consumption by an average of 59.5%59.5\% (geometric mean) and up to 94.7%94.7\% relative to cloud-only processing.
    • Prediction Accuracy: Neurosurgeon selects the exact optimal partition point in 4444 out of 4848 evaluated configurations (8 benchmarks ×\times 3 network types ×\times 2 mobile compute units). Suboptimal choices occur only when adjacent partition candidates exhibit near-identical performance, with Neurosurgeon delivering within 98.5%98.5\% of optimal latency speedup and 98.8%98.8\% of optimal energy reduction.
  6. Knowl 6 — Neurosurgeon Optimal Partition Decisions Across Benchmarks and Networks

    data/table

    The optimal layer boundaries chosen by Neurosurgeon for best end-to-end latency vary systematically across network uplink bandwidth, mobile compute capabilities (CPU vs. GPU), and model layer composition:

    Mobile Network IMC VGG FACE DIG ASR POS NER CHK
    CPU Wi-Fi input input input input input fc3 fc3 fc3
    CPU LTE input input input argmax input fc3 fc3 fc3
    CPU 3G argmax input input argmax input fc3 fc3 fc3
    GPU Wi-Fi pool5 input input argmax input fc3 fc3 fc3
    GPU LTE argmax argmax input argmax input fc3 fc3 fc3
    GPU 3G argmax argmax argmax argmax input fc3 fc3 fc3

    Here input denotes cloud-only execution, argmax or fc3 denotes mobile-only execution, and intermediate names (e.g., pool5 in IMC/AlexNet) indicate an intermediate cut point where front-end feature extraction executes on mobile and back-end classification runs in the cloud. As wireless bandwidth degrades (Wi-Fi →\to LTE →\to 3G) or mobile compute power increases (CPU →\to GPU), optimal partitioning shifts from the cloud toward the mobile edge.

  7. Knowl 7 — Datacenter Throughput Gains Under Collaborative Edge-Cloud Execution

    empirical result

    Offloading front-end or entire DNN computations to mobile edge devices reduces execution cycles required on datacenter servers, directly shortening server query service times and increasing query processing capacity. Simulated using BigHouse with query inter-arrival distributions derived from Google web search traces and an equal mix of the 8 evaluated DNN workloads:

    • Under Wi-Fi connectivity, Neurosurgeon achieves an average datacenter throughput improvement of 1.04×1.04\times.
    • Under LTE connectivity, datacenter throughput increases by 1.43×1.43\times on average.
    • Under 3G connectivity, datacenter throughput increases by 2.36×2.36\times on average.
    • Across scenarios with high mobile GPU adoption (100%100\% mobile GPU clients), datacenter throughput improvements reach up to 6.7×6.7\times over the cloud-only baseline, as greater portions of computation are pushed to mobile edge devices.
  8. Knowl 8 — Data-Centric DNN Partitioning Versus Control-Centric Code Offloading

    empirical result

    Neurosurgeon outperforms general-purpose control-centric computation offloading systems such as MAUI by up to 32×32\times (1.9×1.9\times on average).

    Control-centric frameworks profile execution at the granularity of program functions or methods and use past runtime history of a function to predict future behavior. In DNNs, this assumption fails because layers calling the same underlying function (e.g., convolution) have drastically different computational intensities and data sizes depending on their depth and tensor dimensions. For example, in VGG, the input data size for layer conv1.1 is 0.57 MB0.57\text{ MB}, whereas layer conv1.2 receives 12.25 MB12.25\text{ MB}. MAUI predicts future layer costs based on previous layer invocations, mispredicting the cost of conv1.2 and choosing to offload prior to it, resulting in massive data transfer over LTE and causing a 20.5×20.5\times slowdown relative to the cloud-only baseline. Neurosurgeon's data-centric approach analyzes the layer parameter topology directly, correctly selecting cloud-only execution and achieving a 20.5×20.5\times speedup over MAUI.

  9. Knowl 9 — Dynamic Adaptation to Wireless Network and Datacenter Load Variations

    empirical result

    Neurosurgeon dynamically adapts partition boundaries in response to real-time network fluctuations and diurnal datacenter load shifts:

    • Network Bandwidth Variance: Under continuous T-Mobile LTE bandwidth monitoring, available uplink speed fluctuates between <1 Mbps<1\text{ Mbps} and 5 Mbps5\text{ Mbps}. While cloud-only execution suffers severe latency spikes during low-bandwidth periods, Neurosurgeon dynamically shifts execution from remote cloud processing to intermediate partitioned execution or fully local mobile execution, maintaining consistent low end-to-end latency.
    • Datacenter Load Variance: As server load increases from 10%10\% to 90%90\%, cloud processing latency for AlexNet increases from 105 ms105\text{ ms} to 753 ms753\text{ ms} due to queuing and resource contention. Neurosurgeon dynamically adjusts its partition point across two thresholds: executing entirely in the cloud at low load, partitioning across mobile and cloud at medium load, and executing entirely on the mobile device at peak load, maintaining end-to-end latency strictly below 380 ms380\text{ ms}.
  10. Knowl 10 — Experimental Hardware and Benchmark Setup

    experimental setup

    The experimental evaluation uses the following hardware, software, and benchmark configuration:

    • Mobile Edge Platform: NVIDIA Jetson TK1 development board featuring the Tegra K1 SoC with a quad-core ARM Cortex A15 CPU (up to 2.1 GHz2.1\text{ GHz}), an NVIDIA Kepler mobile GPU with 192 CUDA cores, and 2 GB2\text{ GB} DDR3L memory (933 MHz933\text{ MHz}). Mobile power and energy are measured using a Watts Up? digital power meter.
    • Server Platform: 4U server chassis equipped with two Intel Xeon E5-2620 V2 CPUs (6 cores each, 2.10 GHz6\text{ cores each, } 2.10\text{ GHz}), 256 GB256\text{ GB} DDR3 ECC memory, and an NVIDIA Tesla K40 GPU (12 GB12\text{ GB} GDDR5 PCIe).
    • Software Stack: Modified Caffe deep learning framework utilizing OpenBLAS (NEON-vectorized matrix multiplication) on ARM CPU, cuDNN and CUDA on GPUs, and Apache Thrift RPC for inter-process edge-cloud communication.
    • Benchmark Suite: 8 DNN architectures encompassing Computer Vision (AlexNet/IMC: 24 layers; VGG: 46 layers; DeepFace/FACE: 10 layers; MNIST/DIG: 9 layers), Automatic Speech Recognition (Kaldi/ASR: 13 layers), and Natural Language Processing (SENNA POS: 3 layers; SENNA NER: 3 layers; SENNA CHK: 3 layers).

Coverage note — None was omitted; the extracted knowls comprehensively capture Neurosurgeon's layer-level analysis, prediction models, dynamic partitioning algorithm, system implementation, empirical latency/energy/throughput results, comparative evaluation against prior offloading frameworks, dynamic adaptability, and experimental setup.

References

  1. 1.Wearables market to be worth $25 billion by 2019. http://www.ccsinsight.com/press/company-news/2332-wearables-market-to-be-worth-25-billion-by-2019-reveals-ccs-insight. Accessed: 2017-01.
  2. 2.Rapid Expansion Projected for Smart Home Devices, IHS Markit Says. http://news.ihsmarkit.com/press-release/technology/rapid-expansion-projected-smart-home-devices-ihs-markit-says. Accessed: 2017-01.
  3. 3.Intelligent Virtual Assistant Market Worth $3.07Bn By 2020. https://globenewswire.com/news-release/2015/12/17/796353/0/en/Intelligent-Virtual-Assistant-Market-Worth-3-07Bn-By-2020.html. Accessed: 2016-08.
  4. 4.Intelligent Virtual Assistant Market Analysis And Segment Forecasts 2015 To 2022. https://www.hexaresearch.com/research-report/intelligent-virtual-assistant-industry/. Accessed: 2016-08.
  5. 5.Growing Focus on Strengthening Customer Relations Spurs Adoption of Intelligent Virtual Assistant Technology. http://www.transparencymarketresearch.com/pressrelease/intelligent-virtual-assistant-industry.html/. Accessed: 2016-08.
  6. 6.Google Brain. https://backchannel.com/google-search-will-be-your-next-brain-5207c26e4523#.x9n2ajota. Accessed: 2017-01.
  7. 7.Microsoft Deep Learning Outperforms Humans in Image Recognition. http://www.forbes.com/sites/michaelthomsen/2015/02/19/microsofts-deep-learning-project-outperforms-humans-in-image-recognition/. Accessed: 2016-08.
  8. 8.Baidu Supercomputer. https://gigaom.com/2015/01/14/baidu-has-built-a-supercomputer-for-deep-learning/. Accessed: 2016-08.
  9. 9.Johann Hauswald, Yiping Kang, Michael A. Laurenzano, Quan Chen, Cheng Li, Trevor Mudge, Ronald G. Dreslinski, Jason Mars, and Lingjia Tang. Djinn and tonic: Dnn as a service and its implications for future warehouse scale computers. In Proceedings of the 42nd Annual International Symposium on Computer Architecture (ISCA), ISCA ’15, New York, NY, USA, 2015. ACM.
  10. 10.The ’Google Brain’ is a real thing but very few people have seen it. http://www.businessinsider.com/what-is-google-brain-2016-9. Accessed: 2017-01.
  11. 11.Google supercharges machine learning tasks with TPU custom chip. https://cloudplatform.googleblog.com/2016/05/Google-supercharges-machine-learning-tasks-with-custom-chip.html. Accessed: 2017-01.
  12. 12.Apple’s Massive New Data Center Set To Host Nuance Tech. http://techcrunch.com/2011/05/09/apple-nuance-data-center-deal/. Accessed: 2016-08.
  13. 13.Apple moves to third-generation Siri back-end, built on open-source Mesos platform. http://9to5mac.com/2015/04/27/siri-backend-mesos/. Accessed: 2016-08.
  14. 14.Matthew Halpern, Yuhao Zhu, and Vijay Janapa Reddi. Mobile cpu’s rise to power: Quantifying the impact of generational mobile cpu design trends on performance, energy, and user satisfaction. In High Performance Computer Architecture (HPCA), 2016 IEEE International Symposium on, pages 64–76. IEEE, 2016.
  15. 15.Whitepaper: NVIDIA Tegra X1. Technical report. Accessed: 2017-01.
  16. 16.NVIDIA Jetson TK1 Development Kit: Bringing GPU-accelerated computing to Embedded Systems. Technical report. Accessed: 2017-01.
  17. 17.Nvidia’s Tegra K1 at the Heart of Google’s Nexus 9. http://www.pcmag.com/article2/0,2817,2470740,00.asp. Accessed: 2016-08.
  18. 18.Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014.
  19. 19.Qian Wang, Xianyi Zhang, Yunquan Zhang, and Qing Yi. Augem: automatically generate high performance dense linear algebra kernels on x86 cpus. In Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, page 25. ACM, 2013.
  20. 20.Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. cuDNN: Efficient Primitives for Deep Learning. CoRR, abs/1410.0759, 2014.
  21. 21.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, 2012.
  22. 22.Adam Coates, Brody Huval, Tao Wang, David Wu, Bryan Catanzaro, and Andrew Ng. Deep learning with cots hpc systems. In Proceedings of the 30th international conference on machine learning, pages 1337–1345, 2013.
  23. 23.TestMyNet: Internet Speed Test. http://testmy.net/. Accessed: 2015-02.
  24. 24.Watts Up? Power Meter. https://www.wattsupmeters.com/. Accessed: 2015-05.
  25. 25.Junxian Huang, Feng Qian, Alexandre Gerber, Z Morley Mao, Subhabrata Sen, and Oliver Spatscheck. A close examination of performance and power characteristics of 4g lte networks. In Proceedings of the 10th international conference on Mobile systems, applications, and services, pages 225–238. ACM, 2012.
  26. 26.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  27. 27.Yaniv Taigman, Ming Yang, Marc’Aurelio Ranzato, and Lior Wolf. Deepface: Closing the gap to human-level performance in face verification. In Computer Vision and Pattern Recognition (CVPR), 2014.
  28. 28.Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998.
  29. 29.Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. The kaldi speech recognition toolkit. In Proc. ASRU, 2011.
  30. 30.Ronan Collobert, Jason Weston, Leon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. Natural language processing (almost) from scratch. The Journal of Machine Learning Research, 2011.
  31. 31.Ashkan Nikravesh, David R Choffnes, Ethan Katz-Bassett, Z Morley Mao, and Matt Welsh. Mobile network performance from user devices: A longitudinal, multidimensional analysis. In International Conference on Passive and Active Network Measurement, pages 12–22. Springer, 2014.
  32. 32.David Lo, Liqun Cheng, Rama Govindaraju, Luiz Andre Barroso, and Christos Kozyrakis. Towards energy proportionality for large-scale latency-critical workloads. In ACM SIGARCH Computer Architecture News, volume 42, pages 301–312. IEEE Press, 2014.
  33. 33.Mark Slee, Aditya Agarwal, and Marc Kwiatkowski. Thrift: Scalable cross-language services implementation. Facebook White Paper, 5(8), 2007.
  34. 34.Eduardo Cuervo, Aruna Balasubramanian, Dae-ki Cho, Alec Wolman, Stefan Saroiu, Ranveer Chandra, and Paramvir Bahl. Maui: making smartphones last longer with code offload. In Proceedings of the 8th international conference on Mobile systems, applications, and services, pages 49–62. ACM, 2010.
  35. 35.Mark S Gordon, Davoud Anoushe Jamshidi, Scott A Mahlke, Zhuoqing Morley Mao, and Xu Chen. Comet: Code offload by migrating execution transparently.
  36. 36.Moo-Ryong Ra, Anmol Sheth, Lily Mummert, Padmanabhan Pillai, David Wetherall, and Ramesh Govindan. Odessa: enabling interactive perception applications on mobile devices. In Proceedings of the 9th international conference on Mobile systems, applications, and services, pages 43–56. ACM, 2011.
  37. 37.Byung-Gon Chun, Sunghwan Ihm, Petros Maniatis, Mayur Naik, and Ashwin Patti. Clonecloud: elastic execution between mobile device and cloud. In Proceedings of the sixth conference on Computer systems, pages 301–314. ACM, 2011.
  38. 38.David Meisner, Junjie Wu, and Thomas F. Wenisch. BigHouse: A Simulation Infrastructure for Data Center Systems. ISPASS ’12: International Symposium on Performance Analysis of Systems and Software, April 2012.
  39. 39.Chang-Hong Hsu, Yunqi Zhang, Michael A. Laurenzano, David Meisner, Thomas Wenisch, Lingjia Tang, Jason Mars, and Ronald G. Dreslinski. Adrenaline: Pinpointing and reigning in tail queries with quick voltage boosting. In International Symposium on High Performance Computer Architecture (HPCA), 2015.
  40. 40.Michael A. Laurenzano, Yunqi Zhang, Lingjia Tang, and Jason Mars. Protean code: Achieving near-free online code transformations for warehouse scale computers. In International Symposium on Microarchitecture (MICRO), 2014.
  41. 41.Jason Mars, Lingjia Tang, Robert Hundt, Kevin Skadron, and Mary Lou Soffa. Bubble-up: Increasing utilization in modern warehouse scale computers via sensible co-locations. In International Symposium on Microarchitecture (MICRO), 2011.
  42. 42.Vinicius Petrucci, Michael A. Laurenzano, Yunqi Zhang, John Doherty, Daniel Mosse, Jason Mars, and Lingjia Tang. Octopus-man: Qos-driven task management for heterogeneous multicore in warehouse scale computers. In International Symposium on High Performance Computer Architecture (HPCA), 2015.
  43. 43.Jason Mars and Lingjia Tang. Whare-map: Heterogeneity in ”homogeneous” warehouse-scale computers. In International Symposium on Computer Architecture (ISCA), 2013.
  44. 44.Johann Hauswald, Tom Manville, Qi Zheng, Ronald G. Dreslinski, Chaitali Chakrabarti, and Trevor Mudge. A hybrid approach to offloading mobile image classification. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014.
  45. 45.Johann Hauswald, Michael A. Laurenzano, Yunqi Zhang, Cheng Li, Austin Rovinski, Arjun Khurana, Ronald G. Dreslinski, Trevor Mudge, Vinicius Petrucci, Lingjia Tang, and Jason Mars. Sirius: An open end-to-end voice and vision personal assistant and its implications for future warehouse scale computers. In International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2015.
  46. 46.Hailong Yang, Alex Breslow, Jason Mars, and Lingjia Tang. Bubble-flux: Precise online qos management for increased utilization in warehouse scale computers. In International Symposium on Computer Architecture (ISCA), 2013.
  47. 47.Matt Skach, Manish Arora, Chang-Hong Hsu, Qi Li, Dean Tullsen, Lingjia Tang, and Jason Mars. Thermal time shifting: Leveraging phase change materials to reduce cooling costs in warehouse-scale computers. In Proceedings of the 42nd Annual International Symposium on Computer Architecture (ISCA), ISCA ’15, 2015.
  48. 48.Yunqi Zhang, Michael A. Laurenzano, Jason Mars, and Lingjia Tang. Smite: Precise qos prediction on real system smt processors to improve utilization in warehouse scale computers. In International Symposium on Microarchitecture (MICRO), 2014.
  49. 49.Quan Chen, Hailong Yang, Jason Mars, and Lingjia Tang. Baymax: Qos awareness and increased utilization for nonpreemptive accelerators in warehouse scale computers. In ACM SIGPLAN Notices, volume 51, pages 681–696. ACM, 2016.
  50. 50.Yunqi Zhang, David Meisner, Jason Mars, and Lingjia Tang. Treadmill: Attributing the source of tail latency through precise load testing and statistical inference. In Computer Architecture (ISCA), 2016 ACM/IEEE 43rd Annual International Symposium on, pages 456–468. IEEE, 2016.
  51. 51.Animesh Jain, Michael A Laurenzano, Lingjia Tang, and Jason Mars. Continuous shape shifting: Enabling loop co-optimization via near-free dynamic code rewriting. In Microarchitecture (MICRO), 2016 49th Annual IEEE/ACM International Symposium on, pages 1–12. IEEE, 2016.
  52. 52.Michael A. Laurenzano, Yunqi Zhang, Jiang Chen, Lingjia Tang, and Jason Mars. Powerchop: Identifying and managing non-critical units in hybrid processor architectures. In Proceedings of the 43rd International Symposium on Computer Architecture, ISCA ’16, pages 140–152, Piscataway, NJ, USA, 2016. IEEE Press.
  53. 53.Tianshi Chen, Zidong Du, Ninghui Sun, Jia Wang, Chengyong Wu, Yunji Chen, and Olivier Temam. Diannao: A small-footprint high-throughput accelerator for ubiquitous machine-learning. In Proceedings of the 19th international conference on Architectural support for programming languages and operating systems, pages 269–284. ACM, 2014.
  54. 54.Daofu Liu, Tianshi Chen, Shaoli Liu, Jinhong Zhou, Shengyuan Zhou, Olivier Teman, Xiaobing Feng, Xuehai Zhou, and Yunji Chen. Pudiannao: A polyvalent machine learning accelerator. In Proceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems, pages 369–381. ACM, 2015.
  55. 55.Kalin Ovtcharov, Olatunji Ruwase, Joo-Young Kim, Jeremy Fowers, Karin Strauss, and Eric S Chung. Accelerating deep convolutional neural networks using specialized hardware. Microsoft Research Whitepaper, 2(11), 2015.
  56. 56.Xin Lei, Andrew Senior, Alexander Gruenstein, and Jeffrey Sorensen. Accurate and Compact Large vocabulary speech recognition on mobile devices. In INTERSPEECH, pages 662–665, 2013.
  57. 57.Xin Lei, Andrew Senior, Alexander Gruenstein, and Jeffrey Sorensen. Accurate and compact large vocabulary speech recognition on mobile devices. In INTERSPEECH, pages 662–665, 2013.
  58. 58.Seungyeop Han, Haichen Shen, Matthai Philipose, Sharad Agarwal, Alec Wolman, and Arvind Krishnamurthy. Mcdnn: An execution framework for deep neural networks on resource-constrained devices. In MobiSys, 2016.

Citation

MLA
Kang, Y., et al. “Neurosurgeon”. Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, 2017, pp. 615–29, https://doi.org/10.1145/3037697.3037698.
APA
Kang, Y., Hauswald, J., Gao, C., Rovinski, A., Mudge, T., Mars, J., & Tang, L. (2017). Neurosurgeon. Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, 615–629. https://doi.org/10.1145/3037697.3037698
Chicago
Kang, Y., J. Hauswald, C. Gao, et al. 2017. “Neurosurgeon”. Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, 615–29. https://doi.org/10.1145/3037697.3037698.
Harvard
Kang, Y. et al. (2017) “Neurosurgeon”, Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems. ACM, pp. 615–629. Available at: https://doi.org/10.1145/3037697.3037698.
Vancouver
1. Kang Y, Hauswald J, Gao C, Rovinski A, Mudge T, Mars J, Tang L (2017) Neurosurgeon. In: Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems. ACM, pp 615–629

BibTeX

@inproceedings{Kang_2017, series={ASPLOS ’17}, title={Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge}, url={http://dx.doi.org/10.1145/3037697.3037698}, DOI={10.1145/3037697.3037698}, booktitle={Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems}, publisher={ACM}, author={Kang, Yiping and Hauswald, Johann and Gao, Cao and Rovinski, Austin and Mudge, Trevor and Mars, Jason and Tang, Lingjia}, year={2017}, month=Apr, pages={615–629}, collection={ASPLOS ’17} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF