Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge
Yiping KangJohann HauswaldCao GaoA. RovinskiT. MudgeJason MarsLingjia Tang
Proposes an adaptive runtime system that automatically partitions deep neural network inference layer by layer between mobile devices and datacenters to significantly reduce query latency, lower mobile energy consumption, and increase cloud throughput.
Modern intelligent applications, such as virtual assistants handling speech, natural language, and computer vision queries, rely heavily on deep neural networks. Traditionally, service providers execute these computationally heavy models entirely within cloud datacenters. However, uploading rich input data such as images and audio over wireless networks creates severe latency bottlenecks, drains mobile battery life, and places heavy computational demands on cloud infrastructure. As mobile hardware becomes significantly more capable, continuing with a cloud-only processing model becomes inefficient.
The article aims to evaluate the viability of splitting neural network computation across mobile devices and cloud servers, and it demonstrates an automated system to identify optimal, layer-by-layer partitioning strategies for intelligent applications.
To evaluate this approach, the researchers analyzed the data and compute profiles of eight production-grade neural networks covering vision, speech, and natural language tasks. Using an experimental setup consisting of a mobile development platform and a GPU-accelerated server, the team developed lightweight, platform-specific regression models to predict execution latency and power consumption for individual network layers. These models informed the creation of Neurosurgeon, an automated runtime scheduler that dynamically decides whether to execute each layer locally or in the cloud based on real-time network bandwidth and datacenter load.
The findings show that data transmission over wireless connections often accounts for over 90% of end-to-end response times in cloud-only setups. Furthermore, neural network layers exhibit distinct structural traits: computer vision models generally see data size decrease after front-end layers while computational intensity rises in back-end layers, creating ideal partition points in the middle of the network. Overall, deploying Neurosurgeon improved end-to-end response latency by an average of 3.1 times (and up to 40.7 times), reduced mobile device energy consumption by an average of 59.5% (and up to 94.7%), and increased cloud datacenter query throughput by an average of 1.5 times (and up to 6.7 times) compared to standard cloud-only execution. Neurosurgeon also outperformed existing code-offloading frameworks by an average of 1.9 times by basing decisions on network layer structures rather than code regions.
These results demonstrate that a dynamic, collaborative computing architecture provides substantial commercial and operational benefits. By offloading selected computation to user devices, organizations can drastically cut cloud hosting requirements, enhance user experience through lower latency, and preserve device battery life. Moreover, dynamic partitioning shields users from network volatility and datacenter traffic spikes without requiring model-specific profiling or developer annotations.
Organizations operating large-scale intelligent services should evaluate shifting from pure cloud processing to layer-aware collaborative intelligence systems. Adopting dynamic partitioning allows engineering teams to maximize the utilization of emerging edge hardware while optimizing datacenter capacity. For systems where full deployment is pending, teams should implement real-time network and server load monitoring to identify candidate applications—particularly vision models—that will benefit most from intermediate offloading.
The analysis carries high confidence for deep neural networks with linear topologies across common mobile and server GPU platforms. However, the evaluation focused on a specific mobile system-on-chip and eight benchmark architectures under consistent experimental conditions. Decision-makers should validate these prediction models across wider fleets of heterogeneous edge devices, varying wireless environments, and non-standard network topologies before enterprise-wide rollout.
- Paper: BranchyNet: Fast inference via early exiting from deep neural networks, Surat Teerapittayanon et al. (2016). Introduces dynamic inference routing and multi-tier early exiting across network layers, providing a foundational architectural concept for dynamic layer-level edge-cloud partitioning.
- Paper: DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learning, Tianshi Chen et al. (2014). Establishes how memory transfers and layer-by-layer computational bottlenecks dominate neural network execution on embedded and accelerator hardware.
- Paper: Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation, Emily L. Denton et al. (2014). Analyzes the layer-wise structural traits and computational variations between early and deep convolutional layers that inform layer partitioning decisions.
- Paper: Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding, Song Han et al. (2015). Presents foundational methods for reducing model data footprint and evaluating layer-specific resource demands on mobile devices.
- Paper: Learning both Weights and Connections for Efficient Neural Networks, Song Han et al. (2015). Examines the severe memory and energy trade-offs of deploying deep neural networks to constrained mobile hardware.
- Paper: Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing, Zhi Zhou et al. (2019). Synthesizes multi-level edge-cloud collaborative intelligence paradigms and formalizes taxonomy frameworks that build upon collaborative runtime partitioners.
- Paper: Clipper: a low-latency online prediction serving system, Daniel Crankshaw et al. (2017). Addresses system-level online serving, adaptive batching, and low-latency scheduling for machine learning models across distributed environments.
- Paper: MnasNet: Platform-Aware Neural Architecture Search for Mobile, Mingxing Tan et al. (2018). Extends hardware-aware optimization by automating platform-specific neural architecture design directly for measured mobile execution latency.
- Paper: Once for All: Train One Network and Specialize it for Efficient Deployment, Han Cai et al. (2019). Generalizes multi-device deployment by training a single versatile super-network that dynamically specializes sub-networks to diverse hardware constraints without retraining.
- Paper: ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware, Han Cai et al. (2018). Develops direct hardware-latency prediction models to guide architecture search across heterogeneous mobile and cloud platforms.
