Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing

Zhi ZhouXu ChenEn LiLiekang ZengKe LuoJunshan Zhang

article2019Proceedings of the IEEE2,009 citations

Synthesizes the architectural designs, software frameworks, and emerging technologies required to train and deploy deep learning models efficiently across distributed edge networks.

arXiv: 1905.10083
Cover for Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing

Abstract

With the breakthroughs in deep learning, the recent years have witnessed a booming of artificial intelligence (AI) applications and services, spanning from personal assistant to recommendation systems to video/audio surveillance. More recently, with the proliferation of mobile computing and Internet-of-Things (IoT), billions of mobile and IoT devices are connected to the Internet, generating zillions Bytes of data at the network edge. Driving by this trend, there is an urgent need to push the AI frontiers to the network edge so as to fully unleash the potential of the edge big data. To meet this demand, edge computing, an emerging paradigm that pushes computing tasks and services from the network core to the network edge, has been widely recognized as a promising solution. The resulted new inter-discipline, edge AI or edge intelligence, is beginning to receive a tremendous amount of interest. However, research on edge intelligence is still in its infancy stage, and a dedicated venue for exchanging the recent advances of edge intelligence is highly desired by both the computer system and artificial intelligence communities. To this end, we conduct a comprehensive survey of the recent research efforts on edge intelligence. Specifically, we first review the background and motivation for artificial intelligence running at the network edge. We then provide an overview of the overarching architectures, frameworks and emerging key technologies for deep learning model towards training/inference at the network edge. Finally, we discuss future research opportunities on edge intelligence. We believe that this survey will elicit escalating attentions, stimulate fruitful discussions and inspire further research ideas on edge intelligence.

Table of Contents

  • I Introduction
  • II A Primer on Artificial Intelligence
  • II-A Artificial Intelligence
  • II-B Deep Learning and Deep Neural Networks
  • II-C From Deep Learning to Model Training and Inference
  • II-D Popular Deep Learning Models
  • III Edge Intelligence
  • III-A Motivation and Benefits of Edge Intelligence
  • III-B Scope and Rating of Edge Intelligence
  • IV Edge Intelligence Model Training
  • IV-A Architectures
  • IV-A1 Centralized
  • IV-A2 Decentralized
  • IV-A3 Hybrid
  • IV-B Key Performance Indicators
  • IV-B1 Training Loss
  • IV-B2 Convergence
  • IV-B3 Privacy
  • IV-B4 Communication Cost
  • IV-B5 Latency
  • IV-B6 Energy Efficiency
  • IV-C Enabling Technologies
  • IV-C1 Federated Learning
  • IV-C2 Aggregation Frequency Control
  • IV-C3 Gradient Compression
  • IV-C4 DNN Splitting
  • IV-C5 Knowledge Transfer Learning
  • IV-C6 Gossip Training
  • IV-D Summary of Existing Systems and Frameworks
  • V Edge Intelligence Model Inference
  • V-A Architectures
  • V-A1 Edge-based
  • V-A2 Device-based
  • V-A3 Edge-device
  • V-A4 Edge-cloud
  • V-B Key Performance Indicators
  • V-B1 Latency
  • V-B2 Accuracy
  • V-B3 Energy
  • V-B4 Privacy
  • V-B5 Communication overhead
  • V-B6 Memory Footprint
  • V-C Enabling Technologies
  • V-C1 Model Compression
  • V-C2 Model Partition
  • V-C3 Model Early-Exit
  • V-C4 Edge Caching
  • V-C5 Input Filtering
  • V-C6 Model Selection
  • V-C7 Support for Multi-Tenancy
  • V-C8 Application-Specific Optimization
  • V-D Summary of Existing Systems and Frameworks
  • VI Future Research Directions
  • VI-A Programming and Software Platforms
  • VI-B Resource-friendly Edge AI Model Design
  • VI-C Computation-aware Networking Techniques
  • VI-D Trade-off Design with Various DNN Performance Metrics
  • VI-E Smart Service and Resource Management
  • VI-F Security and Privacy Issues
  • VI-G Incentive and Business Models
  • VII Concluding Remarks
  • References

Knowls

  1. Knowl 1 — Six-Level Rating Framework for Edge Intelligence

    definition

    Edge Intelligence (EI) is defined as the paradigm that exploits data and resources across the full hierarchy of end devices, edge nodes, and cloud datacenters to optimize the training and inference of deep learning models. Rather than restricting EI solely to on-device execution, EI spans six levels categorized by the amount and path length of data offloading:

    • Cloud Intelligence (Level 0): Training and inference of the deep neural network (DNN) model are performed entirely in the cloud datacenter.
    • Level 1 — Cloud–Edge Co-Inference and Cloud Training: The DNN model is trained in the cloud, while inference is executed through cloud-edge cooperation by partially offloading data to the cloud.
    • Level 2 — In-Edge Co-Inference and Cloud Training: The DNN model is trained in the cloud, while inference is executed entirely within the network edge by offloading data to local edge nodes or nearby devices via device-to-device (D2D) communications without cloud involvement.
    • Level 3 — On-Device Inference and Cloud Training: The DNN model is trained in the cloud, but inference is executed completely locally on the end device with zero data offloading.
    • Level 4 — Cloud–Edge Co-Training and Inference: Both training and inference are performed collaboratively across the cloud-edge hierarchy via data offloading.
    • Level 5 — All In-Edge: Both training and inference are executed entirely within the network edge without cloud datacenter intervention.
    • Level 6 — All On-Device: Both training and inference are executed entirely locally on the end device without any external data offloading.

    As the EI level increases from Level 1 to Level 6, the volume and path length of offloaded data decrease, resulting in reduced transmission latency, lower wide-area network (WAN) bandwidth costs, and enhanced privacy, at the expense of higher on-device computation latency and energy consumption.

  2. Knowl 2 — Architectural Modes and Performance Indicators for Distributed Edge DNN Training

    model/method

    Distributed training of Deep Neural Networks (DNNs) across edge computing environments is structured into three fundamental architectural modes:

    1. Centralized Mode: Training data generated at distributed end devices (e.g., mobile phones, cameras, vehicles) is transmitted across wide-area networks to a central cloud datacenter where the complete DNN model is trained.
    2. Decentralized Mode: Each computing node trains its local DNN model using locally stored data and exchanges model updates or gradients directly with peer nodes, eliminating the need for a central cloud coordinator and preserving raw data privacy.
    3. Hybrid Mode (Cloud-Edge-Device): Edge servers serve as local coordination hubs that perform decentralized model updates among themselves or engage in periodic centralized synchronization with the cloud datacenter.

    Distributed edge training systems are evaluated along six Key Performance Indicators (KPIs):

    • Training Loss: Measures how accurately the trained DNN fits the training data by minimizing the objective loss function.
    • Convergence: Evaluates whether and how rapidly decentralized/asynchronous gradient updates reach consensus across distributed nodes.
    • Privacy: The extent to which privacy-sensitive raw datasets remain localized on the originating end devices.
    • Communication Cost: The transmission volume and bandwidth overhead required to exchange gradients, weights, or activations across nodes.
    • Latency: Total time required to complete training rounds, composed of compute latency on edge hardware and communication latency over the network.
    • Energy Efficiency: Total energy consumed by computation and wireless transmission on battery-constrained edge devices.
  3. Knowl 3 — Enabling Technologies for Distributed Edge DNN Training

    model/method

    Distributed DNN training across resource-constrained and heterogeneous edge nodes is enabled by six key technical categories:

    1. Federated Learning: Retains raw training datasets locally on client devices. Clients compute local Stochastic Gradient Descent (SGD) updates and upload parameter updates to a centralized aggregator (or peer nodes) which combines them via algorithms such as Federated Averaging (FedAvg) or Selective SGD (SSGD), and decentralized blockchain architectures (BlockFL).
    2. Aggregation Frequency Control: Regulates the interval between local compute steps and global parameter synchronization (e.g., the Approximate Synchronous Parallel (ASP) model in Gaia, or the FedCS client selection protocol) to balance convergence speed and communication overhead under strict bandwidth/compute budgets.
    3. Gradient Compression: Reduces transmission payloads via:
      • Gradient Quantization: Encoding high-precision floating-point gradient elements into low-precision bit representations.
      • Gradient Sparsification: Transmitting only top-kk or thresholded coordinates combined with residual/momentum error accumulation (e.g., Deep Gradient Compression, eSGD) to ensure convergence matching full-precision SGD.
    4. DNN Splitting: Divides neural network layers between devices and edge servers. The device computes early layers and transmits intermediate feature activations, protecting raw user data while offloading heavy downstream computation (e.g., Arden privacy-preserving partitioning, PipeDream pipelined model parallelism).
    5. Knowledge Transfer Learning: Pretrains a large teacher network on a comprehensive base dataset and transfers generic feature extraction layers to a compact student network deployed at the edge, substantially reducing local edge training compute and sample requirements.
    6. Gossip Training: Replaces centralized all-reduce communication with asynchronous peer-to-peer gossip algorithms (e.g., GoSGD, GossipGraD), where nodes randomly select communication partners to diffuse gradient updates with O(1)O(1) communication complexity.
  4. Knowl 4 — Comparative Taxonomy of Distributed Edge DNN Training Systems

    data/table

    The following table summarizes the architectures, Edge Intelligence (EI) levels, core optimization technologies, and empirical performance metrics of representative distributed edge DNN training systems:

    System / Framework Architecture EI Level Employed Technology Key Effectiveness
    FedAvg Hybrid Level-4 Federated Learning, Iterative model averaging Reduces communication rounds by 10×100×10\times\text{--}100\times relative to standard SGD.
    SSGD Hybrid Level-4 Federated Learning, Selective SGD Preserves client privacy while achieving higher accuracy than local-only training.
    Zoo Hybrid Level-4 Federated Learning, Composable services Constant-time processing per image despite input dataset differences.
    BlockFL Decentralized Level-6 Federated Learning, Blockchain verification Incurs 1.5%\le 1.5\% latency increase to achieve optimal block generation.
    Gaia Centralized Cloud Intel. Aggregation frequency control, ASP model Achieves 1.8×5.35×1.8\times\text{--}5.35\times speedup over conventional WAN-distributed ML systems.
    DGC N/A N/A Gradient sparsification, Quantization, Momentum correction Compresses gradients by 270×600×270\times\text{--}600\times without accuracy loss across CNNs and RNNs.
    eSGD Hybrid Level-4 Selective gradient coordinates, Momentum residual tracking Reaches 91.2%91.2\%, 86.7%86.7\%, and 81.5%81.5\% accuracy on MNIST under 50%50\%, 75%75\%, and 87.5%87.5\% gradient drop ratios.
    Inceptionn Hybrid Level-5 Lossy gradient compression, Gradient-centric aggregation Reduces communication time by 70.9%80.7%70.9\%\text{--}80.7\% with no accuracy drop.
    Arden Centralized Cloud Intel. DNN splitting, Data nullification, Noise addition Achieves 60.10%60.10\% time, 92.07%92.07\% memory, and 77.05%77.05\% energy reductions vs. standard networks.
    PipeDream Hybrid Level-5 DNN splitting, Pipeline parallelism Converges 2.5×3.1×2.5\times\text{--}3.1\times faster than a single machine and 3×3\times faster than data parallelism on VGG16.
    GoSGD Decentralized Level-6 Gossip Training Faster convergence than EASGD without centralized parameter servers.
    Gossiping SGD Decentralized Level-6 Gossip Training, Model partition Gossip step runs faster than a synchronous all-reduce step.
    GossipGraD Decentralized Level-6 Gossip Training, Model partition Achieves 100%\approx 100\% compute efficiency for ResNet-50 across 128 GPUs; reduces complexity from Θ(logp)\Theta(\log p) to O(1)O(1).

    The comparison shows that while decentralized gossip and blockchain approaches provide maximal privacy (Level 6), hybrid pipelined or sparsified approaches (Levels 4 and 5) offer the strongest trade-offs between communication reduction and model convergence speed.

  5. Knowl 5 — Edge-Centric Inference Architectural Modes and Performance Indicators

    model/method

    DNN model inference across edge environments is structured into four primary edge-centric operational modes:

    1. Edge-Based Mode: An end device collects input data (e.g., camera frames, audio) and offloads it across the network to an edge server; the edge server executes the entire DNN inference and returns the output predictions to the device.
    2. Device-Based Mode: An end device downloads the trained DNN model from the edge server and runs inference completely locally without runtime network interactions, providing high reliability and privacy at the cost of consuming on-device hardware resources.
    3. Edge-Device Mode (Collaborative Co-Inference): The DNN model is partitioned dynamically based on environmental conditions (bandwidth, compute availability). The device computes the initial layers (e.g., early convolutions) and sends intermediate feature maps to the edge server, which computes the remaining layers and returns the prediction.
    4. Edge-Cloud Mode: Used for heavily resource-constrained edge devices; the device gathers data while inference computation is split cooperatively between intermediate edge servers and cloud datacenters.

    Edge DNN inference quality is assessed along six Key Performance Indicators (KPIs):

    • Latency: End-to-end turnaround time including preprocessing, data transmission, on-device/edge execution, and postprocessing.
    • Accuracy: Ratio of correct predictions to total inputs, influenced by model precision, compression rates, confidence thresholds, and skipped video frames.
    • Energy: Total energy consumed by on-device computation and wireless radio transmission.
    • Privacy: Degree of raw sensory data protection achieved by running locally or transmitting transformed feature representations.
    • Communication Overhead: Network payload and bandwidth consumed between device, edge, and cloud.
    • Memory Footprint: Peak RAM and memory bandwidth consumed on resource-constrained devices during model storage and activation buffering.
  6. Knowl 6 — Enabling Technologies for Edge Deep Learning Inference

    model/method

    Eight foundational technological categories optimize deep neural network inference on resource-constrained edge infrastructure:

    1. Model Compression: Shrinks memory footprint and compute complexity via magnitude- or energy-aware weight pruning (e.g., EAP), fixed or adaptive precision data quantization (e.g., Number Abstract Data Type), and automated multi-technique orchestration (e.g., AdaDeep, Deep Compression).
    2. Model Partition: Splits DNN computation graphs between devices and servers (e.g., Neurosurgeon regression-based splitting, JALAD integer linear programming, graph min-cut DAG partitioning) or across peer mobile devices (e.g., MoDNN horizontal layer splitting, DeepThings vertical fused-tile partitioning).
    3. Model Early Exit: Embeds early classifier branches into intermediate hidden layers (e.g., BranchyNet, DDNNs, Edgent), allowing high-confidence inputs to terminate execution early to minimize average inference latency.
    4. Edge Caching: Caches intermediate feature embeddings or final classification results (e.g., Glimpse optical flow reuse, Cachier, FoggyCache semantic hashing) at edge nodes to fulfill subsequent matching queries with low latency.
    5. Input Filtering: Utilizes lightweight binary classifiers, specialized small networks, or temporal difference detectors (e.g., NoScope, FFS-VA, ReXCam) to discard uninformative or background video frames prior to executing deep networks.
    6. Model Selection: Pretrains a family of independent DNN models spanning diverse accuracy-resource profiles (e.g., Big/Little models, IF-CNN recognition predictors) and dynamically selects the optimal model online based on input complexity and available device resources.
    7. Support for Multi-Tenancy: Coordinates concurrent multi-task DNN execution on shared edge hardware through capacity nesting (e.g., NestDNN), compiler/runtime graph merging (e.g., HiveMind), and layer-level queuing (e.g., DeepEye).
    8. Application-Specific Optimization: Adaptively tunes domain knobs, such as video resolution and frame rate (e.g., Chameleon, DeepDecision knapsack modeling), to balance real-time resource cost against accuracy.
  7. Knowl 7 — Comparative Taxonomy of Distributed Edge DNN Inference Systems

    data/table

    The following table details the target applications, architectures, Edge Intelligence (EI) levels, optimization technologies, and performance gains of representative edge DNN inference frameworks:

    System / Framework Application Architecture / EI Level Optimization Technology Key Effectiveness
    VideoEdge Video Analytics Cloud-Edge-Device / Level-1 Frame/resolution adaptation, Multi-tenancy, Placement 5.4×25.4×5.4\times\text{--}25.4\times accuracy improvement.
    Chameleon Video Analytics Device-Cloud / Level-1 Frame/resolution adaptation, Model selection 2×3×2\times\text{--}3\times resource reduction.
    DeepX Mobile Sensing On-Device / Level-2 Runtime layer compression, Architecture decomposition 7.12×26.7×7.12\times\text{--}26.7\times energy reduction across mobile processors.
    Edgent General Device-Edge / Level-2 Model early-exit, Model partition Maximizes accuracy under given latency constraints.
    AdaDeep General On-Device / Level-3 DRL-based Model compression selection Latency: 9.8×9.8\times, Energy: 4.3×4.3\times, Storage: 38×38\times reductions.
    DeepIns Industrial IoT Edge-Cloud / Level-1 Model early-exit Latency reduction of 0.98×1.21×0.98\times\text{--}1.21\times.
    Neurosurgeon General Device-Cloud / Level-1 Regression-based Model partition 1.1×40.7×1.1\times\text{--}40.7\times speedup, 59.5%94.7%59.5\%\text{--}94.7\% energy reduction.
    Minerva General On-Device / Level-3 Hardware acceleration, Model compression 8×8\times energy saving.
    FoggyCache Industrial IoT Device-Edge / Level-2 Edge semantic caching (A-LSH, H-kkNN) 3×10×3\times\text{--}10\times latency and energy reduction.
    NoScope Video Analytics Cloud / Cloud Intel. Input filtering (difference detection, specialized models) 265×15,500×265\times\text{--}15,500\times query speedup.
    JALAD General Device-Cloud / Level-1 Model partition (ILP), Lossy feature encoding 1.25×25.1×1.25\times\text{--}25.1\times latency reduction under accuracy constraints.
    DDNNs General Cloud-Edge-Device / Level-1 Hierarchical model early-exit Substantial latency reduction over cloud-only baselines.
    FFS-VA Video Analytics On-Device / Level-3 Multistage input filtering, Multi-tenancy 3×3\times latency reduction, >7×>7\times throughput improvement.
    Cachier General Cloud-Edge / Level-1 Edge semantic caching, LFU cache replacement >3×>3\times throughput improvement.
    DeepDecision Video Analytics Cloud-Edge / Level-1 Knapsack knob-tuning, Model selection 2×10×2\times\text{--}10\times latency reduction.

    The empirical data reveals that combining multiple orthogonal techniques (such as combining early exit with partitioning in Edgent, or input filtering with multi-tenancy in FFS-VA) yields compounding gains in throughput, energy efficiency, and latency.

  8. Knowl 8 — Model Partitioning and Co-Inference Techniques for Edge Systems

    model/method

    Model partitioning divides deep neural network layer execution across heterogeneous physical nodes to minimize inference latency and device energy consumption under bandwidth and compute constraints:

    1. Server-Device Partitioning:

      • Linear Chain Partitioning (e.g., Neurosurgeon): Evaluates layer-by-layer compute and communication latencies using regression models to identify an optimal single cut point where intermediate activation sizes are small and computation can be offloaded to an edge server.
      • DAG Topologies (e.g., Hu et al., JALAD): For complex non-chain networks represented as directed acyclic graphs (DAGs), finding optimal multi-layer partition points is NP-hard. Solutions utilize integer linear programming (ILP) or graph min-cut approximation algorithms to provide guaranteed performance bounds.
      • Lossy Feature Encoding & Incremental Uploading: Compresses intermediate layer activation maps using lossy encoding before wireless transmission (e.g., JALAD, Ko et al.) or incrementally uploads partitioned layers to edge servers on demand to enable immediate collaborative querying (e.g., IONN).
    2. Device-to-Device (D2D) Cluster Partitioning:

      • Horizontal Layer Partitioning (e.g., MoDNN, MeDNN): Partitions individual convolutional layers horizontally across peer mobile devices connected via Wi-Fi Direct, achieving 1.86×4.28×1.86\times\text{--}4.28\times speedups across 2 to 4 mobile nodes.
      • Vertical Fused-Tile Partitioning (e.g., DeepThings): Slices convolutional layers vertically into spatial grid tiles through fused network layers, eliminating intermediate layer feature caching and minimizing memory footprint on memory-constrained IoT microcontrollers.
      • Heterogeneous Hardware Decomposition (e.g., DeepX, LEO): Decomposes DNN sub-blocks across heterogeneous on-chip compute units (CPU, GPU, DSP, sensing co-processors) to maximize hardware utilization and energy efficiency.
  9. Knowl 9 — Model Early-Exit Architectures for Hierarchical Edge Inference

    model/method

    Model early-exit architectures accelerate deep neural network inference by embedding auxiliary classification branches at intermediate hidden layers, allowing samples with high prediction certainty to exit early without executing deeper layers:

    • BranchyNet Mechanism: Modifies standard neural network backbones by attaching side-branch exit classifiers at intermediate positions. During inference, the entropy of normalized softmax outputs is computed at each exit branch. If the prediction confidence exceeds a preset threshold, execution terminates and the early result is returned; otherwise, intermediate activations proceed to the next layer.
    • Distributed Hierarchical Early-Exit (DDNNs): Maps early-exit points directly to the physical edge computing hierarchy:
      • Device Layer: Evaluates the earliest, low-overhead exit branch locally.
      • Edge Server Layer: Executes intermediate layers and mid-level exit points if the device exit is uncertain.
      • Cloud Layer: Executes the remaining deep layers for the hardest classification instances.
    • Multi-Device Activation Aggregation: When multiple edge devices feed intermediate activations into an upstream edge or cloud node, DDNNs aggregate feature vectors using three operations:
      1. Max Pooling (MP): Extracts the component-wise maximum across client intermediate vectors.
      2. Average Pooling (AP): Computes the component-wise arithmetic mean across vectors.
      3. Concatenation (CC): Concatenates client intermediate vectors into a single extended vector.
    • Joint Early-Exit and Partitioning Optimization (Edgent): Leverages regression-based latency prediction to jointly optimize the early-exit threshold and the physical device-edge partition boundary, maximizing accuracy under hard deadline constraints (e.g., sub-100 ms).
  10. Knowl 10 — Semantic Edge Caching for Deep Inference Acceleration

    model/method

    Semantic edge caching accelerates inference by caching and reusing intermediate feature representations or final inference results at edge servers and devices, avoiding redundant DNN evaluations for temporally or spatially correlated inputs:

    • Temporal Visual Caching (e.g., Glimpse): Caches detected object bounding boxes locally on mobile devices and tracks target movement between consecutive frames using lightweight optical flow calculations, achieving 1.6×5.5×1.6\times\text{--}5.5\times speedups in continuous object detection.
    • Edge Server Caching & Prefetching (e.g., Cachier, Precog): Maintains a feature-to-result key-value store on the edge server using Least Frequently Used (LFU) cache eviction. Precog extends this by caching data on both client and edge server while using Markov chain state predictions to prefetch anticipated results to the mobile client.
    • Compact Semantic Hashing (e.g., Shadow Puppets, FoggyCache):
      • Shadow Puppets: Employs a small auxiliary neural network to convert input images into compact semantic hash codes, outperforming standard Locality-Sensitive Hashing (LSH) by capturing human-perceptual similarity and improving lookup latency by 5×10×5\times\text{--}10\times.
      • FoggyCache: Enables cross-device approximate computation reuse for nearby devices observing identical environments by employing Adaptive Locality-Sensitive Hashing (A-LSH) to index unknown input data distributions and Homogenized kk-Nearest Neighbors (H-kkNN) to evaluate semantic similarity, reducing compute latency and energy by 3×10×3\times\text{--}10\times.
  11. Knowl 11 — Multistage Input Filtering for Edge Video Analytics

    model/method

    Input filtering reduces inference compute and bandwidth overhead in continuous edge video analytics by discarding uninformative, static, or non-target video frames before invoking resource-heavy DNN models:

    • Temporal Difference Classifiers (e.g., NoScope): Employs lightweight binary classifiers and reference-frame difference detectors to identify inter-frame pixel variance and object presence. Frames without detected changes or target classes are skipped, accelerating query processing by 265×15,500×265\times\text{--}15,500\times.
    • Multistage Filtering Pipelines (e.g., FFS-VA): Evaluates video streams across a three-stage filter hierarchy:
      1. Stream-Specialized Difference Detector (SDD): Eliminates frames containing only background scenes.
      2. Stream-Specialized Network Model (SNM): Detects the presence of specific target-object classes.
      3. Tiny-YOLO-Voc (T-YOLO): Eliminates frames where the total target object count falls below a predefined threshold.
    • Semantic Buffer Distance Filtering (e.g., Canel et al.): Extracts intermediate semantic feature vectors from video frames and stores them in a directed acyclic graph (DAG) frame buffer, using Euclidean distance metrics across embeddings to select only the top-kk most informative frames for downstream processing.
    • Spatiotemporal Cross-Camera Filtering (e.g., ReXCam): Exploits learned spatiotemporal correlations across multi-camera networks to eliminate redundant cross-view video feeds, reducing computation load by 4.6×4.6\times while boosting overall detection accuracy by 27%27\%.
  12. Knowl 12 — Open Research Challenges for Edge Intelligence Platforms and Systems

    limitation

    Six fundamental technical and architectural challenges currently limit the widespread deployment of Edge Intelligence (EI):

    1. EI as a Service (EIaaS) Software Platforms: Public cloud Machine Learning as a Service (MLaaS) focuses on server selection for cloud training, whereas EIaaS must support heterogeneous edge hardware, multi-framework model portability (TensorFlow, PyTorch, Caffe), lightweight container/function virtualization, and runtime migration across dispersive edge providers.
    2. Resource-Aware Neural Architecture Search (NAS): Existing deep learning models are over-parameterized. Designing native edge models requires AutoML and NAS techniques (guided by reinforcement learning, genetic algorithms, or Bayesian optimization) that directly incorporate edge CPU, GPU, memory, and battery constraints into the search space.
    3. Computation-Aware Networking: Incorporating 5G Ultra-Reliable Low-Latency Communication (URLLC), software-defined networking (SDN), network function virtualization (NFV), over-the-air computation, and gradient coding into edge systems to mitigate wireless channel dynamics and straggler nodes during distributed training.
    4. Dynamic Multitenancy Management: Lack of adaptive runtime orchestration algorithms capable of jointly optimizing CPU, GPU, memory bandwidth, and wireless channels when multiple concurrent DNN workloads compete on the same edge node.
    5. Decentralized Security and Privacy: Vulnerability of distributed edge environments to malicious nodes, requiring lightweight decentralized trust verification, secure routing, and formal privacy techniques (differential privacy, homomorphic encryption, secure multi-party computation) during parameter sharing.
    6. Incentive and Business Models: Designing smart pricing, fair revenue allocation, and lightweight blockchain-based consensus mechanisms to incentivize edge device owners to contribute local compute, bandwidth, and sensor data to shared EI consortiums.

Coverage note — None was omitted; the knowls fully capture the paper's core contributions including its six-level rating framework, taxonomies of training and inference architectures, key performance indicators, enabling technologies, comparative system matrices, detailed partitioning/early-exit/caching/filtering mechanisms, and identified open challenges.

References

  1. 1.Y. LeCun, Y. Bengio, and G. Hinton, —Deep learning,— Nature, vol. 521, no. 7553, p. 436, 2015.
  2. 2.L. Deng and D. Yu, —Deep learning: Methods and applications,— Found. Trends Signal Process., vol. 7, nos. 3–4, pp. 197–387, Jun. 2014.
  3. 3.Cisco Global Cloud Index: Forecast and Methodology, 2016–2021, White Paper. [Online]. Available: https://www.cisco.com/c/en/us/solutions/collateral/service-provider/global-cloud-index-gci/white-paper-c11-738085.html
  4. 4.B. Heintz, A. Chandra, and R. K. Sitaraman, —Optimizing grouped aggregation in geo-distributed streaming analytics,— in Proc. ACM HPDC, 2015, pp. 133–144.
  5. 5.Q. Pu et al., —Low latency geo-distributed data analytics,— in Proc. ACM SIGCOMM, 2015, pp. 421–434.
  6. 6.W. Shi, J. Cao, Q. Zhang, Y. Li, and L. Xu, —Edge computing: Vision and challenges,— IEEE Internet Things J., vol. 3, no. 5, pp. 637–646, Oct. 2016.
  7. 7.X. Chen, L. Pu, L. Gao, W. Wu, and D. Wu, —Exploiting massive D2D collaboration for energy-efficient mobile edge computing,— IEEE Wireless Commun., vol. 24, no. 4, pp. 64–71, Aug. 2017.
  8. 8.Y. Mao, C. You, J. Zhang, K. Huang, and K. B. Letaief, —A survey on mobile edge computing: The communication perspective,— IEEE Commun. Surveys Tuts., vol. 19, no. 4, pp. 2322–2358, 4th Quart., 2017.
  9. 9.X. Wang, Y. Han, C. Wang, Q. Zhao, X. Chen, and M. Chen, —In-edge Ai: Intelligentizing mobile edge computing, caching and communication by federated learning,— 2018, arXiv:1809.07857. [Online]. Available: https://arxiv.org/abs/1809.07857
  10. 10.E. Li, Z. Zhou, and X. Chen, —Edge intelligence: On-demand deep learning model co-inference with device-edge synergy,— in Proc. Workshop Mobile Edge Commun. (MECOMM), 2018, pp. 31–36.
  11. 11.(2018). 5 Trends Emerge in the Gartner Hype Cycle for Emerging Technologies. [Online]. Available: https://www.gartner.com/smarterwithgartner/5-trends-emerge-in-gartner-hype-cycle-for-emerging-technologies-2018/
  12. 12.G. Ananthanarayanan et al., —Real-time video analytics: The killer app for edge computing,— Computer, vol. 50, no. 10, pp. 58–67, 2017.
  13. 13.K. Ha, Z. Chen, W. Hu, W. Richter, P. Pillai, and M. Satyanarayanan, —Towards wearable cognitive assistance,— in Proc. ACM Mobisys, 2014, pp. 68–81.
  14. 14.C. Jie, L. Xu, R. Abdallah, and W. Shi, —EdgeOS_h: A home operating system for Internet of everything,— in Proc. IEEE ICDCS, Jun. 2017, pp. 1756–1764.
  15. 15.L. Li, K. Ota, and M. Dong, —Deep learning for smart industry: Efficient manufacture inspection system with fog computing,— IEEE Trans. Ind. Informat., vol. 14, no. 10, pp. 4665–4673, Oct. 2018.
  16. 16.D. Svozil, V. Kvasnicka, and J. Pospichal, —Introduction to multi-layer feed-forward neural networks,— Chemometrics Intell. Lab. Syst., vol. 39, no. 1, pp. 43–62, 1997.
  17. 17.R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa, —Natural language processing (almost) from scratch,— J. Mach. Learn. Res., vol. 12 pp. 2493–2537, Aug. 2011.
  18. 18.A. Krizhevsky, I. Sutskever, and G. E. Hinton, —Imagenet classification with deep convolutional neural networks,— in Proc. NIPS, 2012, pp. 1097–1105.
  19. 19.K. Simonyan and A. Zisserman, —Very deep convolutional networks for large-scale image recognition,— 2014, arXiv:1409.1556. [Online]. Available: https://arxiv.org/abs/1409.1556
  20. 20.K. He, X. Zhang, S. Ren, and J. Sun, —Deep residual learning for image recognition,— in Proc. IEEE CVPR, Jun. 2016, pp. 770–778.
  21. 21.A. G. Howard et al., —MobileNets: Efficient convolutional neural networks for mobile vision applications,— 2017, arXiv:1704.04861. [Online]. Available: https://arxiv.org/abs/1704.04861
  22. 22.H. Mao, S. Yao, T. Tang, B. Li, J. Yao, and Y. Wang, —Towards real-time object detection on embedded systems,— IEEE Trans. Emerging Topics Comput., vol. 6, no. 3, pp. 417–431, Aug. 2018.
  23. 23.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, —You only look once: Unified, real-time object detection,— in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2016, pp. 779–788.
  24. 24.W. Liu et al., —SSD: Single shot multibox detector,— in Proc. Eur. Conf. Comput. Vis. 2016, pp. 21–37.
  25. 25.L. Bottou, —Large-scale machine learning with stochastic gradient descent,— in Proc. COMPSTAT 2010, pp. 177–186.
  26. 26.D. E. Rumelhart, G. E. Hinton, and R. J. Williams, —Learning representations by back-propagating errors,— Nature, vol. 323, no. 6088, p. 533, 1986.
  27. 27.Y. Chauvin and D. E. Rumelhart, Backpropagation: Theory, Architectures, and Applications. Psychology Press, 2013.
  28. 28.C. Szegedy et al., —Going deeper with convolutions,— in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2015, pp. 1–9.
  29. 29.P. J. Werbos, —Backpropagation through time: What it does and how to do it,— Proc. IEEE, vol. 78, no. 10, pp. 1550–1560, Oct. 1990.
  30. 30.S. Hochreiter and J. Schmidhuber, —Long short-term memory,— Neural Comput., vol. 9, no. 8, pp. 1735–1780, 1997.
  31. 31.I. Goodfellow et al., —Generative adversarial nets,— in Proc. Adv. Neural Inf. Process. Syst., 2014, pp. 2672–2680.
  32. 32.V. Mnih et al., —Human-level control through deep reinforcement learning,— Nature, vol. 518, no. 7540, p. 529, 2015.
  33. 33.3 AI Trends for Enterprise Computing. [Online]. Available: https://www.gartner.com/smarterwithgartner/3-ai-trends-for-enterprise-computing/
  34. 34.Democratizing AI. [Online]. Available: https://news.microsoft.com/features/democratizing-ai/
  35. 35.M. Satyanarayanan, P. Bahl, R. Caceres, and N. Davies, —The case for VM-based cloudlets in mobile computing,— IEEE Pervasive Comput., no. 4, pp. 14–23, Oct. 2009.
  36. 36.Microsoft Interactive Cloud Gaming. [Online]. Available: https://azure.microsoft.com/en-us/solutions/gaming/
  37. 37.H. Zhang, G. Ananthanarayanan, P. Bodik, M. Philipose, P. Bahl, and M. J. Freedman, —Live video analytics at scale with approximation and delay-tolerance,— in Proc. USENIX NSDI, 2017, pp. 377–392.
  38. 38.C.-C. Hung et al., —VideoEdge: Processing camera streams using hierarchical clusters,— in Proc. IEEE/ACM Symp. Edge Comput. (SEC), Oct. 2018, pp. 115–131.
  39. 39.I. Stoica et al., —A Berkeley view of systems challenges for AI,— 2017, arXiv:1712.05855. [Online]. Available: https://arxiv.org/abs/1712.05855
  40. 40.(2018). 5 Trends Emerge in the Gartner Hype Cycle for Emerging Technologies. [Online]. Available: https://www.gartner.com/smarterwithgartner/5-trends-emerge-in-gartner-hype-cycle-for-emerging-technologies-2018/
  41. 41.
    1. IEC White Paper Edge Intelligence. [Online]. Available: https://www.iec.ch/whitepaper/edgeintelligence/
  42. 42.Accelerating AI on the Intelligent Edge. [Online]. Available: https://azure.microsoft.com/en-us/blog/accelerating-ai-on-the-intelligent-edge-microsoft-and-qualcomm-create-vision-ai-developer-kit/
  43. 43.Edge Intelligence for Industrial Internet of Things. [Online]. Available: https://www.comsoc.org/publications/magazines/ieee-network/cfp/edge-intelligence-industrial-internet-things
  44. 44.Q. Chen, Z. Zheng, C. Hu, D. Wang, and F. Liu, —Data-driven task allocation for multi-task transfer learning on the edge,— in Proc. IEEE 39th Int. Conf. Distrib. Comput. Syst. (ICDCS), 2019.
  45. 45.H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. Y. Arcas, —Communication-efficient learning of deep networks from decentralized data,— 2016, arXiv:1602.05629. [Online]. Available: https://arxiv.org/abs/1602.05629
  46. 46.R. Shokri and V. Shmatikov, —Privacy-preserving deep learning,— in Proc. 22nd ACM SIGSAC Conf. Comput. Commun. Secur., 2015, pp. 1310–1321.
  47. 47.J. Konećny, H. B. McMahan, F. X. Yu, P. Richt‡rik, A. T. Suresh, and D. Bacon, —Federated learning: Strategies for improving communication efficiency,— 2016, arXiv:1610.05492.
  48. 48.A. Lalitha, S. Shekhar, T. Javidi, and F. Koushanfar, —Peer-to-peer federated learning on graphs,— 2019, arXiv:1901.11173. [Online]. Available: https://arxiv.org/abs/1901.11173
  49. 49.H. Kim, J. Park, M. Bennis, and S.-L. Kim, —On-device federated learning via blockchain and its latency analysis,— 2018, arXiv:1808.03949. [Online]. Available: https://arxiv.org/abs/1808.03949
  50. 50.K. Hsieh, A. Harlap, N. Vijaykumar, D. Konomis, G. R. Ganger, and P. B. Gibbons, —Gaia: Geo-distributed machine learning approaching LAN speeds,— in Proc. NSDI, 2017, pp. 629–647.
  51. 51.S. Wang et al., —Adaptive federated learning in resource constrained edge computing systems,— IEEE J. Sel. Areas Commun., vol. 37, no. 3, pp. 1205–1221, Jun. 2019.
  52. 52.T. Nishio and R. Yonetani, —Client selection for federated learning with heterogeneous resources in mobile edge,— 2018, arXiv:1804.08333. [Online]. Available: https://arxiv.org/abs/1804.08333
  53. 53.Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally, —Deep gradient compression: Reducing the communication bandwidth for distributed training,— 2017, arXiv:1712.01887. [Online]. Available: https://arxiv.org/abs/1712.01887
  54. 54.Z. Tao and Q. Li, —eSGD: Communication efficient distributed deep learning on the edge,— in Proc. USENIX Workshop Hot Topics Edge Comput. (HotEdge), Boston, MA, USA, 2018.
  55. 55.S. U. Stich, J.-B. Cordonnier, and M. Jaggi, —Sparsified sgd with memory,— in Proc. Adv. Neural Inf. Process. Syst., 2018, pp. 4452–4463.
  56. 56.H. Tang, S. Gan, C. Zhang, T. Zhang, and J. Liu, —Communication compression for decentralized training,— in Proc. Adv. Neural Inf. Process. Syst., 2018, pp. 7663–7673.
  57. 57.M. M. Amiri and D. Gunduz, —Machine learning at the wireless edge: Distributed stochastic gradient descent over-the-air,— 2019, arXiv:1901.00844. [Online]. Available: https://arxiv.org/abs/1901.00844
  58. 58.Y. Mao, S. Yi, Q. Li, J. Feng, F. Xu, and S. Zhong, —A privacy-preserving deep learning approach for face recognition with edge computing,— in Proc. USENIX, 2018.
  59. 59.J. Wang, J. Zhang, W. Bao, X. Zhu, B. Cao, and P. S. Yu, —Not just privacy: Improving performance of private deep learning in mobile cloud,— in Proc. 24th ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, 2018, pp. 2407–2416.
  60. 60.S. A. Osia et al., —A hybrid deep learning architecture for privacy-preserving mobile analytics,— 2017, arXiv:1703.02952. [Online]. Available: https://arxiv.org/abs/1703.02952
  61. 61.A. Harlap, D. Narayanan, A. Phanishayee, V. Seshadri, G. R. Ganger, and P. B. Gibbons, —PipeDream: Fast and efficient pipeline parallel DNN training,— 2018, arXiv:1806.03377. [Online]. Available: https://arxiv.org/abs/1806.03377
  62. 62.R. Sharma, S. Biookaghazadeh, B. Li, and M. Zhao, —Are existing knowledge transfer techniques effective for deep learning with edge devices?— in Proc. IEEE Int. Conf. Edge Comput. (EDGE), Jul. 2018, pp. 42–49.
  63. 63.S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah, —Randomized gossip algorithms,— IEEE Trans. Inf. Theory, vol. 52, no. 6, pp. 2508–2530, Jun. 2006.
  64. 64.M. Blot, D. Picard, M. Cord, and N. Thome, —Gossip training for deep learning,— 2016, arXiv:1611.09726. [Online]. Available: https://arxiv.org/abs/1611.09726
  65. 65.P. H. Jin, Q. Yuan, F. Iandola, and K. Keutzer, —How to scale distributed deep learning?— 2016, arXiv:1611.04581. [Online]. Available: https://arxiv.org/abs/1611.04581
  66. 66.J. Daily, A. Vishnu, C. Siegel, T. Warfel, and V. Amatya, —GossipGraD: Scalable deep learning using gossip communication based asynchronous gradient descent,— 2018, arXiv:1803.05880. [Online]. Available: https://arxiv.org/abs/1803.05880
  67. 67.J. Zhao, R. Mortier, J. Crowcroft, and L. Wang, —Privacy-preserving machine learning based data analytics on edge devices,— in Proc. AIES, 2018, pp. 341–346.
  68. 68.Y. Li et al., —A network-centric hardware/algorithm co-design to accelerate distributed training of deep neural networks,— in Proc. 51st Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), Oct. 2018, pp. 175–188.
  69. 69.C.-J. Wu et al., —Machine learning at facebook: Understanding inference at the edge,— in Proc. IEEE Int. Symp. High Perform. Comput. Archit. (HPCA), Feb. 2019, pp. 331–344.
  70. 70.S. Han, J. Pool, J. Tran, and W. Dally, —Learning both weights and connections for efficient neural network,— in Proc. Adv. Neural Inf. Process. Syst., 2015, pp. 1135–1143.
  71. 71.S. Han, H. Mao, and W. J. Dally, —Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding,— 2015, arXiv:1510.00149. [Online]. Available: https://arxiv.org/abs/1510.00149
  72. 72.Y.-H. Chen, J. Emer, and V. Sze, —Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,— ACM SIGARCH Comput. Archit. News, vol. 44, no. 3, pp. 367–379, 2016.
  73. 73.T.-J. Yang, Y.-H. Chen, and V. Sze, —Designing energy-efficient convolutional neural networks using energy-aware pruning,— in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2017, pp. 5687–5695.
  74. 74.B. Reagen et al., —Minerva: Enabling low-power, highly-accurate deep neural network accelerators,— ACM SIGARCH Comput. Archit. News, vol. 44, no. 3, pp. 267–278, 2016.
  75. 75.S. Liu, Y. Lin, Z. Zhou, K. Nan, H. Liu, and J. Du, —On-demand deep model compression for mobile devices: A usage-driven model selection framework,— in Proc. 16th Annu. Int. Conf. Mobile Syst., Appl., Services, 2018, pp. 389–400.
  76. 76.Y. H. Oh et al., —A portable, automatic data qantizer for deep neural networks,— in Proc. ACM PACT, 2018, p. 17.
  77. 77.N. D. Lane et al., —DeepX: A software accelerator for low-power deep learning inference on mobile devices,— in Proc. 15th Int. Conf. Inf. Process. Sensor Netw., 2016, p. 23.
  78. 78.L. Zeng, E. Li, Z. Zhou, and X. Chen, —Boomerang: On-demand cooperative deep neural network inference for edge intelligence on industrial Internet of Things,— IEEE Netw., 2019.
  79. 79.Y. Kang et al., —Neurosurgeon: Collaborative intelligence between the cloud and mobile edge,— ACM SIGPLAN Notices, vol. 52, no. 4, pp. 615–629, 2017.
  80. 80.C. Hu, W. Bao, D. Wang, and F. Liu, —Dynamic adaptive DNN surgery for inference acceleration on the edge,— in Proc. IEEE INFOCOM, 2019.
  81. 81.H. Li, C. Hu, J. Jiang, Z. Wang, Y. Wen, and W. Zhu, —JALAD: Joint accuracy-and latency-aware deep structure decoupling for edge-cloud execution,— 2018, arXiv:1812.10027. [Online]. Available: https://arxiv.org/abs/1812.10027
  82. 82.H.-J. Jeong, H.-J. Lee, C. H. Shin, and S.-M. Moon, —IONN: Incremental offloading of neural network computations from mobile devices to edge servers,— in Proc. ACM Symp. Cloud Comput., 2018, pp. 401–411.
  83. 83.J. H. Ko, T. Na, M. F. Amir, and S. Mukhopadhyay, —Edge-host partitioning of deep neural networks with feature space encoding for resource-constrained Internet-of-Things platforms,— 2018, arXiv:1802.03835. [Online]. Available: https://arxiv.org/abs/1802.03835
  84. 84.J. Mao et al., —MeDNN: A distributed mobile system with enhanced partition and deployment for large-scale DNNs,— in Proc. 36th Int. Conf. Comput.-Aided Design, Nov. 2017, pp. 751–756.
  85. 85.J. Mao, X. Chen, K. W. Nixon, C. Krieger, and Y. Chen, —MoDNN: Local distributed mobile computing system for deep neural network,— in Proc. Design, Autom. Test Eur. Conf. Exhib. (DATE), 2017, pp. 1396–1401.
  86. 86.Z. Zhao, K. M. Barijough, and A. Gerstlauer, —DeepThings: Distributed adaptive deep learning inference on resource-constrained IoT edge clusters,— IEEE Trans. Comput.-Aided Design Integr. Circuits Syst., vol. 37, no. 11, pp. 2348–2359, Nov. 2018.
  87. 87.S. Teerapittayanon, B. McDanel, and H. Kung, —BranchyNet: Fast inference via early exiting from deep neural networks,— in Proc. 23rd Int. Conf. Pattern Recognit. (ICPR), 2016, pp. 2464–2469.
  88. 88.S. Teerapittayanon, B. McDanel, and H.-T. Kung, —Distributed deep neural networks over the cloud, the edge and end devices,— in Proc. IEEE 37th Int. Conf. Distrib. Comput. Syst. (ICDCS), Jun. 2017, pp. 328–339.
  89. 89.T. Bolukbasi, J. Wang, O. Dekel, and V. Saligrama, —Adaptive neural networks for efficient inference,— 2017, arXiv:1702.07811. [Online]. Available: https://arxiv.org/abs/1702.07811
  90. 90.C. Lo, Y.-Y. Su, C.-Y. Lee, and S.-C. Chang, —A dynamic deep neural network design for efficient workload allocation in edge computing,— in Proc. IEEE Int. Conf. Comput. Design (ICCD), Nov. 2017, pp. 273–280.
  91. 91.S. Leroux et al., —The cascading neural network: Building the Internet of smart things,— Knowl. Inf. Syst., vol. 52, no. 3, pp. 791–814, 2017.
  92. 92.T. Y.-H. Chen, L. Ravindranath, S. Deng, P. Bahl, and H. Balakrishnan, —Glimpse: Continuous, real-time object recognition on mobile devices,— in Proc. ACM Sensys, 2015.
  93. 93.U. Drolia, K. Guo, J. Tan, R. Gandhi, and P. Narasimhan, —Cachier: Edge-caching for recognition applications,— in Proc. IEEE ICDCS, Jun. 2017, pp. 276–286.
  94. 94.U. Drolia, K. Guo, and P. Narasimhan, —Precog: Prefetching for image recognition applications at the edge,— in Proc. ACM/IEEE Symp. Edge Comput. (SEC), Oct. 2017, p. 17.
  95. 95.P. Guo, B. Hu, R. Li, and W. Hu, —FoggyCache: Cross-device approximate computation reuse,— in Proc. ACM Mobicom, 2018, pp. 19–34.
  96. 96.S. Venugopal, M. Gazzetti, Y. Gkoufas, and K. Katrinis, —Shadow puppets: Cloud-level accurate AI inference at the speed and economy of edge,— in Proc. USENIX Workshop Hot Topics in Edge Comput. (HotEdge), 2018.
  97. 97.D. Kang, J. Emmons, F. Abuzaid, P. Bailis, and M. Zaharia, —Noscope: Optimizing neural network queries over video at scale,— Proc. VLDB Endowment, vol. 10, no. 11, pp. 1586–1597, 2017.
  98. 98.J. Wang et al., —Bandwidth-efficient live video analytics for drones via edge computing,— in Proc. IEEE/ACM Symp. Edge Comput. (SEC), Oct. 2018, pp. 159–173.
  99. 99.S. Jain, J. Jiang, Y. Shu, G. Ananthanarayanan, and J. Gonzalez, —ReXCam: Resource-efficient, cross-camera video analytics at enterprise scale,— 2018, arXiv:1811.01268. [Online]. Available: https://arxiv.org/abs/1811.01268
  100. 100.C. Zhang, Q. Cao, H. Jiang, W. Zhang, J. Li, and J. Yao, —FFS-VA: A fast filtering system for large-scale video analytics,— in Proc. ACM ICPP, 2018, p. 85.
  101. 101.C. Canel et al., —Picking interesting frames in streaming video,— to be published.
  102. 102.E. Park et al., —Big/little deep neural network for ultra low power inference,— in Proc. 10th Int. Conf. Hardw./Softw. Codesign Syst. Synth., 2015, pp. 124–132.
  103. 103.B. Taylor, V. S. Marco, W. Wolff, Y. Elkhatib, and Z. Wang, —Adaptive deep learning model selection on embedded systems,— in Proc. ACM LCTES, 2018, pp. 31–43.
  104. 104.J. Jiang, G. Ananthanarayanan, P. Bodik, S. Sen, and I. Stoica, —Chameleon: Scalable adaptation of video analytics,— in Proc. ACM SIGCOMM, 2018, pp. 253–266.
  105. 105.G. Shu, W. Liu, X. Zheng, and J. Li, —IF-CNN: Image-aware inference framework for CNN with the collaboration of mobile devices and cloud,— IEEE Access, vol. 6, pp. 621–633, 2018.
  106. 106.D. Stamoulis et al., —Designing adaptive neural networks for energy-constrained image classification,— in Proc. ACM ICCAD, 2018, Art. no. 23.
  107. 107.B. Fang, X. Zeng, and M. Zhang, —NestDNN: Resource-aware multi-tenant on-device deep learning for continuous mobile vision,— in Proc. ACM Mobicom, 2018, pp. 115–127.
  108. 108.A. Mathur, N. D. Lane, S. Bhattacharya, A. Boran, C. Forlivesi, and F. Kawsar, —DeepEye: Resource efficient local execution of multiple deep vision models using wearable commodity hardware,— in Proc. ACM Mobisys, 2017, pp. 68–81.
  109. 109.Z. Fang, M. Luo, T. Yu, O. J. Mengshoel, M. B. Srivastava, and R. K. Gupta, —Mitigating multi-tenant interference in continuous mobile offloading,— in Proc. Int. Conf. Cloud Comput. Springer, 2018, pp. 20–36.
  110. 110.A. H. Jiang et al., —Mainstream: Dynamic stem-sharing for multi-tenant video processing,— in Proc. USENIX ATC, 2018, pp. 29–42.
  111. 111.D. Narayanan, K. Santhanam, A. Phanishayee, and M. Zaharia, —Accelerating deep learning workloads through efficient multi-model execution,— in Proc. NIPS Workshop Syst. Mach. Learn., Dec. 2018.
  112. 112.X. Ran, H. Chen, X. Zhu, Z. Liu, and J. Chen, —Deepdecision: A mobile deep learning framework for edge video analytics,— in Proc. IEEE INFOCOM, Apr. 2018, pp. 1421–1429.
  113. 113.P. Georgiev, N. D. Lane, K. K. Rachuri, and C. Mascolo, —Leo: Scheduling sensor inference algorithms across heterogeneous mobile processors and network resources,— in Proc. 22nd Annu. Int. Conf. Mobile Computing Netw., 2016, pp. 320–333.
  114. 114.V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, —Efficient processing of deep neural networks: A tutorial and survey,— Proc. IEEE, vol. 105, no. 12, pp. 2295–2329, Dec. 2017.
  115. 115.X. Zhang, Y. Wang, and W. Shi, —pCAMP: Performance comparison of machine learning packages on the edges,— in Proc. USENIX Workshop Hot Topics Edge Comput. (HotEdge), 2018.
  116. 116.Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han, —AMC: Automl for model compression and acceleration on mobile devices,— in Proc. Eur. Conf. Comput. Vis. Springer, 2018, pp. 815–832.
  117. 117.B. Zoph and Q. V. Le, —Neural architecture search with reinforcement learning,— 2016, arXiv:1611.01578. [Online]. Available: https://arxiv.org/abs/1611.01578
  118. 118.R. Tandon, Q. Lei, A. G. Dimakis, and N. Karampatziakis, —Gradient coding: Avoiding stragglers in distributed learning,— in Proc. Int. Conf. Mach. Learn., 2017, pp. 3368–3376.
  119. 119.G. Zhu, Y. Wang, and K. Huang, —Low-latency broadband analog aggregation for federated edge learning,— 2018, arXiv:1812.11494. [Online]. Available: https://arxiv.org/abs/1812.11494
  120. 120.J. Huang et al., —Speed/accuracy trade-offs for modern convolutional object detectors,— in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Jul. 2017, pp. 7310–7311.
  121. 121.D. Li, Z. Zhang, W. Liao, and Z. Xu, —KLRA: A kernel level resource auditing tool for IoT operating system security,— in Proc. IEEE/ACM Symp. Edge Comput. (SEC), Oct. 2018, pp. 427–432.
  122. 122.M. Du, K. Wang, Y. Chen, X. Wang, and Y. Sun, —Big data privacy preserving in multi-access edge computing for heterogeneous Internet of Things,— IEEE Commun. Mag., vol. 56, no. 8, pp. 62–67, Aug. 2018.

Citation

MLA
Zhou, Z., et al. “Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing”. arXiv, 2019, http://arxiv.org/abs/1905.10083v1.
APA
Zhou, Z., Chen, X., Li, E., Zeng, L., Luo, K., & Zhang, J. (2019). Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing. arXiv. http://arxiv.org/abs/1905.10083v1
Chicago
Zhou, Z., X. Chen, E. Li, L. Zeng, K. Luo, and J. Zhang. 2019. “Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing”. arXiv. http://arxiv.org/abs/1905.10083v1.
Harvard
Zhou, Z. et al. (2019) “Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1905.10083v1.
Vancouver
1. Zhou Z, Chen X, Li E, Zeng L, Luo K, Zhang J (2019) Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing. arXiv

BibTeX

@article{zhou2019edge,
  title = {Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing},
  author = {Zhou, Zhi and Chen, Xu and Li, En and Zeng, Liekang and Luo, Ke and Zhang, Junshan},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1905.10083v1},
  eprint = {1905.10083}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF