Applications of Deep Reinforcement Learning in Communications and Networking: A Survey

Nguyen Cong LuongDinh Thai HoangShimin GongDusit NiyatoPing WangYing-Chang LiangDong In Kim

article2018IEEE Communications Surveys and Tutorials1,838 citations

Surveys the application of deep reinforcement learning across modern communication networks, providing a structured analysis of how advanced models solve complex decision-making problems in dynamic spectrum access, wireless caching, traffic routing, and network security.

Listen

Modern wireless communication systems, such as cellular networks, the Internet of Things, and autonomous drone networks, are increasingly decentralized, dynamic, and complex. Network entities like mobile devices, base stations, and sensors must make real-time decisions regarding spectrum access, data transmission rates, content caching, computation offloading, and cybersecurity defenses without complete knowledge of their operating environment. While traditional reinforcement learning allows devices to learn optimal behaviors through trial and error, it struggles with slow convergence and computational unmanageability in large-scale systems with massive decision spaces. This article set out to evaluate how deep reinforcement learningwhich combines reinforcement learning with deep neural networkscan overcome these scalability challenges and optimize autonomous decision-making across diverse networking applications.

The article conducts a comprehensive review and taxonomy of deep reinforcement learning applications across communications and networking, examining algorithmic frameworks such as deep Q-learning, double deep Q-networks, dueling architectures, and actor-critic models. It synthesizes findings from extensive studies, simulations, and real-world datasets across cellular infrastructure, cognitive radio, heterogeneous networks, vehicular networks, and satellite systems. Through comparative analysis, the article evaluates how these techniques solve complex sequential decision processes and multi-agent game-theoretic formulations in uncertain, time-varying environments without requiring complete channel models or centralized coordination.

The findings demonstrate that deep reinforcement learning consistently achieves near-optimal performance across critical networking domains while drastically accelerating learning speed compared to traditional reinforcement learning. In dynamic spectrum access and user association, deep learning algorithms increased network throughput by 24% to 28% over standard methods and achieved up to double the throughput of traditional random access protocols. For adaptive video streaming and traffic rate control, deep reinforcement learning architectures improved user Quality of Experience by up to 25% and reduced video freezing by roughly 33%. In mobile edge computing and wireless caching, deep learning strategies improved cache hit rates and cut operational delay and energy costs by up to 55% compared to static allocation policies. In network security and anti-jamming scenarios, deep reinforcement learning converged up to 83% faster than standard methods, cutting transmission error rates by approximately 47% and enhancing communication secrecy.

These findings indicate that integrating deep neural networks into autonomous decision-making provides a viable pathway toward self-organizing next-generation communication architectures. By enabling decentralized network entities to learn optimal policies locally with minimal information exchange, deep reinforcement learning reduces communication overhead, enhances operational robustness, and lowers latency. For network operators and technology providers, this translates into reduced infrastructure energy costs, improved spectrum efficiency, and superior service reliability without the need to solve computationally prohibitive optimization problems in real time.

Decision-makers and network engineers should consider deploying deep reinforcement learning incrementally, prioritizing high-impact applications such as proactive edge caching and adaptive video bitrate control where performance gains are substantial. Organizations should establish hybrid architectures where computationally intensive training occurs on centralized or edge servers while local agents execute low-complexity inference. However, stakeholders should remain cautious: deep reinforcement learning models often depend heavily on the quality and stability of training data, demand significant hardware resources, and can experience instability in rapidly shifting environments. Further validation through large-scale physical testbeds and real-world pilot deployments is recommended before deploying fully autonomous control mechanisms in mission-critical networks.

arXiv: 1810.07862
Cover for Applications of Deep Reinforcement Learning in Communications and Networking: A Survey

Abstract

This paper presents a comprehensive literature review on applications of deep reinforcement learning in communications and networking. Modern networks, e.g., Internet of Things (IoT) and Unmanned Aerial Vehicle (UAV) networks, become more decentralized and autonomous. In such networks, network entities need to make decisions locally to maximize the network performance under uncertainty of network environment. Reinforcement learning has been efficiently used to enable the network entities to obtain the optimal policy including, e.g., decisions or actions, given their states when the state and action spaces are small. However, in complex and large-scale networks, the state and action spaces are usually large, and the reinforcement learning may not be able to find the optimal policy in reasonable time. Therefore, deep reinforcement learning, a combination of reinforcement learning with deep learning, has been developed to overcome the shortcomings. In this survey, we first give a tutorial of deep reinforcement learning from fundamental concepts to advanced models. Then, we review deep reinforcement learning approaches proposed to address emerging issues in communications and networking. The issues include dynamic network access, data rate control, wireless caching, data offloading, network security, and connectivity preservation which are all important to next generation networks such as 5G and beyond. Furthermore, we present applications of deep reinforcement learning for traffic routing, resource sharing, and data collection. Finally, we highlight important challenges, open issues, and future research directions of applying deep reinforcement learning.

Table of Contents

  • I Introduction
  • II Deep Reinforcement Learning: An Overview
  • II-A Markov Decision Processes
  • II-A1 Partially Observable Markov Decision Process
  • II-A2 Markov Games
  • II-B Reinforcement Learning
  • II-B1 QQ-Learning Algorithm
  • II-B2 SARSA: An Online Q-Learning Algorithm
  • II-B3 Q-Learning for Markov Games
  • II-C Deep Learning
  • II-D Deep QQ-Learning
  • II-E Advanced Deep QQ-Learning Models
  • II-E1 Double Deep QQ-Learning
  • II-E2 Deep QQ-Learning with Prioritized Experience Replay
  • II-E3 Dueling Deep QQ-Learning
  • II-E4 Asynchronous Multi-step Deep Q-Learning
  • II-E5 Distributional Deep Q-learning
  • II-E6 Deep QQ-learning with Noisy Nets
  • II-E7 Rainbow Deep QQ-learning
  • II-F Deep Q-Learning for Extensions of MDPs
  • II-F1 Deep Deterministic Policy Gradient Q-Learning for Continuous Action
  • II-F2 Deep Recurrent Q-Learning for POMDPs
  • II-F3 Deep SARSA Learning
  • II-F4 Deep QQ-Learning for Markov Games
  • III Network Access and Rate Control
  • III-A Network Access
  • III-A1 Dynamic Spectrum Access
  • III-A2 Joint User Association and Spectrum Access
  • III-B Adaptive Rate Control
  • IV Caching and Offloading
  • IV-A Wireless Proactive Caching
  • IV-A1 QoS-Aware Caching
  • IV-A2 Joint Caching and Transmission Control
  • IV-A3 Joint Caching, Networking, and Computation
  • IV-B Data and Computation Offloading
  • V Network Security and Connectivity Preservation
  • V-A Network Security
  • V-A1 Jamming Attack
  • V-A2 Cyber-Physical Attack
  • V-B Connectivity Preservation
  • VI Miscellaneous Issues
  • VI-A Traffic Engineering and Routing
  • VI-B Resource Sharing and Scheduling
  • VI-C Power Control and Data Collection
  • VII Challenges, Open Issues, and Future Research Directions
  • VII-A Challenges
  • VII-A1 State Determination in Density Networks
  • VII-A2 Knowledge of Jammers’ Channel Information
  • VII-A3 Multi-agent DRL in Dynamic HetNets
  • VII-A4 Training and Performance Evaluation of DRL Framework
  • VII-B Open Issues
  • VII-B1 Distributed DRL Framework in Wireless Networks
  • VII-B2 Balance between Information Quality and Learning Performance
  • VII-C Future Research Directions
  • VII-C1 DRL for Channel Estimation in Wireless Systems
  • VII-C2 DRL for Crowdsensing Service Optimization
  • VII-C3 DRL for Cryptocurrency Management in Wireless Networks
  • VII-C4 DRL for Auction
  • VIII Conclusions
  • References

Knowls

  1. Knowl 1 — Taxonomy of Deep Reinforcement Learning in Communications and Networking

    model/method

    A comprehensive taxonomy of Deep Reinforcement Learning (DRL) applications across modern communication and networking systems categorizes decision and optimization problems into four primary domains:

    1. Network Access and Rate Control:

      • Dynamic Spectrum Access: Distributed multi-channel selection under partial observability without explicit transition probability matrices.
      • Joint User Association and Spectrum Allocation: Non-convex matching of mobile users to base stations (BSs) and frequency bands.
      • Adaptive Rate Control: Dynamic bitrate selection for Dynamic Adaptive Streaming over HTTP (DASH) and High Volume Flexible Time (HVFT) cellular IoT traffic.
    2. Caching and Computation Offloading:

      • Wireless Proactive Caching: Edge content replacement, Time-To-Live (TTL) expiration learning, and joint caching and interference alignment / transmission control.
      • Computation Offloading: Multi-user mobile edge computing (MEC) workload partitioning, base station sleep/(de)activation control, and container migration in fog networks.
    3. Network Security and Connectivity Preservation:

      • Anti-Jamming: Frequency hopping, power control, and relay-assisted communications under dynamic radio frequency (RF) jamming attacks.
      • Cyber-Physical Security: Sensor weight adaptation in autonomous vehicle platooning and dynamic watermarking authentication against data injection attacks in IoT.
      • Connectivity Preservation: Velocity/trajectory optimization in multi-robot/UAV formations and handover control in ultra-dense networks (UDNs).
    4. Miscellaneous Network Operations:

      • Traffic Engineering and Routing: Model-free network utility maximization and delay-sensitive path splitting in Software-Defined Networks (SDNs).
      • Resource Sharing and Slicing: Virtualized Network Function (VNF) chaining, radio slicing, and cloud-RAN beamforming.
      • Power Control and Crowdsensing: Power allocation in massive MIMO and incentive/payment design in mobile crowdsensing under strategic user behavior.
  2. Knowl 2 — Deep Q-Learning with Experience Replay and Fixed Target Network

    algorithm

    Standard Q-learning with non-linear function approximation suffers from divergence due to correlated data samples and non-stationary target values. Deep Q-Learning (DQL) stabilizes training of a Deep Q-Network (DQN) with weights θ\theta through two key mechanisms:

    1. Experience Replay: Transition tuples (st,at,rt,st+1)(s_t, a_t, r_t, s_{t+1}) are saved into a replay buffer D\mathcal{D}. Mini-batches of transitions are sampled uniformly to break temporal autocorrelation.
    2. Fixed Target Network: A secondary target network Q^\hat{Q} parameterized by θ\theta' computes target Q-values. Its weights θ\theta' are periodically updated by copying the online network parameters (i.e., θθ\theta' \leftarrow \theta) every fixed number of steps.

    The loss function minimized via stochastic gradient descent with respect to θ\theta is:

    L(θ)=E(sj,aj,rj,sj+1)D[(rj+γmaxaQ^(sj+1,a;θ)Q(sj,aj;θ))2]L(\theta) = \mathbb{E}_{(s_j, a_j, r_j, s_{j+1}) \sim \mathcal{D}} \left[ \left( r_j + \gamma \max_{a'} \hat{Q}(s_{j+1}, a'; \theta') - Q(s_j, a_j; \theta) \right)^2 \right]

    where γ[0,1]\gamma \in [0, 1] is the discount factor.

    Input: Replay memory capacity NDN_D, discount factor γ\gamma, exploration rate ϵ\epsilon, target network update frequency CC, total training episodes TT
    Initialize replay memory D\mathcal{D} to capacity NDN_D
    Initialize primary Q-network QQ with random weights θ\theta
    Initialize target Q-network Q^\hat{Q} with weights θ=θ\theta' = \theta
    for episode = 1 to TT do
        Receive initial state s1s_1
        for t=1t = 1 to TmaxT_{\text{max}} do
            With probability ϵ\epsilon select a random action atAa_t \in \mathcal{A}
            otherwise select at=argmaxaQ(st,a;θ)a_t = \arg\max_a Q(s_t, a; \theta)
            Execute action ata_t, observe immediate reward rtr_t and next state st+1s_{t+1}
            Store transition tuple (st,at,rt,st+1)(s_t, a_t, r_t, s_{t+1}) in D\mathcal{D}
            Sample random minibatch of transitions (sj,aj,rj,sj+1)(s_j, a_j, r_j, s_{j+1}) from D\mathcal{D}
            Set yj=rj+γmaxaQ^(sj+1,a;θ)y_j = r_j + \gamma \max_{a'} \hat{Q}(s_{j+1}, a'; \theta')
            Perform a gradient descent step on (yjQ(sj,aj;θ))2(y_j - Q(s_j, a_j; \theta))^2 with respect to parameters θ\theta
            Every CC steps, reset Q^=Q\hat{Q} = Q by setting θθ\theta' \leftarrow \theta
            stst+1s_t \leftarrow s_{t+1}
        end for
    end for
    Output: Optimal action-value function Q(s,a;θ)Q^*(s, a; \theta) and policy π(s)=argmaxaQ(s,a;θ)\pi^*(s) = \arg\max_a Q^*(s, a; \theta)
  3. Knowl 3 — Comparative Tradeoffs of Advanced Deep Q-Learning Architectures

    data/table

    Advanced extensions of the Deep Q-Network (DQN) framework mitigate overestimation bias, enhance sample efficiency, handle large action/state spaces, or capture uncertainty. The table compares the core architectural mechanisms, advantages, limitations, and suitable network problem structures:

    Algorithm Key Features Pros Cons Target Applications
    DQN Deep Neural Network as Q-function approximator with replay memory and target net Fast convergence; straightforward implementation Overestimation of action values Small discrete action spaces in MDP settings
    Double DQN (DDQN) Decouples action selection (via θ\theta) from target action evaluation (via θ\theta') Mitigates positive overestimation bias; fast convergence Does not exploit unique problem structure Dynamic spectrum access; resource allocation
    Prioritized DDQN Samples transitions proportional to TD-error priority instead of uniform replay Faster convergence on critical transitions Memory overhead for priority tree management Heterogeneous networks with rare, critical states
    Dueling DDQN Decomposes Q(s,a)Q(s, a) into state-value V(s)V(s) and advantage A(s,a)A(s, a) streams Accelerates learning when actions have similar state values High complexity on small state spaces Problems with large action spaces (e.g., joint user-BS association)
    Asynchronous DQL (e.g., A3C) Parallel multi-agent gradient updates directly to a global network model Highly scalable; lock-free asynchronous CPU training High hardware and communication requirements Ultra-dense networks; handover control
    Distributional DQL Models full return distribution Z(s,a)Z(s, a) instead of scalar expected value Accurate under multimodal or stochastic reward distributions Sensitive to discretization/definition of ZZ Random interference channels with collision tails
    Noisy Nets DQL Injects parametric Gaussian noise directly into dense neural layers Parameterized exploration replacing ϵ\epsilon-greedy heuristics Increased training parameter complexity Large state-action spaces with difficult exploration
    Rainbow DQL Integrates multi-step distributional loss, DDQN, dueling, prioritized replay, and Noisy Nets Combines strengths of all specialized DQL extensions High architectural complexity and hyperparameter tuning Complex cross-layer networking optimizations
  4. Knowl 4 — DRL for Dynamic Spectrum Access and User Association

    model/method

    In dynamic wireless access control, network entities select communication channels and base stations (BSs) under channel state uncertainty:

    1. Dynamic Spectrum Access:

      • Single-Agent POMDP: A sensor agent selects one of MM channels. The state is formed by historical actions and observations (rewards +1+1 for low interference, 1-1 for collision). A DQN with experience replay learns the channel selection policy without requiring prior transition probability knowledge.
      • Multi-User Random Access Game: Multiple users share KK orthogonal channels without carrier sensing or online coordination. Each user's local policy maps past actions and observations to transmission attempt probabilities. Training Double Dueling DQN models offline achieves a subgame perfect equilibrium and doubles aggregate throughput compared to slotted-Aloha.
      • Energy Harvesting Channel Access: A BS schedules transmission channels for energy-harvesting sensors using a two-layer LSTM-DQN. The first layer predicts future sensor battery levels, and the second layer determines the multi-access policy using the predicted battery states and Channel State Information (CSI).
    2. Joint User Association and Spectrum Allocation:

      • In heterogeneous networks (HetNets) with macro and femto BSs, each mobile user acts as an agent selecting a BS and frequency band to maximize data rate subject to a Signal-to-Interference-plus-Noise Ratio (SINR) threshold γth\gamma_{\text{th}}. If SINRγth\text{SINR} \ge \gamma_{\text{th}}, the agent receives a rate-based utility reward; otherwise, it incurs an action selection cost penalty. Double Dueling DQN converges to the optimal joint association with higher capacity than standard tabular Q-learning.
  5. Knowl 5 — DRL for Adaptive Video Streaming and Cellular Rate Control

    model/method

    Deep reinforcement learning optimizes transmission bitrates in dynamic environments without requiring explicit predictive channel models:

    1. Dynamic Adaptive Streaming over HTTP (DASH):

      • Problem Formulation: A video client chooses the bitrate representation RtR_t for chunk tt to maximize user Quality of Experience (QoE).

      • State: Includes last downloaded segment video quality, current playback buffer occupancy, rebuffering duration, and past download channel capacities.

      • Reward: Combines visual quality, rebuffering penalty, and smoothness:

        rt=q(Rt)μTtλq(Rt)q(Rt1)r_t = q(R_t) - \mu T_t - \lambda |q(R_t) - q(R_{t-1})|

        where q(Rt)q(R_t) denotes the perceptual quality of bitrate RtR_t, TtT_t is the rebuffering time, and q(Rt)q(Rt1)|q(R_t) - q(R_{t-1})| penalizes abrupt quality switches.

      • Architectures: Standard LSTM-DQN with peephole connections manages client buffers dynamically and converges in 3\sim 3 episodes (compared to 180\sim 180 episodes for tabular Q-learning). Asynchronous Advantage Actor-Critic (A3C) allows parallel client training, reducing rebuffering by 32.8% and increasing average QoE by up to 25%. A pre-processing CNN-RNN video quality prediction network extracts features directly from raw video chunks, reducing state space dimension and cutting transmission delay by 45%.

    2. High Volume Flexible Time (HVFT) Rate Control:

      • A cellular BS schedules background IoT traffic using an A3C-LSTM framework based on cell congestion load, connection count, and cell efficiency. This doubles transmitted HVFT traffic compared to heuristic schedulers while preventing throughput degradation to foreground mobile users.
  6. Knowl 6 — DRL-Based Wireless Proactive Caching and Transmission Control

    model/method

    Wireless proactive caching reduces backhaul link congestion by pre-fetching popular contents to edge base stations (BSs) or UAVs:

    1. QoS-Aware Content Placement and Expiration:

      • Content Replacement: A BS observes request frequency features over short-, medium-, and long-term windows. To manage large discrete action sets, Deep Deterministic Policy Gradient (DDPG) is combined with the Wolpertinger architecture (using K-Nearest Neighbors to map continuous proto-actions to valid discrete file replacement decisions). This yields significantly higher cache hit rates (0.5 vs 0.4) than First-In First-Out (FIFO) policies with bounded execution complexity.
      • Time-To-Live (TTL) Control: For database query caching, Continuous DQL with Normalized Advantage Functions (NAFs) and Delayed Experience Injection (DEI) adjusts cache expiration duration according to query miss rates and capacity utilization.
    2. Joint Caching and Transmission Control:

      • In multi-user MIMO networks, a central scheduler uses a CNN-based DQN with experience replay to select active user subsets and precoding/caching parameters, avoiding costly real-time CSI feedback and explicit matrix optimization while increasing sum-throughput (e.g., 240 Mbps vs 200 Mbps at SNR=15 dB\text{SNR} = 15\text{ dB}).
      • In Virtual Reality (VR) networks served by cache-enabled UAVs, combining Liquid State Machines (LSM) for spiking temporal memory with Echo State Networks (ESN) as output layers improves VR delivery reliability by 25.4% over standard Q-learning.
  7. Knowl 7 — DRL for Computation and Task Offloading in Mobile Edge Computing

    model/method

    In Mobile Edge Computing (MEC) and Fog Computing, resource-constrained IoT devices offload computational workloads to edge servers, base stations, or neighboring mobile cloudlets:

    1. Cellular and Multi-BS Offloading:

      • State Formulation: Local battery level, pending task queue length, wireless channel state, and computational capacity/queue state at available MEC servers.
      • Cost Objective: Minimization of the long-term expected cost defined as a weighted sum of execution delay, mobile energy consumption, handover delay, and task dropping penalties.
      • Implementations:
        • In cellular-to-WLAN offloading, CNN-DQN reduces mobile device energy consumption by up to 500 Joules compared to dynamic programming with imperfect transition probabilities.
        • In ultra-dense networks (UDNs), Double DQN (DDQN) combined with additive Q-function decomposition and SARSA online updating learns user-BS association and offloading fractions without knowing global network dynamics.
        • For mobile malware detection games, CNN-DQN with hotbooting initialization (transferring Q-values from related scenarios) reduces detection delay by 24.6% to 35.3% over random/standard Q-learning.
    2. Container Migration in Fog Computing:

      • Fog container reallocation is structured as a multi-dimensional MDP solved via DDQN with Prioritized Experience Replay (PER). Grouping nodes into under-, normal-, and over-utilization sets enables power-saving server shutdowns with minimal migration delay.
  8. Knowl 8 — DRL for Anti-Jamming Communications and Physical Layer Security

    model/method

    In non-cooperative wireless environments, radio frequency (RF) jammers cause intentional interference to legitimate transmissions. DRL enables autonomous transceivers to counter dynamic jamming strategies without prior knowledge of channel or jammer models:

    1. Two-Dimensional Frequency Hopping and Mobility Defense:

      • In a Cognitive Radio Network (CRN) under dynamic jamming, a secondary user (SU) agent selects either a frequency channel or moves to an alternate BS area. The state consists of active primary users and discretized signal SINR. A CNN-DQN learns optimal hopping and mobility policies, increasing utility by 8.3% over tabular Q-learning.
      • When raw SINR is corrupted by noise, a Recursive Convolutional Neural Network (RCNN) pre-processing architecture filters spectrum waterfall representations, achieving near-optimal throughput under dynamic jamming where tabular Q-learning fails to converge.
    2. Power Control and Relay-Assisted Anti-Jamming:

      • Power Adaptation: An IoT transmitter learns transmit power levels against an observing jammer using CNN-DQN, improving transmitter utility by 17.7% and reducing jammer utility by 18.1% compared to Q-learning.
      • UAV Relay: When a direct BS link is jammed, a relay UAV uses CNN-DQN to adapt relay transmit power based on received SINR and Bit Error Rate (BER), converging 83.3% faster and reducing BER by 46.6% compared to hill-climbing relay baselines.
      • Physical Layer Secrecy: Transmit UAVs adapt multi-channel power allocation against smart eavesdroppers/jammers to maximize secrecy capacity, yielding a 13% utility gain over Win or Learn Faster-Policy Hill Climbing (WoLF-PHC).
  9. Knowl 9 — DRL for Traffic Engineering and Multi-UAV Connectivity Preservation

    model/method

    Deep Reinforcement Learning enables model-free routing and autonomous formation control in complex dynamic network topologies:

    1. Traffic Engineering and Dynamic Routing:

      • In Software-Defined Networks (SDNs), a centralized controller observes traffic demand matrices between source-destination pairs and determines path split ratios to minimize average network delay.
      • Incorporating Traffic-Engineering-aware exploration (leveraging shortest-path and Network Utility Maximization (NUM) baselines) and Prioritized Experience Replay (PER) into an Actor-Critic architecture significantly reduces packet delay and enhances network utility compared to continuous DDPG and static NUM solvers.
    2. Multi-Robot and UAV Connectivity Preservation:

      • In cooperative multi-UAV networks, global communication graph connectivity is quantified by the algebraic connectivity λ2(L)>0\lambda_2(\mathbf{L}) > 0, where λ2\lambda_2 is the second-smallest eigenvalue of the graph Laplacian matrix L\mathbf{L}.

      • A ground base station agent controls follower velocities using an A3C neural network. The reward function is:

        rt={+1,if Δλ20 (connectivity maintained/improved)1penalty(dmin),if Δλ2<0 or inter-robot collision distance is breachedr_t = \begin{cases} +1, & \text{if } \Delta \lambda_2 \ge 0 \text{ (connectivity maintained/improved)} \\ -1 - \text{penalty}(d_{\text{min}}), & \text{if } \Delta \lambda_2 < 0 \text{ or inter-robot collision distance is breached} \end{cases}

      • For partially observable navigation in complex environments, Recurrent Deterministic Policy Gradient (RDPG) using RNNs guides UAV trajectories to destinations while avoiding obstacles.

  10. Knowl 10 — Fundamental Limitations and Open Challenges of DRL in Wireless Networks

    limitation

    The application of Deep Reinforcement Learning to next-generation wireless communications faces several foundational challenges:

    1. State Determination in Dense Networks: DRL policies typically rely on local state representations such as Received Signal Strength Indicators (RSSIs). In ultra-dense networks (UDNs), RSSI values from adjacent base stations become indistinguishable, making accurate state identification difficult.
    2. Jammer Channel State Unobservability: Security and anti-jamming DRL formulations often require knowledge of the adversary's channel conditions to construct reward functions (e.g., secrecy capacity), which is unfeasible in non-cooperative adversarial settings.
    3. Multi-Agent Non-Stationarity in Dynamic HetNets: In decentralized 5G HetNets, simultaneous adaptation of multiple local DRL agents alters the transition dynamics from each individual agent's perspective, violating the Markov property and causing policy instability or slow convergence.
    4. Scarcity of Real Training Data: Unlike computer vision or natural language processing, wireless networking lacks large-scale, standardized referential measurement datasets. Over-reliance on synthetic stochastic channel models introduces a simulation-to-reality gap where trained policies fail to generalize to actual propagation dynamics.
    5. Information Quality versus Signaling Overhead Tradeoff: Collecting global network state information (e.g., synchronized CSI, queue states, node locations) across distributed nodes improves DRL decision quality but incurs substantial backhaul latency, energy consumption, and signaling overhead.

Coverage note — None was omitted; the knowls fully capture the structural taxonomies, algorithmic models, comparative tables, domain-specific formulations, and open challenges presented in the survey.

References

  1. 1.R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction. MIT press Cambridge, 1998.
  2. 2.I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio, Deep learning. MIT press Cambridge, 2016.
  3. 3.(2016, Jan.) Google achieves ai “breakthrough” by beating go champion. BBC. [Online]. Available: https://www.bbc.com/news/technology-35420579
  4. 4.M. L. Puterman, Markov decision processes: discrete stochastic dy-namic programming. John Wiley & Sons, 2014.
  5. 5.D. P. Bertsekas, D. P. Bertsekas, D. P. Bertsekas, and D. P. Bertsekas, Dynamic programming and optimal control. Athena scientific Bel-mont, MA, 2005, vol. 1, no. 3.
  6. 6.R. Bellman, Dynamic programming. Mineola, NY: Courier Corpora-tion, 2013.
  7. 7.Y. Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274, 2017.
  8. 8.K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “A brief survey of deep reinforcement learning,” IEEE Signal Process-ing Magazine, vol. 34, no. 6, pp. 26–38, Nov. 2017.
  9. 9.Y. Xin, L. Kong, Z. Liu, Y. Chen, Y. Li, H. Zhu, M. Gao, H. Hou, and C. Wang, “Machine learning and deep learning methods for cybersecurity,” IEEE Access, to appear.
  10. 10.Z. Fadlullah, F. Tang, B. Mao, N. Kato, O. Akashi, T. Inoue, and K. Mizutani, “State-of-the-art deep learning: Evolving machine intelli-gence toward tomorrow’s intelligent network traffic control systems,” IEEE Communications Surveys & Tutorials, vol. 19, no. 4, pp. 2432–2455, 2017.
  11. 11.Q. Mao, F. Hu, and Q. Hao, “Deep learning for intelligent wireless networks: A comprehensive survey,” IEEE Communications Surveys & Tutorials, vol. 20, no. 4, pp. 2595–2621, 2018.
  12. 12.M. Chen, U. Challita, W. Saad, C. Yin, and M. Debbah, “Machine learning for wireless networks with artificial intelligence: A tutorial on neural networks,” arXiv preprint arXiv:1710.02913, 2017.
  13. 13.W. Wang, A. Kwasinski, D. Niyato, and Z. Han, “A survey on applica-tions of model-free strategy learning in cognitive wireless networks,” IEEE Communications Surveys & Tutorials, vol. 18, no. 3, pp. 1717–1757, 2016.
  14. 14.R. Lowe, Y. Wu, A. Tamar, J. Harb, O. P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environ-ments,” in Advances in Neural Information Processing Systems, 2017, pp. 6379–6390.
  15. 15.G. E. Monahan, “State of the art-a survey of partially observable markov decision processes: theory, models, and algorithms,” Manage-ment Science, vol. 28, no. 1, pp. 1–16, 1982.
  16. 16.L. S. Shapley, “Stochastic games,” Proceedings of the national academy of sciences, vol. 39, no. 10, pp. 1095–1100, 1953.
  17. 17.J. Hu and M. P. Wellman, “Nash q-learning for general-sum stochastic games,” Journal of machine learning research, vol. 4, no. Nov, pp. 1039–1069, 2003.
  18. 18.W. C. Dabney, “Adaptive step-sizes for reinforcement learning,” Ph.D. dissertation, 2014.
  19. 19.C. J. Watkins and P. Dayan, “Q-learning,” Machine learning, vol. 8, no. 3-4, pp. 279–292, 1992.
  20. 20.J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venu-gopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 2625–2634.
  21. 21.V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, p. 529, 2015.
  22. 22.Y. Lin, X. Dai, L. Li, and F.-Y. Wang, “An efficient deep rein-forcement learning model for urban traffic control,” arXiv preprint arXiv:1808.01876, 2018.
  23. 23.S. Gu, E. Holly, T. Lillicrap, and S. Levine, “Deep reinforcement learn-ing for robotic manipulation with asynchronous off-policy updates,” in IEEE International Conference on Robotics and Automation (ICRA), 2017, pp. 3389–3396.
  24. 24.S. Thrun and A. Schwartz, “Issues in using function approximation for reinforcement learning,” in Proceedings of Connectionist Models Summer School Hillsdale, NJ. Lawrence Erlbaum, 1993.
  25. 25.H. V. Hasselt, “Double q-learning,” in Advances in Neural Information Processing Systems, 2010, pp. 2613–2621.
  26. 26.H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning.” in AAAI, vol. 2, Phoenix, AZ, Feb. 2016, pp. 2094–2100.
  27. 27.O. Naparstek and K. Cohen, “Deep multi-user reinforcement learning for dynamic spectrum access in multichannel wireless networks,” arXiv preprint arXiv:1704.02613, 2017.
  28. 28.N. Zhao, Y.-C. Liang, D. Niyato, Y. Pei, M. Wu, and Y. Jiang, “Deep reinforcement learning for user association and resource allocation in heterogeneous networks,” in IEEE GLOBECOM, Abu Dhabi, UAE, Dec. 2018, pp. 1–6.
  29. 29.T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” arXiv preprint arXiv:1511.05952, 2015.
  30. 30.Z. Wang, T. Schaul, M. Hessel, H. Van Hasselt, M. Lanctot, and N. De Freitas, “Dueling network architectures for deep reinforcement learning,” in International Conference on Machine Learning, New York, NY, Jun. 2016.
  31. 31.V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep rein-forcement learning,” in International Conference on Machine Learning, New York City, New York, Jun. 2016, pp. 1928–1937.
  32. 32.Z. Wang, Y. Xu, L. Li, H. Tian, and S. Cui, “Handover control in wireless systems via asynchronous multi-user deep reinforcement learning,” arXiv preprint arXiv:1801.02077, 2018.
  33. 33.M. G. Bellemare, W. Dabney, and R. Munos, “A distributional per-spective on reinforcement learning,” arXiv preprint arXiv:1707.06887, 2017.
  34. 34.M. Fortunato, M. G. Azar, B. Piot, J. Menick, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin et al., “Noisy networks for exploration,” in International Conference on Learning Representations, 2018.
  35. 35.M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver, “Rainbow: Combining improvements in deep reinforcement learning,” in The Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
  36. 36.T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforce-ment learning,” San Juan, Puerto Rico, USA, May 2016.
  37. 37.D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Ried-miller, “Deterministic policy gradient algorithms,” in ICML, 2014.
  38. 38.M. Hausknecht and P. Stone, “Deep recurrent q-learning for partially observable mdps,” CoRR, abs/1507.06527, vol. 7, no. 1, 2015.
  39. 39.D. Zhao, H. Wang, K. Shao, and Y. Zhu, “Deep reinforcement learning with experience replay based on sarsa,” in IEEE Symposium Series on Computational Intelligence (SSCI), 2016, pp. 1–6.
  40. 40.W. Wang, J. Hao, Y. Wang, and M. Taylor, “Towards cooperation in sequential prisoner’s dilemmas: a deep multiagent reinforcement learning approach,” arXiv preprint arXiv:1803.00162, 2018.
  41. 41.J. Heinrich and D. Silver, “Deep reinforcement learning from self-play in imperfect-information games,” arXiv preprint arXiv:1603.01121, 2016.
  42. 42.D. Fooladivanda and C. Rosenberg, “Joint resource allocation and user association for heterogeneous wireless cellular networks,” IEEE Transactions on Wireless Communications, vol. 12, no. 1, pp. 248–257, 2013.
  43. 43.Y. Lin, W. Bao, W. Yu, and B. Liang, “Optimizing user association and spectrum allocation in hetnets: A utility perspective,” IEEE Journal on Selected Areas in Communications, vol. 33, no. 6, pp. 1025–1039, 2015.
  44. 44.S. Wang, H. Liu, P. H. Gomes, and B. Krishnamachari, “Deep rein-forcement learning for dynamic multichannel access,” in International Conference on Computing, Networking and Communications (ICNC), 2017.
  45. 45.Q. Zhao, B. Krishnamachari, and K. Liu, “On myopic sensing for multi-channel opportunistic access: structure, optimality, and performance,” IEEE Transactions on Wireless Communications, vol. 7, no. 12, pp. 5431–5440, December 2008.
  46. 46.V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller, “Playing atari with deep reinforce-ment learning,” CoRR, vol. abs/1312.5602, 2013.
  47. 47.R. Govindan. Tutornet: A low power wireless iot testbed. [Online]. Available: http://anrg.usc.edu/www/tutornet/
  48. 48.S. Wang, H. Liu, P. H. Gomes, and B. Krishnamachari, “Deep reinforcement learning for dynamic multichannel access in wireless networks,” IEEE Transactions on Cognitive Communications and Net-working, to appear.
  49. 49.J. Zhu, Y. Song, D. Jiang, and H. Song, “A new deep-q-learning-based transmission scheduling mechanism for the cognitive internet of things,” IEEE Internet of Things Journal, 2017.
  50. 50.J. Fearnley, “Strategy iteration algorithms for games and markov decision processes,” Ph.D. dissertation, University of Warwick, 2010.
  51. 51.M. Chu, H. Li, X. Liao, and S. Cui, “Reinforcement learning based multi-access control and battery prediction with energy harvesting in iot systems,” arXiv preprint arXiv:1805.05929, 2018.
  52. 52.X. Di, K. Xiong, P. Fan, H.-C. Yang, and K. B. Letaief, “Optimal resource allocation in wireless powered communication networks with user cooperation,” IEEE Transactions on Wireless Communications, vol. 16, no. 12, pp. 7936–7949, 2017.
  53. 53.H. Ye and G. Y. Li, “Deep reinforcement learning for resource allocation in v2v communications,” arXiv preprint arXiv:1711.00968, 2017.
  54. 54.U. Challita, L. Dong, and W. Saad, “Proactive resource manage-ment in lte-u systems: A deep learning perspective,” arXiv preprint arXiv:1702.07031, 2017.
  55. 55.M. Balazinska and P. Castro. (2003) Ibm watson research center. [Online]. Available: https://crawdad.org/ibm/watson/20030219
  56. 56.Z. Wang, T. Schaul, M. Hessel, H. Van Hasselt, M. Lanctot, and N. De Freitas, “Dueling network architectures for deep reinforcement learning,” 2015.
  57. 57.H. Li, “Multiagent learning for aloha-like spectrum access in cognitive radio systems,” EURASIP Journal on Wireless Communications and Networking, vol. 2010, no. 1, pp. 1–15, May 2010.
  58. 58.S. Liu, X. Hu, and W. Wang, “Deep reinforcement learning based dynamic channel allocation algorithm in multibeam satellite systems,” IEEE ACCESS, vol. 6, pp. 15 733–15 742, 2018.
  59. 59.A. R. Elsherif, W.-P. Chen, A. Ito, and Z. Ding, “Resource allocation and inter-cell interference management for dual-access small cells,” IEEE Journal on Selected Areas in Communications, vol. 33, no. 6, pp. 1082–1096, 2015.
  60. 60.M. Chen, W. Saad, and C. Yin, “Liquid state machine learning for resource allocation in a network of cache-enabled lte-u uavs,” in IEEE GLOBECOM, 2017, pp. 1–6.
  61. 61.W. Maass, “Liquid state machines: motivation, theory, and applica-tions,” in Computability in context: computation and logic in the real world. World Scientific, 2011, pp. 275–296.
  62. 62.M. Chen, M. Mozaffari, W. Saad, C. Yin, M. Debbah, and C. S. Hong, “Caching in the sky: Proactive deployment of cache-enabled unmanned aerial vehicles for optimized quality-of-experience,” IEEE Journal on Selected Areas in Communications, vol. 35, no. 5, pp. 1046–1061, 2017.
  63. 63.I. Szita, V. Gyenes, and A. Lorincz, “Reinforcement learning with Ȅ echo state networks,” in International Conference on Artificial Neural Networks. Springer, 2006, pp. 830–839.
  64. 64.Tyouku of china network video index. [Online]. Available: http://index.youku.com/
  65. 65.T. Stockhammer, “Dynamic adaptive streaming over http–: standards and design principles,” in Proceedings of the second annual ACM conference on Multimedia systems. ACM, 2011, pp. 133–144.
  66. 66.M. Gadaleta, F. Chiariotti, M. Rossi, and A. Zanella, “D-dash: A deep q-learning framework for dash video streaming,” IEEE Transactions on Cognitive Communications and Networking, vol. 3, no. 4, pp. 703–718, 2017.
  67. 67.J. Klaue, B. Rathke, and A. Wolisz, “Evalvid–a framework for video transmission and quality evaluation,” in International conference on modelling techniques and tools for computer performance evaluation, 2003, pp. 255–272.
  68. 68.H. Mao, R. Netravali, and M. Alizadeh, “Neural adaptive video streaming with pensieve,” in Proceedings of the Conference of the ACM Special Interest Group on Data Communication. ACM, 2017, pp. 197–210.
  69. 69.H. Riiser, P. Vigmostad, C. Griwodz, and P. Halvorsen, “Commute path bandwidth traces from 3g networks: analysis and applications,” in Proceedings of the 4th ACM Multimedia Systems Conference. ACM, 2013, pp. 114–118.
  70. 70.X. Yin, A. Jindal, V. Sekar, and B. Sinopoli, “A control-theoretic approach for dynamic adaptive video streaming over http,” in ACM SIGCOMM Computer Communication Review, vol. 45, no. 4. ACM, 2015, pp. 325–338.
  71. 71.T. Huang, R.-X. Zhang, C. Zhou, and L. Sun, “Qarc: Video quality aware rate control for real-time video streaming based on deep rein-forcement learning,” arXiv preprint arXiv:1805.02482, 2018.
  72. 72.(2016) Measuring fixed broadband report. [Online]. Available: https://www.fcc.gov/reports-research/reports/measuring-broadband-america/raw-data-measuring-broadband-america-2016
  73. 73.S. Chinchali, P. Hu, T. Chu, M. Sharma, M. Bansal, R. Misra, M. Pavone, and K. Sachin, “Cellular network traffic scheduling with deep reinforcement learning,” in National Conference on Artificial Intelligence (AAAI), 2018.
  74. 74.Z. Zhang, Y. Zheng, M. Hua, Y. Huang, and L. Yang, “Cache-enabled dynamic rate allocation via deep self-transfer reinforcement learning,” arXiv preprint arXiv:1803.11334, 2018.
  75. 75.P. V. R. Ferreira, R. Paffenroth, A. M. Wyglinski, T. M. Hackett, S. G. Bilén, R. C. Reinhart, and D. J. Mortensen, “Multi-objective reinforcement learning for cognitive satellite communications using deep neural network ensembles,” IEEE Journal on Selected Areas in Communications, 2018.
  76. 76.D. Tarchi, G. E. Corazza, and A. Vanelli-Coralli, “Adaptive coding and modulation techniques for next generation hand-held mobile satellite communications,” in IEE ICC, 2013, pp. 4504–4508.
  77. 77.M. T. Hagan and M. B. Menhaj, “Training feedforward networks with the marquardt algorithm,” IEEE transactions on Neural Networks, vol. 5, no. 6, pp. 989–993, 1994.
  78. 78.C. Zhong, M. C. Gursoy, and S. Velipasalar, “A deep reinforce-ment learning-based framework for content caching,” arXiv preprint arXiv:1712.08132, 2017.
  79. 79.T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” CoRR, vol. abs/1509.02971, 2015. [Online]. Available: http://arxiv.org/abs/1509.02971
  80. 80.G. Dulac-Arnold, R. Evans, P. Sunehag, and B. Coppin, “Reinforcement learning in large discrete action spaces,” CoRR, vol. abs/1512.07679, 2015. [Online]. Available: http://arxiv.org/abs/1512.07679
  81. 81.L. Lei, L. You, G. Dai, T. X. Vu, D. Yuan, and S. Chatzinotas, “A deep learning approach for optimizing content delivering in cache-enabled HetNet,” in Int’l Sym. Wireless Commun. Systems (ISWCS), Aug. 2017, pp. 449–453.
  82. 82.M. Schaarschmidt, F. Gessert, V. Dalibard, and E. Yoneki, “Learning runtime parameters in computer systems with delayed experience injection,” arXiv preprint arXiv:1610.09903, 2016.
  83. 83.B. F. Cooper, A. Silberstein, E. Tam, R. Ramakrishnan, and R. Sears, “Benchmarking cloud serving systems with YCSB,” in proc. 1st ACM Sym. Cloud Comput., 2010, pp. 143–154.
  84. 84.M. Deghel, E. Bastug, M. Assaad, and M. Debbah, “On the benefits of edge caching for mimo interference alignment,” in IEEE International Workshop on Signal Processing Advances in Wireless Communications (SPAWC), 2015, pp. 655–659.
  85. 85.Y. He and S. Hu, “Cache-enabled wireless networks with opportunistic interference alignment,” arXiv preprint arXiv:1706.09024, 2017.
  86. 86.Y. He, C. Liang, F. R. Yu, N. Zhao, and H. Yin, “Optimization of cache-enabled opportunistic interference alignment wireless networks: A big data deep reinforcement learning approach,” in IEEE ICC, 2017, pp. 1–6.
  87. 87.Y. He, Z. Zhang, F. R. Yu, N. Zhao, H. Yin, V. C. Leung, and Y. Zhang, “Deep-reinforcement-learning-based optimization for cache-enabled opportunistic interference alignment wireless networks,” IEEE Transactions on Vehicular Technology, vol. 66, no. 11, pp. 10 433–10 445, 2017.
  88. 88.X. He, K. Wang, H. Huang, T. Miyazaki, Y. Wang, and S. Guo, “Green resource allocation based on deep reinforcement learning in content-centric iot,” IEEE Transactions on Emerging Topics in Computing, to appear.
  89. 89.Q. Wu, Z. Li, and G. Xie, “CodingCache: Multipath-aware CCN cache with network coding,” in proc. ACM SIGCOMM Workshop on Information-centric Networking, 2013, pp. 41–42.
  90. 90.M. Chen, W. Saad, and C. Yin, “Echo-liquid state deep learning for 360 content transmission and caching in wireless vr networks with cellular-connected uavs,” arXiv preprint arXiv:1804.03284, 2018.
  91. 91.M. Chen, U. Challita, W. Saad, C. Yin, and M. Debbah, “Machine learning for wireless networks with artificial intelligence: A tutorial on neural networks,” CoRR, vol. abs/1710.02913, 2017. [Online]. Available: http://arxiv.org/abs/1710.02913
  92. 92.M. Chen, W. Saad, and C. Yin, “Liquid state machine learning for resource allocation in a network of cache-enabled LTE-U UAVs,” in IEEE GLOBECOM, Dec. 2017.
  93. 93.Y. He, Z. Zhang, and Y. Zhang, “A big data deep reinforcement learning approach to next generation green wireless networks,” in IEEE GLOBECOM, 2017, pp. 1–6.
  94. 94.Y. He, C. Liang, Z. Zhang, F. R. Yu, N. Zhao, H. Yin, and Y. Zhang, “Resource allocation in software-defined and information-centric vehic-ular networks with mobile edge computing,” in Vehicular Technology Conference (VTC-Fall), 2017 IEEE 86th, 2017, pp. 1–5.
  95. 95.Y. He, F. R. Yu, N. Zhao, H. Yin, and A. Boukerche, “Deep reinforce-ment learning (drl)-based resource management in software-defined and virtualized vehicular ad hoc networks,” in Proceedings of the 6th ACM Symposium on Development and Analysis of Intelligent Vehicular Networks and Applications, 2017, pp. 47–54.
  96. 96.Y. He, N. Zhao, and H. Yin, “Integrated networking, caching, and computing for connected vehicles: A deep reinforcement learning approach,” IEEE Transactions on Vehicular Technology, vol. 67, no. 1, pp. 44–55, 2018.
  97. 97.T. L. Thanh and R. Q. Hu, “Mobility-aware edge caching and comput-ing framework in vehicle networks: A deep reinforcement learning,” IEEE Transactions on Vehicular Technology, Jan. 2018.
  98. 98.Y. He, F. R. Yu, N. Zhao, V. C. Leung, and H. Yin, “Software-defined networks with mobile edge computing and caching for smart cities: A big data deep reinforcement learning approach,” IEEE Communications Magazine, vol. 55, no. 12, pp. 31–37, 2017.
  99. 99.B. Han, V. Gopalakrishnan, L. Ji, and S. Lee, “Network function virtualization: Challenges and opportunities for innovations,” IEEE Communications Magazine, vol. 53, no. 2, pp. 90–97, 2015.
  100. 100.Y. He, F. R. Yu, N. Zhao, and H. Yin, “Secure social networks in 5g systems with mobile edge computing, caching and device-to-device (d2d) communications,” IEEE Wireless Communications, vol. 25, no. 3, pp. 103–109, Jun. 2018.
  101. 101.Y. He, C. Liang, F. R. Yu, and Z. Han, “Trust-based social networks with computing, caching and communications: A deep reinforcement learning approach,” IEEE Transactions on Network Science and Engi-neering, to appear.
  102. 102.C. Zhang, B. Gu, Z. Liu, K. Yamori, and Y. Tanaka, “Cost-and energy-aware multi-flow mobile data offloading using markov decision process,” IEICE Transactions on Communications, 2017.
  103. 103.C. Zhang, Z. Liu, B. Gu, K. Yamori, and Y. Tanaka, “A deep rein-forcement learning based approach for cost-and energy-aware multi-flow mobile data offloading,” IEICE Transactions on Communications, pp. 2017–2025.
  104. 104.L. Ji, G. Hui, L. Tiejun, and L. Yueming, “Deep reinforcement learning based computation offloading and resource allocation for mec,” in IEEE WCNC, 2018, pp. 1–5.
  105. 105.X. Chen, H. Zhang, C. Wu, S. Mao, Y. Ji, and M. Bennis, “Perfor-mance optimization in mobile-edge computing via deep reinforcement learning,” arXiv preprint arXiv:1804.00514, 2018.
  106. 106.X. Chen, H. Zhang, C. Wu, S. Mao, Y. Ji, and M. Bennis, “Optimized computation offloading performance in virtual edge computing systems via deep reinforcement learning,” arXiv preprint arXiv:1805.06146, 2018.
  107. 107.J. Ye and Y.-J. A. Zhang, “DRAG: Deep reinforcement learning based base station activation in heterogeneous networks,” arXiv:1809.02159, Sep. 2018.
  108. 108.H. Li, H. Gao, T. Lv, and Y. Lu, “Deep q-learning based dynamic resource allocation for self-powered ultra-dense networks,” in IEEE ICC (ICC Workshops), 2018, pp. 1–6.
  109. 109.J. Liu, B. Krishnamachari, S. Zhou, and Z. Niu, “Deepnap: Data-driven base station sleeping operations through deep reinforcement learning,” IEEE Internet of Things Journal, 2018.
  110. 110.X. Wan, G. Sheng, Y. Li, L. Xiao, and X. Du, “Reinforcement learning based mobile offloading for cloud-based malware detection,” in IEEE GLOBECOM, 2017, pp. 1–6.
  111. 111.L. Xiao, X. Wan, C. Dai, X. Du, X. Chen, and M. Guizani, “Security in mobile edge caching with reinforcement learning,” CoRR, vol. abs/1801.05915, 2018. [Online]. Available: http://arxiv.org/abs/1801.05915
  112. 112.A. S. Shamili, C. Bauckhage, and T. Alpcan, “Malware detection on mobile devices using distributed machine learning,” in proc. Int’l Conf. Pattern Recognition, Aug. 2010, pp. 4348–4351.
  113. 113.Y. Li, J. Liu, Q. Li, and L. Xiao, “Mobile cloud offloading for malware detections with learning,” in IEEE INFOCOM Workshops, Apr. 2015, pp. 197–201.
  114. 114.M. Min, D. Xu, L. Xiao, Y. Tang, and D. Wu, “Learning-based computation offloading for iot devices with energy harvesting,” arXiv preprint arXiv:1712.08768, 2017.
  115. 115.L. Quan, Z. Wang, and F. Ren, “A novel two-layered reinforcement learning for task offloading with tradeoff between physical machine utilization rate and delay,” Future Internet, vol. 10, no. 7, 2018.
  116. 116.D. V. Le and C. Tham, “Quality of service aware computation of-floading in an ad-hoc mobile cloud,” IEEE Transactions on Vehicular Technology, pp. 1–1, 2018.
  117. 117.D. V. Le and C.-K. Tham, “A deep reinforcement learning based offloading scheme in ad-hoc mobile clouds,” in Proceedings of IEEE INFOCOM IECCO Workshop, Honolulu, USA., apr 2018.
  118. 118.S. Yu, X. Wang, and R. Langar, “Computation offloading for mobile edge computing: A deep learning approach,” in IEEE PIMRC, Oct. 2017.
  119. 119.Z. Tang, X. Zhou, F. Zhang, W. Jia, and W. Zhao, “Migration modeling and learning algorithms for containers in fog computing,” IEEE Transactions on Services Computing, 2018.
  120. 120.P. Popovski, H. Yomo, and R. Prasad, “Strategies for adaptive fre-quency hopping in the unlicensed bands,” IEEE Wireless Communica-tions, vol. 13, no. 6, pp. 60–67, 2006.
  121. 121.G. Han, L. Xiao, and H. V. Poor, “Two-dimensional anti-jamming communication based on deep reinforcement learning,” in Proceedings of the 42nd IEEE International Conference on Acoustics, Speech and Signal Processing,, 2017.
  122. 122.L. Xiao, D. Jiang, X. Wan, W. Su, and Y. Tang, “Anti-jamming under-water transmission with mobility and learning,” IEEE Communications Letters, 2018.
  123. 123.X. Liu, Y. Xu, L. Jia, Q. Wu, and A. Anpalagan, “Anti-jamming com-munications using spectrum waterfall: A deep reinforcement learning approach,” IEEE Communications Letters, 2018.
  124. 124.Y. Chen, Y. Li, D. Xu, and L. Xiao, “Dqn-based power control for iot transmission against jamming,” in IEEE 87th Vehicular Technology Conference (VTC Spring), 2018, pp. 1–5.
  125. 125.X. Lu, L. Xiao, and C. Dai, “Uav-aided 5g communications with deep reinforcement learning against jamming,” arXiv preprint arXiv:1805.06628, 2018.
  126. 126.L. Xiao, X. Lu, D. Xu, Y. Tang, L. Wang, and W. Zhuang, “Uav relay in vanets against smart jamming with reinforcement learning,” IEEE Transactions on Vehicular Technology, vol. 67, no. 5, pp. 4087–4097, 2018.
  127. 127.S. Lv, L. Xiao, Q. Hu, X. Wang, C. Hu, and L. Sun, “Anti-jamming power control game in unmanned aerial vehicle networks,” in IEEE GLOBECOM, 2017, pp. 1–6.
  128. 128.L. Xiao, C. Xie, M. Min, and W. Zhuang, “User-centric view of unmanned aerial vehicle transmission against smart attacks,” IEEE Transactions on Vehicular Technology, vol. 67, no. 4, pp. 3420–3430, 2018.
  129. 129.M. Bowling and M. Veloso, “Multiagent learning using a variable learning rate,” Artificial Intelligence, vol. 136, no. 2, pp. 215–250, 2002.
  130. 130.Y. Chen, S. Kar, and J. M. Moura, “Cyber-physical attacks with control objectives,” IEEE Transactions on Automatic Control, vol. 63, no. 5, pp. 1418–1425, 2018.
  131. 131.A. Ferdowsi, U. Challita, W. Saad, and N. B. Mandayam, “Robust deep reinforcement learning for security and safety in autonomous vehicle systems,” arXiv preprint arXiv:1805.00983, 2018.
  132. 132.M. Brackstone and M. McDonald, “Car-following: a historical review,” Transportation Research Part F: Traffic Psychology and Behaviour, vol. 2, no. 4, pp. 181–196, 1999.
  133. 133.A. Ferdowsi and W. Saad, “Deep learning-based dynamic watermarking for secure signal authentication in the internet of things,” in IEEE ICC, 2018, pp. 1–6.
  134. 134.B. Satchidanandan and P. R. Kumar, “Dynamic watermarking: Active defense of networked cyber–physical systems,” Proceedings of the IEEE, vol. 105, no. 2, pp. 219–240, 2017.
  135. 135.A. Ferdowsi and W. Saad, “Deep learning for signal authentication and security in massive internet of things systems,” arXiv preprint arXiv:1803.00916, 2018.
  136. 136.P. Vadakkepat, K. C. Tan, and W. Ming-Liang, “Evolutionary artificial potential fields and their application in real time robot path planning,” in Proceedings of the 2000 Congress on Evolutionary Computation, vol. 1, 2000, pp. 256–263.
  137. 137.W. Huang, Y. Wang, and X. Yi, “Deep q-learning to preserve connec-tivity in multi-robot systems,” in Proceedings of the 9th International Conference on Signal Processing Systems. ACM, 2017, pp. 45–50.
  138. 138.W. Huang, Y. Wang, and X. Yi, “A deep reinforcement learning approach to preserve connectivity for multi-robot systems,” in In-ternational Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), 2017, pp. 1–7.
  139. 139.H. A. Poonawala, A. C. Satici, H. Eckert, and M. W. Spong, “Collision-free formation control with decentralized connectivity preservation for nonholonomic-wheeled mobile robots,” IEEE Transactions on control of Network Systems, vol. 2, no. 2, pp. 122–130, 2015.
  140. 140.C. Wang, J. Wang, X. Zhang, and X. Zhang, “Autonomous naviga-tion of uav in large-scale unknown complex environment with deep reinforcement learning,” in IEEE GlobalSIP, 2017, pp. 858–862.
  141. 141.C. Shen, C. Tekin, and M. van der Schaar, “A non-stochastic learning approach to energy efficient mobility management,” IEEE Journal on Selected Areas in Communications, vol. 34, no. 12, pp. 3854–3868, 2016.
  142. 142.M. Faris and E. Brian, “Deep q-learning for self-organizing net-works fault management and radio performance improvement,” https://arxiv.org/abs/1707.02329, 2018.
  143. 143.G. Stampa, M. Arias, D. Sanchez-Charles, V. Muntes-Mulero, and A. Cabellos, “A deep-reinforcement learning approach for software-defined networking routing optimization,” arXiv preprint arXiv:1709.07080, 2017.
  144. 144.M. Roughan, “Simplifying the synthesis of internet traffic matrices,” ACM SIGCOMM Computer Communication Review, vol. 35, no. 5, pp. 93–96, 2015. [Online]. Available: http://arxiv.org/abs/1710.02913
  145. 145.A. Varga and R. Hornig, “An overview of the OMNeT++ simulation environment,” in proc. Int’l Conf. Simulation Tools and Techniques for Communications, Networks and Systems & Workshops, 2008.
  146. 146.Z. Xu, J. Tang, J. Meng, W. Zhang, Y. Wang, C. H. Liu, and D. Yang, “Experience-driven networking: A deep reinforcement learning based approach,” arXiv preprint arXiv:1801.05757, 2018.
  147. 147.K. Winstein and H. Balakrishnan, “TCP ex Machina: Computer-generated congestion control,” in ACM SIGCOMM, 2013, pp. 123–134.
  148. 148.R. G.F. and H. T.R., Modeling and Tools for Network Simulation. Springer, Berlin, Heidelberg, 2010, ch. The ns-3 Network Simulator.
  149. 149.A. Medina, A. Lakhina, I. Matta, and J. Byers, “BRITE: an approach to universal topology generation,” in IEEE MASCOTS, Aug. 2001, pp. 346–353.
  150. 150.R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in proc. 12th Int’l Conf. Neural Inform. Process. Syst., 1999, pp. 1057–1063.
  151. 151.U. Challita, W. Saad, and C. Bettstetter, “Deep reinforcement learning for interference-aware path planning of cellular connected uavs,” in IEEE ICC, Kansas City, MO, May 2018, pp. 1–6.
  152. 152.U. Challita, W. Saad, and C. Bettstetter, “Cellular-connected uavs over 5g: Deep reinforcement learning for interference management,” arXiv preprint arXiv:1801.05500, 2018.
  153. 153.L. Zhu, Y. He, F. R. Yu, B. Ning, T. Tang, and N. Zhao, “Communication-based train control system performance optimization using deep reinforcement learning,” IEEE Transactions on Vehicular Technology, vol. 66, no. 12, pp. 10 705–10 717, 2017.
  154. 154.Y. Yang, Y. Li, K. Li, S. Zhao, R. Chen, J. Wang, and S. Ci, “Decco: Deep-learning enabled coverage and capacity optimization for massive mimo systems,” IEEE Access, to appear.
  155. 155.Z. Xu, Y. Wang, J. Tang, J. Wang, and M. C. Gursoy, “A deep reinforcement learning based framework for power-efficient resource allocation in cloud rans,” in IEEE ICC, 2017, pp. 1–6.
  156. 156.T. M. Hackett, S. G. Bilén, P. V. R. Ferreira, A. M. Wyglinski, and R. C. Reinhart, “Implementation of a space communications cognitive engine,” in Cognitive Communications for Aerospace Applications Workshop (CCAA), 2017, pp. 1–7.
  157. 157.K. Shen and W. Yu, “Fractional programming for communication systemsÑpart i: Power control and beamforming,” IEEE Transactions on Signal Processing, vol. 66, no. 10, pp. 2616–2630, 2018.
  158. 158.Y. S. Nasir and D. Guo, “Deep reinforcement learning for dis-tributed dynamic power allocation in wireless networks,” arXiv preprint arXiv:1808.00490, 2018.
  159. 159.X. Foukas, G. Patounas, A. Elmokashfi, and M. K. Marina, “Network slicing in 5g: Survey and challenges,” IEEE Communications Maga-zine, vol. 55, no. 5, pp. 94–100, 2017.
  160. 160.X. Chen, Z. Li, Y. Zhang, R. Long, H. Yu, X. Du, and M. Guizani, “Re-inforcement learning based qos/qoe-aware service function chaining in software-driven 5g slices,” arXiv preprint arXiv:1804.02099, 2018.
  161. 161.P. Reichl, S. Egger, R. Schatz, and A. D’Alconzo, “The logarithmic nature of qoe and the role of the weber-fechner law in qoe assessment,” in IEEE ICC, Cape Town, South Africa, May 2010, pp. 1–5.
  162. 162.M. Fiedler, T. Hossfeld, and P. Tran-Gia, “A generic quantitative relationship between quality of experience and quality of service,” IEEE Network, vol. 24, no. 2, pp. 36–41, 2014.
  163. 163.Z. Zhao, R. Li, Q. Sun, Y. Yang, X. Chen, M. Zhao, H. Zhang et al., “Deep reinforcement learning for network slicing,” arXiv preprint arXiv:1805.06591, 2018.
  164. 164.Y. Zhou, Z. M. Fadlullah, B. Mao, and N. Kato, “A deep-learning-based radio resource assignment technique for 5g ultra dense networks,” IEEE Network, vol. 32, no. 6, pp. 28–34, Nov. 2018.
  165. 165.H. Mao, M. Alizadeh, I. Menache, and S. Kandula, “Resource man-agement with deep reinforcement learning,” in Proceedings of the 15th ACM Workshop on Hot Topics in Networks, 2016, pp. 50–56.
  166. 166.T. Li, Z. Xu, J. Tang, and Y. Wang, “Model-free control for distributed stream data processing using deep reinforcement learning,” Proceed-ings of the VLDB Endowment, vol. 11, no. 6, pp. 705–718, 2018.
  167. 167.G. D. Arnold, R. Evans, H. v. Hasselt, P. Sunehag, T. Lillicrap, J. Hunt, T. Mann, T. Weber, T. Degris, and B. Coppin, “Deep reinforcement learning in large discrete action spaces,” arXiv: 1512.07679, 2016.
  168. 168.X. Li, J. Fang, W. Cheng, H. Duan, Z. Chen, and H. Li, “Intelligent power control for spectrum sharing in cognitive radios: A deep rein-forcement learning approach,” arXiv preprint arXiv:1712.07365, 2017.
  169. 169.A. Zappone, M. Debbah, and Z. Altman, “Online energy-efficient power control in wireless networks by deep neural networks,” in IEEE 19th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC), 2018, pp. 1–5.
  170. 170.A. Zappone, E. Björnson, L. Sanguinetti, and E. Jorswieck, “Globally optimal energy-efficient power control and receiver design in wireless networks,” IEEE Transactions on Signal Processing, vol. 65, no. 11, pp. 2844–2859, 2017.
  171. 171.L. Sanguinetti, A. Zappone, and M. Debbah, “Deep learning power allocation in massive mimo,” arXiv preprint arXiv:1812.03640, 2018.
  172. 172.S. Boyd, V. Balakrishnan, and P. Kabamba, “A bisection method for computing the h◦ norm of a transfer matrix and related problems,” Mathematics of Control, Signals and Systems, vol. 2, no. 3, pp. 207–219, 1989.
  173. 173.E. Björnson, J. Hoydis, and L. Sanguinetti, “Massive mimo has unlimited capacity,” IEEE Transactions on Wireless Communications, vol. 17, no. 1, pp. 574–590, 2018.
  174. 174.T. Oda, R. Obukata, M. Ikeda, L. Barolli, and M. Takizawa, “Design and implementation of a simulation system based on deep q-network for mobile actor node control in wireless sensor and actor networks,” in International Conference on Advanced Information Networking and Applications Workshops (WAINA), 2017, pp. 195–200.
  175. 175.L. Wang, W. Liu, D. Zhang, Y. Wang, E. Wang, and Y. Yang, “Cell selection with deep reinforcement learning in sparse mobile crowdsensing,” arXiv preprint arXiv:1804.07047, 2018.
  176. 176.L. Wang, D. Zhang, A. Pathak, C. Chen, H. Xiong, D. Yang, and Y. Wang, “Ccs-ta: quality-guaranteed online task allocation in com-pressive crowdsensing,” in Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing. ACM, 2015, pp. 683–694.
  177. 177.F. Ingelrest, G. Barrenetxea, G. Schaefer, M. Vetterli, O. Couach, and M. Parlange., “SensorScope: Application-specific sensor network for environmental monitoring,” ACM Transactions on Sensor Networks, vol. 6, no. 2, pp. 1–32, 2010.
  178. 178.Y. Zheng, F. Liu, and H. P. Hsieh, “U-Air: when urban air quality inference meets big data,” in ACM SIGKDD Int’l Conf. Knowledge Discovery and Data Mining, 2013.
  179. 179.B. Zhang, C. H. Liu, J. Tang, Z. Xu, J. Ma, and W. Wang, “Learning-based energy-efficient data collection by unmanned vehicles in smart cities,” IEEE Transactions on Industrial Informatics, vol. 14, no. 4, pp. 1666–1676, 2018.
  180. 180.L. Bracciale, M. Bonola, P. Loreti, G. Bianchi, R. Amici, and A. Rabuffi. (2014, Jul.) CRAWDAD dataset roma/taxi (v. 2014-07-17). [Online]. Available: http://crawdad.org/roma/taxi/20140717
  181. 181.L. Xiao, Y. Li, G. Han, H. Dai, and H. V. Poor, “A secure mobile crowdsensing game with deep reinforcement learning,” IEEE Transac-tions on Information Forensics and Security, vol. 13, no. 1, pp. 35–47, Jan. 2018.
  182. 182.Y. Zhang, B. Song, and P. Zhang, “Social behavior study under pervasive social networking based on decentralized deep reinforcement learning,” Journal of Network and Computer Applications, vol. 86, pp. 72–81, 2017.
  183. 183.M. Mohammadi, A. Al-Fuqaha, M. Guizani, and J.-S. Oh, “Semisu-pervised deep reinforcement learning in support of iot and smart city services,” IEEE Internet of Things Journal, vol. 5, no. 2, pp. 624–635, 2018.
  184. 184.D. P. Kingma, S. Mohamed, D. J. Rezende, and M. Welling, “Semi-supervised learning with deep generative models,” in Advances in Neural Information Processing Systems, 2014, pp. 3581–3589.
  185. 185.H. Huang, J. Yang, H. Huang, Y. Song, and G. Gui, “Deep learning for super-resolution channel estimation and doa estimation based massive mimo system,” IEEE Transactions on Vehicular Technology, vol. 67, no. 9, pp. 8549–8560, Sep. 2018.
  186. 186.H. Ye, G. Y. Li, and B. Juang, “Power of deep learning for channel estimation and signal detection in ofdm systems,” IEEE Wireless Communications Letters, vol. 7, no. 1, pp. 114–117, Feb 2018.
  187. 187.A. Zappone, L. Sanguinetti, and M. Debbah, “User association and load balancing for massive mimo through deep learning,” arXiv preprint arXiv:1812.06905, 2018.
  188. 188.J. Desrosiers and M. E. Lübbecke, “Branch-price-and-cut algorithms,” Encyclopedia of Operations Research and Management Science. John Wiley & Sons, Chichester, pp. 109–131, 2011.
  189. 189.X. Wang, L. Gao, S. Mao, and S. Pandey, “Csi-based fingerprinting for indoor localization: A deep learning approach,” IEEE Transactions on Vehicular Technology, vol. 66, no. 1, pp. 763–776, 2017.
  190. 190.Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, “Greedy layer-wise training of deep networks,” in Advances in neural information processing systems, 2007, pp. 153–160.
  191. 191.J. Xiao, K. Wu, Y. Yi, and L. M. Ni, “Fifs: Fine-grained indoor fingerprinting system.” in ICCCN. Citeseer, 2012, pp. 1–7.
  192. 192.J. Vieira, E. Leitinger, M. Sarajlic, X. Li, and F. Tufvesson, “Deep convolutional neural networks for massive mimo fingerprint-based positioning,” in Annual International Symposium on Personal, Indoor, and Mobile Radio Communications, 2017, pp. 1–6.
  193. 193.D. L. Donoho, Y. Tsaig, I. Drori, and J.-L. Starck, “Sparse solution of underdetermined systems of linear equations by stagewise orthogonal matching pursuit,” IEEE Transactions on Information Theory, vol. 58, no. 2, pp. 1094–1121, 2012.
  194. 194.Y. Bai, B. Ai, and W. Chen, “Deep learning based fast multiuser detection for massive machine-type communication,” arXiv preprint arXiv:1807.00967, 2018.
  195. 195.G. Cao, Z. Lu, X. Wen, T. Lei, and Z. Hu, “Aif: An artificial intelligence framework for smart wireless network management,” IEEE Communications Letters, vol. 22, no. 2, pp. 400–403, 2018.
  196. 196.Y. Zhan, Y. Xia, J. Zhang, T. Li, and Y. Wang, “Crowdsensing game with demand uncertainties: A deep reinforcement learning approach,” submitted.
  197. 197.N. C. Luong, P. Wang, D. Niyato, Y. Wen, and Z. Han, “Resource management in cloud networking using economic analysis and pricing models: a survey,” IEEE Communications Surveys & Tutorials, vol. 19, no. 2, pp. 954–1001, Jan. 2017.
  198. 198.N. C. Luong, D. T. Hoang, P. Wang, D. Niyato, D. I. Kim, and Z. Han, “Data collection and wireless communication in internet of things (iot) using economic analysis and pricing models: A survey,” IEEE Communications Surveys & Tutorials, vol. 18, no. 4, pp. 2546–2590, Jun. 2016.
  199. 199.F. Shi, Z. Qin, and J. A. McCann, “Oppay: Design and implementation of a payment system for opportunistic data services,” in IEEE Inter-national Conference on Distributed Computing Systems, Atlanta, GA, Jul. 2017, pp. 1618–1628.
  200. 200.Z. Jiang and J. Liang, “Cryptocurrency portfolio management with deep reinforcement learning,” in Intelligent Systems Conference (IntelliSys), 2017, pp. 905–913.
  201. 201.N. C. Luong, P. Wang, D. Niyato, Y.-C. Liang, F. Hou, and Z. Han, “Ap-plications of economic and pricing models for resource management in 5g wireless networks: A survey,” IEEE Communications Surveys and Tutorials, to appear.
  202. 202.J. Zhao, G. Qiu, Z. Guan, W. Zhao, and X. He, “Deep reinforce-ment learning for sponsored search real-time bidding,” arXiv preprint arXiv:1803.00259, 2018.

Citation

MLA
Luong, N. C., et al. “Applications of Deep Reinforcement Learning in Communications and Networking: A Survey”. arXiv, 2018, http://arxiv.org/abs/1810.07862v1.
APA
Luong, N. C., Hoang, D. T., Gong, S., Niyato, D., Wang, P., Liang, Y.-C., & Kim, D. I. (2018). Applications of Deep Reinforcement Learning in Communications and Networking: A Survey. arXiv. http://arxiv.org/abs/1810.07862v1
Chicago
Luong, N. C., D. T. Hoang, S. Gong, et al. 2018. “Applications of Deep Reinforcement Learning in Communications and Networking: A Survey”. arXiv. http://arxiv.org/abs/1810.07862v1.
Harvard
Luong, N.C. et al. (2018) “Applications of Deep Reinforcement Learning in Communications and Networking: A Survey”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1810.07862v1.
Vancouver
1. Luong NC, Hoang DT, Gong S, Niyato D, Wang P, Liang Y-C, Kim DI (2018) Applications of Deep Reinforcement Learning in Communications and Networking: A Survey. arXiv

BibTeX

@article{luong2018applications,
  title = {Applications of Deep Reinforcement Learning in Communications and Networking: A Survey},
  author = {Luong, Nguyen Cong and Hoang, Dinh Thai and Gong, Shimin and Niyato, Dusit and Wang, Ping and Liang, Ying-Chang and Kim, Dong In},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1810.07862v1},
  eprint = {1810.07862}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF