A Joint Learning and Communications Framework for Federated Learning Over Wireless Networks

Mingzhe ChenZhaohui YangW. SaadChangchuan YinH. PoorShuguang Cui

article2019IEEE Transactions on Wireless Communications1,469 citations2023 IEEE Marconi Prize Paper Award in Wireless Communications

Develops an analytical framework for wireless federated learning that derives the mathematical impact of channel packet errors on convergence and jointly optimizes user selection, transmit power, and resource allocation to minimize model loss.

Listen

Federated learning enables mobile devices to collaboratively train artificial intelligence models using local data, protecting user privacy and avoiding the massive bandwidth costs of centralizing raw datasets. However, deploying federated learning over commercial wireless cellular networks introduces operational challenges. Wireless channels suffer from interference, packet transmission errors, latency, and limited spectrum and battery energy. When device updates are corrupted or dropped, the accuracy and convergence speed of the shared machine learning model degrade significantly. Addressing these communication bottlenecks is essential for making decentralized intelligence practical on mobile and edge devices.

The main objective of the article is to establish and evaluate a joint learning and communication framework that minimizes machine learning training loss over wireless networks by systematically optimizing user selection, wireless resource allocation, and device transmit power.

To achieve this, the article analyzes a cellular uplink network where distributed devices transmit local model updates to a central base station using orthogonal frequency division multiple access. The authors derive a closed-form mathematical expression for the expected convergence rate of the federated learning algorithm under packet transmission errors. Using this theoretical convergence bound, they simplify the training loss minimization problem into a bipartite matching formulation that strictly enforces per-round transmission delay and device energy limits. The base station determines optimal device transmit powers analytically and allocates wireless resource blocks using the standard Hungarian matching algorithm. The framework is validated through numeric simulations using synthetic linear regression tasks and image classification on the MNIST handwritten digit dataset across varying user counts and bandwidth allocations.

The findings show that wireless impairments create a quantifiable performance gap between practical federated learning and idealized error-free training, but this gap can be minimized through joint network and learning optimization. In image identification tests, the proposed joint framework improves model accuracy by up to 1.4% compared to optimal user selection with random resource allocation, by 3.5% compared to standard wireless federated learning with random user selection, and by 4.1% compared to conventional wireless resource management that minimizes packet errors without considering machine learning parameters. The theoretical convergence rate closely aligns with simulation results, exhibiting less than a 9% discrepancy. Furthermore, the analysis demonstrates that while increasing user participation supplies more data to improve model quality, performance gains plateau once the base station receives sufficient data samples (e.g., beyond approximately 30 samples per user in linear regression), meaning network managers do not need to schedule every device in every round.

These results indicate that treating wireless network management and machine learning design as separate domains leads to suboptimal artificial intelligence performance. Cellular operators and edge computing architects can implement learning-aware scheduling at the base station to achieve higher prediction accuracy and faster model convergence without requiring additional hardware or spectral bandwidth. By accounting for device sample counts, channel reliability, and local energy constraints simultaneously, network operators reduce operational risks such as training stalls, battery drain, and excessive latency in time-sensitive edge applications.

Organizations deploying federated learning over wireless infrastructure should integrate joint user selection and resource allocation algorithms into base station schedulers. Operators should prioritize scheduling devices that possess both high data quality and favorable channel conditions to maximize learning efficiency per round. Further development should focus on deploying small-scale field pilots in live cellular networks and extending the framework to handle non-orthogonal multiple access, asynchronous model updates, and multi-cell interference environments.

The primary analytical derivations rely on standard convex optimization assumptions regarding the loss functions and assume that corrupted packets are discarded rather than retransmitted. Nevertheless, simulation results on non-convex convolutional neural networks confirm that the optimization framework remains robust and effective in practical deep learning scenarios, providing high confidence in the operational conclusions.

Cover for A Joint Learning and Communications Framework for Federated Learning Over Wireless Networks

Abstract

In this paper, the problem of training federated learning (FL) algorithms over a realistic wireless network is studied. In the considered model, wireless users execute an FL algorithm while training their local FL models using their own data and transmitting the trained local FL models to a base station (BS) that generates a global FL model and sends the model back to the users. Since all training parameters are transmitted over wireless links, the quality of training is affected by wireless factors such as packet errors and the availability of wireless resources. Meanwhile, due to the limited wireless bandwidth, the BS needs to select an appropriate subset of users to execute the FL algorithm so as to build a global FL model accurately. This joint learning, wireless resource allocation, and user selection problem is formulated as an optimization problem whose goal is to minimize an FL loss function that captures the performance of the FL algorithm. To seek the solution, a closed-form expression for the expected convergence rate of the FL algorithm is first derived to quantify the impact of wireless factors on FL. Then, based on the expected convergence rate of the FL algorithm, the optimal transmit power for each user is derived, under a given user selection and uplink resource block (RB) allocation scheme. Finally, the user selection and uplink RB allocation is optimized so as to minimize the FL loss function.

Table of Contents

  • I Introduction
  • I-A Related Works
  • I-B Contributions
  • II System Model and Problem Formulation
  • II-A Machine Learning Model
  • II-B Transmission Model
  • II-C Packet Error Rates
  • II-D Energy Consumption Model
  • II-E Problem Formulation
  • III Analysis of the FL Convergence Rate
  • IV Optimization of FL Training Loss
  • IV-A Optimal Transmit Power
  • IV-B Optimal Uplink Resource Block Allocation
  • IV-C Implementation and Complexity
  • V Simulation Results and Analysis
  • V-A FL for Linear Regression
  • V-B FL for Handwritten Digit Identification
  • VI Conclusion
  • VI-A Proof of Theorem
  • VI-B Proof of Proposition
  • References

Knowls

  1. Knowl 1 — Upper Bound on Expected Convergence Rate of Wireless Federated Learning

    theoretical result

    Consider a federated learning (FL) system executed over an orthogonal frequency division multiple access (OFDMA) wireless cellular network with a base station (BS) and a set U\mathcal{U} of UU users. Let KiK_i be the number of training samples at user ii, K=∑i=1UKiK = \sum_{i=1}^U K_i, ai∈{0,1}a_i \in \{0, 1\} be the indicator of whether user ii is selected (ai=1a_i = 1) or not (ai=0a_i = 0), rir_i be the resource block (RB) allocation vector for user ii, and PiP_i be its transmit power. Let qi(ri,Pi)q_i(r_i, P_i) denote user ii's uplink packet error rate.

    Assume the global loss function F(g)=1K∑i=1U∑k=1Kif(g,xik,yik)F(g) = \frac{1}{K}\sum_{i=1}^U \sum_{k=1}^{K_i} f(g, x_{ik}, y_{ik}) is LL-smooth and μ\mu-strongly convex, and local gradients satisfy ∥∇f(gt,xik,yik)∥2≤ζ1+ζ2∥∇F(gt)∥2\|\nabla f(g_t, x_{ik}, y_{ik})\|^2 \le \zeta_1 + \zeta_2 \|\nabla F(g_t)\|^2 for non-negative constants ζ1,ζ2\zeta_1, \zeta_2. Using a gradient descent update step with learning rate λ=1/L\lambda = 1/L, the expected suboptimality gap of the global model gt+1g_{t+1} after t+1t+1 communication rounds satisfies the upper bound:

    E[F(gt+1)−F(g∗)]≤AtE[F(g0)−F(g∗)]+2ζ1LK∑i=1UKi(1−ai+aiqi(ri,Pi))1−At1−A\mathbb{E}\left[F(g_{t+1}) - F(g^*)\right] \le A^t \mathbb{E}\left[F(g_0) - F(g^*)\right] + \frac{2\zeta_1}{LK} \sum_{i=1}^U K_i \left(1 - a_i + a_i q_i(r_i, P_i)\right) \frac{1 - A^t}{1 - A}

    where g∗g^* is the optimal global model in an error-free centralized setting, g0g_0 is the initial model, and the convergence factor AA is defined as:

    A=1−μL+4μζ2LK∑i=1UKi(1−ai+aiqi(ri,Pi))A = 1 - \frac{\mu}{L} + \frac{4\mu\zeta_2}{LK} \sum_{i=1}^U K_i \left(1 - a_i + a_i q_i(r_i, P_i)\right)

    When packet transmission is error-free (qi=0q_i = 0) and all users participate (ai=1a_i = 1), the second additive gap term vanishes and the rate reduces to the standard linear convergence bound E[F(gt+1)−F(g∗)]≤(1−μ/L)tE[F(g0)−F(g∗)]\mathbb{E}\left[F(g_{t+1}) - F(g^*)\right] \le (1 - \mu/L)^t \mathbb{E}\left[F(g_0) - F(g^*)\right].

  2. Knowl 2 — Global Model Aggregation under Wireless Packet Transmission Errors

    model/method

    In a wireless federated learning system with UU distributed users and one base station (BS), each selected user ii trains its local parameter vector wiw_i and transmits it over an assigned wireless uplink resource block (RB). Due to channel fading, noise, and co-channel interference, the uplink transmission of wiw_i (transmitted as a single packet) experiences packet errors verified via cyclic redundancy check (CRC).

    The global FL model g(a,P,R)g(a, P, R) aggregated at the BS at each round is defined as:

    g(a,P,R)=∑i=1UKiaiwiC(wi)∑i=1UKiaiC(wi)g(a, P, R) = \frac{\sum_{i=1}^U K_i a_i w_i C(w_i)}{\sum_{i=1}^U K_i a_i C(w_i)}

    where:

    • KiK_i is the number of local data samples collected by user ii.
    • ai∈{0,1}a_i \in \{0, 1\} is the user selection variable (ai=1a_i = 1 if user ii is selected, ai=0a_i = 0 otherwise).
    • C(wi)∈{0,1}C(w_i) \in \{0, 1\} is the packet reception indicator, modeled as a Bernoulli random variable where C(wi)=1C(w_i) = 1 with probability 1−qi(ri,Pi)1 - q_i(r_i, P_i) (correct reception) and C(wi)=0C(w_i) = 0 with probability qi(ri,Pi)q_i(r_i, P_i) (packet error).
    • qi(ri,Pi)=∑n=1Rri,nqi,nq_i(r_i, P_i) = \sum_{n=1}^R r_{i,n} q_{i,n} is the packet error rate experienced by user ii, with ri,n∈{0,1}r_{i,n} \in \{0, 1\} indicating allocation of RB n∈{1,…,R}n \in \{1, \dots, R\}, and qi,n=Ehi[1−exp⁡(−m(In+BUN0)Pihi)]q_{i,n} = \mathbb{E}_{h_i}\left[1 - \exp\left(-\frac{m(I_n + B^U N_0)}{P_i h_i}\right)\right], where PiP_i is user transmit power, hih_i is the channel gain, InI_n is inter-cell interference on RB nn, BUB^U is the RB bandwidth, N0N_0 is noise power spectral density, and mm is a waterfall threshold.

    When a packet contains errors (C(wi)=0C(w_i) = 0), the BS discards the local update without requesting retransmission and updates the global model using only the correctly received updates.

  3. Knowl 3 — Joint Learning, User Selection, and Resource Allocation Optimization Problem

    model/method

    The joint optimization of user selection a=[a1,…,aU]a = [a_1, \dots, a_U], transmit power P=[P1,…,PU]P = [P_1, \dots, P_U], and uplink resource block (RB) allocation R=[r1,…,rU]R = [r_1, \dots, r_U] to minimize the global federated learning training loss under wireless delay and energy constraints is formulated as:

    min⁡a,P,R1K∑i=1U∑k=1Kif(g(a,P,R),xik,yik)\min_{a, P, R} \frac{1}{K} \sum_{i=1}^U \sum_{k=1}^{K_i} f(g(a, P, R), x_{ik}, y_{ik})

    s.t.ai,ri,n∈{0,1},∀i∈U,  n=1,…,R\text{s.t.} \quad a_i, r_{i,n} \in \{0, 1\}, \quad \forall i \in \mathcal{U}, \; n = 1, \dots, R

    ∑n=1Rri,n=ai,∀i∈U\sum_{n=1}^R r_{i,n} = a_i, \quad \forall i \in \mathcal{U}

    liU(ri,Pi)+liD≤γT,∀i∈Ul_i^U(r_i, P_i) + l_i^D \le \gamma_T, \quad \forall i \in \mathcal{U}

    ei(ri,Pi)≤γE,∀i∈Ue_i(r_i, P_i) \le \gamma_E, \quad \forall i \in \mathcal{U}

    ∑i∈Uri,n≤1,∀n=1,…,R\sum_{i \in \mathcal{U}} r_{i,n} \le 1, \quad \forall n = 1, \dots, R

    0≤Pi≤Pmax⁡,∀i∈U0 \le P_i \le P_{\max}, \quad \forall i \in \mathcal{U}

    where:

    • K=∑i=1UKiK = \sum_{i=1}^U K_i is the total dataset size across all users U\mathcal{U}, and (xik,yik)(x_{ik}, y_{ik}) denotes data sample kk of user ii.
    • liU(ri,Pi)=Z(wi)ciU(ri,Pi)l_i^U(r_i, P_i) = \frac{Z(w_i)}{c_i^U(r_i, P_i)} is the uplink transmission delay with data size Z(wi)Z(w_i) bits and rate ciU(ri,Pi)=∑n=1Rri,nBUEhi[log⁡2(1+PihiIn+BUN0)]c_i^U(r_i, P_i) = \sum_{n=1}^R r_{i,n} B^U \mathbb{E}_{h_i}\left[\log_2\left(1 + \frac{P_i h_i}{I_n + B^U N_0}\right)\right].
    • liD=Z(g)ciDl_i^D = \frac{Z(g)}{c_i^D} is the downlink transmission delay with rate ciD=BDEhi[log⁡2(1+PBhiID+BDN0)]c_i^D = B^D \mathbb{E}_{h_i}\left[\log_2\left(1 + \frac{P_B h_i}{I^D + B^D N_0}\right)\right], where BDB^D is downlink bandwidth, PBP_B is BS transmit power, and IDI^D is interference from other BSs.
    • ei(ri,Pi)=ςωiϑ2Z(wi)+PiliU(ri,Pi)e_i(r_i, P_i) = \varsigma \omega_i \vartheta^2 Z(w_i) + P_i l_i^U(r_i, P_i) is user ii's total energy consumption per round (local computing plus uplink transmission), with CPU frequency ϑ\vartheta, compute cycles per bit ωi\omega_i, and hardware energy coefficient ς\varsigma.
    • γT\gamma_T and γE\gamma_E denote the per-round latency and energy consumption limits, and Pmax⁡P_{\max} is the maximum user transmit power.
  4. Knowl 4 — Optimal Transmit Power Allocation for Wireless FL Users

    theoretical result

    Given an assigned uplink resource block (RB) allocation vector rir_i for user ii, the per-round energy consumption ei(ri,Pi)=ςωiϑ2Z(wi)+PiZ(wi)ciU(ri,Pi)e_i(r_i, P_i) = \varsigma \omega_i \vartheta^2 Z(w_i) + \frac{P_i Z(w_i)}{c_i^U(r_i, P_i)} is strictly monotonically increasing with respect to transmit power Pi>0P_i > 0. Because the packet error rate qi(ri,Pi)q_i(r_i, P_i) strictly decreases as transmit power increases, minimizing the FL loss gap requires maximizing transmit power up to the energy constraint ei(ri,Pi)≤γEe_i(r_i, P_i) \le \gamma_E and the device power ceiling Pmax⁡P_{\max}.

    The optimal transmit power Pi∗(ri)P_i^*(r_i) of user ii is uniquely given by:

    Pi∗(ri)=min⁡{Pmax⁡,Pi,γE}P_i^*(r_i) = \min\{P_{\max}, P_{i, \gamma_E}\}

    where Pi,γEP_{i, \gamma_E} is the unique positive root of the equality:

    ςωiϑ2Z(wi)+Pi,γEZ(wi)ciU(ri,Pi,γE)=γE\varsigma \omega_i \vartheta^2 Z(w_i) + \frac{P_{i, \gamma_E} Z(w_i)}{c_i^U(r_i, P_{i, \gamma_E})} = \gamma_E

    with Z(wi)Z(w_i) being the local model size in bits, ciU(ri,Pi)c_i^U(r_i, P_i) the uplink data rate on the assigned RB, ϑ\vartheta the user CPU clock frequency, ωi\omega_i the required CPU cycles per bit, and ς\varsigma the device chip energy coefficient.

  5. Knowl 5 — Asymptotic FL Loss Simplification and Bipartite Matching Reformulation

    model/method

    In the asymptotic regime (t→∞t \to \infty) where the convergence coefficient satisfies A<1A < 1 (such that At→0A^t \to 0), minimizing the expected convergence gap 2ζ1LK∑i=1UKi(1−ai+aiqi(ri,Pi∗))11−A\frac{2\zeta_1}{LK}\sum_{i=1}^U K_i(1 - a_i + a_i q_i(r_i, P_i^*))\frac{1}{1-A} reduces to minimizing ∑i=1UKi(1−∑n=1Rri,n+∑n=1Rri,nqi,n)\sum_{i=1}^U K_i\left(1 - \sum_{n=1}^R r_{i,n} + \sum_{n=1}^R r_{i,n} q_{i,n}\right).

    This optimization problem is equivalently mapped to a minimum weight bipartite matching problem on the graph A=(U×R,E)\mathcal{A} = (\mathcal{U} \times \mathcal{R}, \mathcal{E}), where U\mathcal{U} represents the set of UU users, R\mathcal{R} represents the set of RR orthogonal resource blocks, and E\mathcal{E} is the set of user-RB candidate edges. The weight ψin\psi_{in} assigned to edge χin\chi_{in} connecting user ii and RB nn is defined as:

    ψin={Ki(qi,n−1),if liU(ri,n,Pi∗)+liD≤γT and ei(ri,n,Pi∗)≤γE,0,otherwise\psi_{in} = \begin{cases} K_i (q_{i,n} - 1), & \text{if } l_i^U(r_{i,n}, P_i^*) + l_i^D \le \gamma_T \text{ and } e_i(r_{i,n}, P_i^*) \le \gamma_E, \\ 0, & \text{otherwise} \end{cases}

    where KiK_i is user ii's dataset size, qi,nq_{i,n} is the packet error rate over RB nn, liUl_i^U and liDl_i^D are uplink and downlink delays, and eie_i is energy consumption. If delay or energy constraints cannot be met when assigning RB nn to user ii, setting ψin=0\psi_{in} = 0 prevents matching.

    The global minimum is solved via the Hungarian algorithm with a worst-case computational complexity of O(U2R)\mathcal{O}(U^2 R) and best-case complexity of O(UR)\mathcal{O}(UR). The resulting matching yields the optimal RB allocation R∗R^* and optimal user selection ai∗=∑n=1Rri,n∗a_i^* = \sum_{n=1}^R r_{i,n}^*.

  6. Knowl 6 — Condition on Gradient Variance Bound for Convergence of Wireless FL

    theoretical result

    Let the local sample loss gradients satisfy ∥∇f(gt,xik,yik)∥2≤ζ1+ζ2∥∇F(gt)∥2\|\nabla f(g_t, x_{ik}, y_{ik})\|^2 \le \zeta_1 + \zeta_2 \|\nabla F(g_t)\|^2 with constants ζ1,ζ2≥0\zeta_1, \zeta_2 \ge 0. Under learning rate λ=1/L\lambda = 1/L, strong convexity parameter μ\mu, and Lipschitz constant LL, to guarantee convergence of the wireless federated learning algorithm (i.e., ensuring the convergence contraction parameter A<1A < 1), the parameter ζ2\zeta_2 must strictly satisfy:

    0<ζ2<Kmax⁡P,R4∑i=1UKiqi(ri,Pi)0 < \zeta_2 < \frac{K}{\max_{P, R} 4 \sum_{i=1}^U K_i q_i(r_i, P_i)}

    where K=∑i=1UKiK = \sum_{i=1}^U K_i is the total dataset size across all UU users, and qi(ri,Pi)q_i(r_i, P_i) is the packet error rate of user ii under transmit power vector PP and resource block allocation matrix RR.

  7. Knowl 7 — Joint User Selection, Resource Allocation, and Transmit Power Optimization Algorithm

    algorithm

    The joint federated learning and wireless resource allocation framework determines the optimal user selection vector a∗a^*, RB allocation matrix R∗R^*, and user transmit power vector P∗P^* to minimize global FL loss while respecting per-round latency and energy budgets.

    Input: Uplink rates ciU(ri,Pi)c_i^U(r_i, P_i), downlink rates ciDc_i^D, model data size Z(wi)Z(w_i), packet error rates qi(ri,Pi)q_i(r_i, P_i) for all i∈Ui \in \mathcal{U}, delay limit γT\gamma_T, energy limit γE\gamma_E
    Output: Optimal user selection a∗a^*, RB allocation matrix R∗R^*, transmit power vector P∗P^*
    1: for each user i∈Ui \in \mathcal{U} and each RB n∈{1,…,R}n \in \{1, \dots, R\} do
    2: Calculate optimal transmit power Pi∗(ri,n)=min⁡{Pmax⁡,Pi,γE}P_i^*(r_{i,n}) = \min\{P_{\max}, P_{i, \gamma_E}\}
    3: Compute delay liU(ri,n,Pi∗)+liDl_i^U(r_{i,n}, P_i^*) + l_i^D and energy ei(ri,n,Pi∗)e_i(r_{i,n}, P_i^*)
    4: if liU+liD≤γTl_i^U + l_i^D \le \gamma_T and ei≤γEe_i \le \gamma_E then
    5: Set edge weight ψin=Ki(qi,n−1)\psi_{in} = K_i (q_{i,n} - 1)
    6: else
    7: Set edge weight ψin=0\psi_{in} = 0
    8: end if
    9: end for
    10: Solve the bipartite matching problem with weight matrix [ψin][\psi_{in}] using the Hungarian algorithm to obtain optimal RB assignment R∗R^*
    11: Set optimal user selection ai∗=∑n=1Rri,n∗a_i^* = \sum_{n=1}^R r_{i,n}^* for each user i∈Ui \in \mathcal{U}
    12: Set optimal transmit power Pi∗=Pi∗(ri∗)P_i^* = P_i^*(r_i^*) for each user i∈Ui \in \mathcal{U}
    13: Execute FL training rounds using a∗a^*, R∗R^*, and P∗P^*

    The BS centralized execution requires O(UR)\mathcal{O}(UR) evaluations to set up the bipartite graph weights and O(U2R)\mathcal{O}(U^2 R) operations in the worst case to solve the Hungarian assignment.

  8. Knowl 8 — Global and Local Loss Function Assumptions for Wireless FL Convergence

    assumption

    Let F(g)=1K∑i=1U∑k=1Kif(g,xik,yik)F(g) = \frac{1}{K}\sum_{i=1}^U \sum_{k=1}^{K_i} f(g, x_{ik}, y_{ik}) denote the global empirical loss function over model parameters gg, where K=∑i=1UKiK = \sum_{i=1}^U K_i. The theoretical convergence analysis of the wireless FL framework rests on four mathematical assumptions:

    1. Uniform Lipschitz Continuous Gradient: The gradient ∇F(g)\nabla F(g) is Lipschitz continuous with constant L>0L > 0: ∥∇F(gt+1)−∇F(gt)∥≤L∥gt+1−gt∥\|\nabla F(g_{t+1}) - \nabla F(g_t)\| \le L \|g_{t+1} - g_t\|

    2. Strong Convexity: F(g)F(g) is strongly convex with parameter μ>0\mu > 0: F(gt+1)≥F(gt)+(gt+1−gt)T∇F(gt)+μ2∥gt+1−gt∥2F(g_{t+1}) \ge F(g_t) + (g_{t+1} - g_t)^T \nabla F(g_t) + \frac{\mu}{2}\|g_{t+1} - g_t\|^2

    3. Twice Continuous Differentiability: F(g)F(g) is twice-continuously differentiable, which implies: μI⪯∇2F(g)⪯LI\mu I \preceq \nabla^2 F(g) \preceq L I where II is the identity matrix.

    4. Bounded Gradient Dissimilarity: For all users ii and samples kk, local sample gradients satisfy: ∥∇f(gt,xik,yik)∥2≤ζ1+ζ2∥∇F(gt)∥2\|\nabla f(g_t, x_{ik}, y_{ik})\|^2 \le \zeta_1 + \zeta_2 \|\nabla F(g_t)\|^2 where ζ1,ζ2≥0\zeta_1, \zeta_2 \ge 0 are non-negative constants.

  9. Knowl 9 — Empirical Performance and Accuracy Gains of Wireless FL Framework

    empirical result

    The proposed joint learning and communications framework is evaluated on linear regression (fitting y=−2x+1+0.4ny = -2x + 1 + 0.4n, n∼N(0,1)n \sim \mathcal{N}(0, 1) using a 20-neuron feedforward neural network) and MNIST handwritten digit identification (50-neuron FNN and CNN, cross-entropy loss) in a circular cell of radius r=500r = 500 m with U=15U = 15 users and R∈[3,15]R \in [3, 15] resource blocks.

    Key simulation parameters include user transmit power limit Pmax⁡=0.01P_{\max} = 0.01 W, BS power PB=1P_B = 1 W, downlink bandwidth BD=20B^D = 20 MHz, uplink RB bandwidth BU=1B^U = 1 MHz, noise density N0=−174N_0 = -174 dBm/Hz, delay limit γT=500\gamma_T = 500 ms, energy limit γE=0.003\gamma_E = 0.003 J, ϑ=109\vartheta = 10^9 Hz, ς=10−27\varsigma = 10^{-27}, and ωi=40\omega_i = 40 cycles/bit.

    Key empirical findings include:

    • Accuracy Gains: In MNIST handwritten digit identification with R=9R = 9 RBs, the proposed algorithm improves classification accuracy by up to 1.4% over an optimal user selection baseline with random RB allocation, by up to 3.5% over standard FL with random user selection and random RB allocation, and by up to 4.1% over a wireless optimization baseline that minimizes sum packet error rates while ignoring FL parameters.
    • CNN Generalization: On non-convex CNN training with MNIST, the proposed algorithm correctly classifies 30 out of 36 sample digits compared to 27 out of 36 for standard FL with random selection and resource allocation.
    • Theoretical Validation: The theoretical asymptotic convergence gap matches empirical simulation values with less than 9% difference across varying user counts.

Coverage note — Intermediate algebraic proof steps for Theorem 1 in Appendix A and Proposition 2 in Appendix B were omitted in accordance with the requirement to exclude proof derivations.

References

  1. 1.M. Chen, Z. Yang, W. Saad, C. Yin, H. V. Poor, and S. Cui, “Performance optimization of federated learning over wireless networks,” in Proc. IEEE Global Commun. Conf., Waikoloa, HI, USA, December 2019.
  2. 2.M. Chen, U. Challita, W. Saad, C. Yin, and M. Debbah, “Artificial neural networks-based machine learning for wireless networks: A tutorial,” IEEE Commun. Surveys Tut., vol. 21, no. 4, pp. 3039–3071, Fourthquarter 2019.
  3. 3.Y. Sun, M. Peng, Y. Zhou, Y. Huang, and S. Mao, “Application of machine learning in wireless networks: Key techniques and open issues,” IEEE Commun. Surveys Tut., vol. 21, no. 4, pp. 3072–3108, Fourthquarter 2019.
  4. 4.Y. Liu, S. Bi, Z. Shi, and L. Hanzo, “When machine learning meets big data: A wireless communication perspective,” IEEE Veh. Technol. Mag., vol. 15, no. 1, pp. 63–72, March 2020.
  5. 5.K. Bonawitz, H. Eichner, W. Grieskamp, D. Huba, A. Ingerman, V. Ivanov, C. M. Kiddon, J. Konecny, S. Mazzocchi, B. McMahan, T. V. Overveldt, D. Petrou, D. Ramage, and J. Roselander, “Towards federated learning at scale: System design,” in Proc. Systems and Machine Learning Conference, Stanford, CA, USA, 2019.
  6. 6.V. Smith, C. K. Chiang, M. Sanjabi, and A. S. Talwalkar, “Federated multi-task learning,” in Proc. Advances in Neural Information Processing Systems, Long Beach, CA, USA, Dec. 2017.
  7. 7.X. Wang, Y. Han, C. Wang, Q. Zhao, X. Chen, and M. Chen, “In-edge AI: Intelligentizing mobile edge computing, caching and communication by federated learning,” IEEE Netw., vol. 33, no. 5, pp. 156–165, Sep. 2019.
  8. 8.W. Saad, M. Bennis, and M. Chen, “A vision of 6G wireless systems: Applications, trends, technologies, and open research problems,” IEEE Netw., vol. 34, no. 3, pp. 134–142, May 2020.
  9. 9.E. Jeong, S. Oh, J. Park, H. Kim, M. Bennis, and S. L. Kim, “Multi-hop federated private data augmentation with sample compression,” in Proc. International Joint Conference on Artificial Intelligence Workshop on Federated Machine Learning for User Privacy and Data Confidentiality, Macao, China, Aug. 2019.
  10. 10.Z. Yang, M. Chen, W. Saad, C. S. Hong, and M. Shikh-Bahaei, “Energy efficient federated learning over wireless communication networks,” arXiv preprint arXiv:1911.02417, 2019.
  11. 11.T. Li, A. K. Sahu, A. Talwalkar, and V. Smith, “Federated learning: Challenges, methods, and future directions,” IEEE Signal Process. Mag., vol. 37, no. 3, pp. 50–60, May 2020.
  12. 12.J. Konećnȃ y, H. B. McMahan, D. Ramage, and P. Richtárik, “Federated optimization: Distributed machine learning for ` on-device intelligence,” arXiv preprint arXiv:1610.02527, Oct. 2016.
  13. 13.H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Proc. Conf. Machine Learning Research, Fort Lauderdale, FL, USA, Apr. 2017.
  14. 14.M. Chen, O. Semiari, W. Saad, X. Liu, and C. Yin, “iterated echo state learning for minimizing breaks in presence in wireless virtual reality networks,” IEEE Trans. Wireless Commun., vol. 19, no. 1, pp. 177–191, Jan. 2020.
  15. 15.J. Konećnȃ y, B. McMahan, and D. Ramage, “iterated optimization: Distributed optimization beyond the datacenter,” ` arXiv preprint arXiv:1511.03575, Nov. 2015.
  16. 16.S. Samarakoon, M. Bennis, W. Saad, and M. Debbah, “Distributed federated learning for ultra-reliable low-latency vehicular communications,” IEEE Trans. Commun., vol. 68, no. 2, pp. 1146–1159, Feb. 2020.
  17. 17.S. Ha, J. Zhang, O. Simeone, and J. Kang, “iterated federated computing in wireless networks with straggling devices and imperfect CSI,” in Proc. IEEE Int. Symp. Inf. Theory (ISIT), Paris, France, Jul. 2019.
  18. 18.O. Habachi, M. A. Adjif, and J. P. Cances, “iterated uplink grant for NOMA: A federated learning based approach,” arXiv preprint arXiv:1904.07975, Mar. 2019.
  19. 19.J. Park, S. Samarakoon, M. Bennis, and M. Debbah, “iterated network intelligence at the edge,” Proc. IEEE, vol. 107, no. 11, pp. 2204–2239, Nov. 2019.
  20. 20.Q. Zeng, Y. Du, K. Huang, and K. K. Leung, “iterated-efficient radio resource allocation for federated edge learning,” in Proc. IEEE Int. Conf. Commun. Workshop, Dublin, Ireland, Jun. 2020.
  21. 21.S. Wang, T. Tuor, T. Salonidis, K. K. Leung, C. Makaya, T. He, and K. Chan, “iterated federated learning in resource constrained edge computing systems,” IEEE J. Sel. Areas Commun., vol. 37, no. 6, pp. 1205–1221, June 2019.
  22. 22.Z. Zhao, C. Feng, H. H. Yang, and X. Luo, “iterated-learning-enabled intelligent fog radio access networks: Fundamental theory, key techniques, and future trends,” IEEE Wireless Commun., vol. 27, no. 2, pp. 22–28, April 2020.
  23. 23.T. T. Vu, D. T. Ngo, N. H. Tran, H. Q. Ngo, M. N. Dao, and R. H. Middleton, “iterated-free massive MIMO for wireless federated learning,” IEEE Tans. Wireless Commun., to appear, 2020.
  24. 24.N. H. Tran, W. Bao, A. Zomaya, N. Minh N.H., and C. S. Hong, “iterated learning over wireless networks: Optimization model design and analysis,” in Proc. IEEE Int. Conf. Compt. Commun., Paris, France, April 2019.
  25. 25.M. Chen, H. V. Poor, W. Saad, and S. Cui, “Convergence time optimization for federated learning over wireless networks,” arXiv preprint arXiv:2001.07845, 2020.
  26. 26.H. H. Yang, Z. Liu, T. Q. S. Quek, and H. V. Poor, “iterated policies for federated learning in wireless networks,” IEEE Tans. Commun., vol. 68, no. 1, pp. 317–333, Jan. 2020.
  27. 27.S. Bi, J. Lyu, Z. Ding, and R. Zhang, “iterated radio maps for wireless resource management,” IEEE Wireless Commun., vol. 26, no. 2, pp. 133–141, April 2019.
  28. 28.C. Hennig and M. Kutlukaya, “iterated thoughts about the design of loss functions,” Revstat Stat. J., vol. 5, no. 1, pp. 19–39, March 2007.
  29. 29.J. Konecny, H. B. McMahan, F. X. Yu, P. Richtarik, A. Theertha Suresh, and D. Bacon, “iterated learning: Strategies for improving communication efficiency,” in Proc. NIPS Workshop on Private Multi-Party Machine Learning, Barcelona, Spain, Dec. 2016.
  30. 30.Y. Xi, A. Burr, J. Wei, and D. Grace, “iterated general upper bound to evaluate packet error rate over quasi-static fading channels,” IEEE Tans. Wireless Commun., vol. 10, no. 5, pp. 1373–1377, May 2011.
  31. 31.Y. Pan, C. Pan, Z. Yang, and M. Chen, “iterated allocation for D2D communications underlaying a NOMA-based cellular network,” IEEE Wireless Commun. Lett., vol. 7, no. 1, pp. 130–133, Feb 2018.
  32. 32.M. P. Friedlander and M. Schmidt, “iterated deterministic-stochastic methods for data fitting,” SIAM J. Sci. Comput., vol. 34, no. 3, pp. A1380–A1405, May 2012.
  33. 33.M. Mahdian and Q. Yan, “iterated bipartite matching with random arrivals: An approach based on strongly factor-revealing LPs,” in Proc. ACM Symposium on Theory of Computing, San Jose, California, USA, June 2011.
  34. 34.R. Jonker and T. Volgenant, “iterated the Hungarian assignment algorithm,” Oper. Res. Lett., vol. 5, no. 4, pp. 171–175, 1986.
  35. 35.N. Landman K. Moore and J. Khim, “iterated maximum matching algorithm,” https://brilliant.org/wiki/ hungarian-matching/.
  36. 36.Y. LeCun, “iterated MNIST database of handwritten digits,” http://yann.lecun.com/exdb/mnist/.
  37. 37.S. Boyd and L. Vandenberghe, Convex Optimization, Cambridge University Press, 2004.

Citation

MLA
Chen, M., et al. “A Joint Learning and Communications Framework for Federated Learning over Wireless Networks”. arXiv, 2019, http://arxiv.org/abs/1909.07972v4.
APA
Chen, M., Yang, Z., Saad, W., Yin, C., Poor, H. V., & Cui, S. (2019). A Joint Learning and Communications Framework for Federated Learning over Wireless Networks. arXiv. http://arxiv.org/abs/1909.07972v4
Chicago
Chen, M., Z. Yang, W. Saad, C. Yin, H. V. Poor, and S. Cui. 2019. “A Joint Learning and Communications Framework for Federated Learning over Wireless Networks”. arXiv. http://arxiv.org/abs/1909.07972v4.
Harvard
Chen, M. et al. (2019) “A Joint Learning and Communications Framework for Federated Learning over Wireless Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1909.07972v4.
Vancouver
1. Chen M, Yang Z, Saad W, Yin C, Poor HV, Cui S (2019) A Joint Learning and Communications Framework for Federated Learning over Wireless Networks. arXiv

BibTeX

@article{chen2019joint,
  title = {A Joint Learning and Communications Framework for Federated Learning over Wireless Networks},
  author = {Chen, Mingzhe and Yang, Zhaohui and Saad, Walid and Yin, Changchuan and Poor, H. Vincent and Cui, Shuguang},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1909.07972v4},
  eprint = {1909.07972}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF