An Aggregation-Free Federated Learning for Tackling Data Heterogeneity

Yuan WangHuazhu FuRenuga KanagaveluQingsong WeiYong LiuRick Siow Mong Goh

article2024CVPR74 citations

Proposes FedAF, an aggregation-free federated learning framework where clients generate synthetic condensed data and soft labels via Sliced Wasserstein Distance regularization, eliminating client drift and significantly boosting convergence speed and accuracy on heterogeneous datasets.

Listen

Federated learning enables multiple distributed clients to train a shared global artificial intelligence model without sharing their raw, private datasets. However, real-world deployments frequently suffer from significant data heterogeneity, where client data distributions differ sharply in class labels or visual features. In standard frameworks, this diversity causes client drift and catastrophic forgetting, as local updates diverge from the central learning goal, severely degrading overall model accuracy and slowing convergence.

The article introduces FedAF, an aggregation-free federated learning framework designed to eliminate client drift across heterogeneous networks. FedAF replaces the standard cycle of averaging local model weights with a framework where clients condense their local datasets into compact synthetic samples and share them, along with predictive soft labels, for the central server to train the global model directly.

To evaluate this framework, the authors conducted extensive benchmark experiments across image classification datasets, including Fashion-MNIST, CIFAR-10, CIFAR-100, and DomainNet. They simulated harsh label-skew conditions using varying Dirichlet distribution parameters across ten clients and evaluated feature-skew across six distinct visual domains. FedAF was benchmarked against traditional aggregate-then-adapt methods, such as FedAvg, FedProx, FedDyn, and MOON, as well as an existing aggregation-free condensation method, FedDM.

The findings show that FedAF consistently outperforms existing methods in both final model accuracy and training efficiency. In label-skew settings, FedAF improved accuracy by up to 25.44% on CIFAR-10, 17.91% on CIFAR-100, and 31.03% on Fashion-MNIST compared to standard baseline approaches. It also improved accuracy by up to 4.87% over existing aggregation-free methods. In terms of training speed, FedAF accelerated convergence by up to 80%, reaching target accuracy thresholds in two to three rounds where prior methods required ten to fifteen rounds. Furthermore, under domain-level feature skew, FedAF achieved the highest overall average accuracy across six domains, ranking first or second in every individual domain.

These results demonstrate that shifting to a data condensation paradigm allows organizations to train robust global models across non-uniform data environments without sacrificing communication efficiency or privacy. By sharing condensed synthetic data rather than vulnerable gradient updates, organizations mitigate model divergence while maintaining resistance to data inference attacks. The framework also retains historical knowledge better between training rounds, delivering predictable model stabilization.

For engineering and research teams looking to deploy this framework, the article recommends configuring synthetic dataset sizes to around 50 images per class, as this provides an optimal balance among accuracy, communication overhead, and privacy preservation before returns diminish. Teams should also implement parameter re-sampling around a 0.9 retention ratio to prevent overfitting during local data synthesis.

Confidence in these findings is high for standard image classification benchmarks across simulated client splits. However, decision-makers should note that evaluations were conducted in simulated research settings with ten or fewer clients. Further testing on production-scale edge hardware, real-world network latency, and non-image modalities is advised before wide operational deployment.

arXiv: 2404.18962
Cover for An Aggregation-Free Federated Learning for Tackling Data Heterogeneity

Abstract

The performance of Federated Learning (FL) hinges on the effectiveness of utilizing knowledge from distributed datasets. Traditional FL methods adopt an aggregate-then-adapt framework, where clients update local models based on a global model aggregated by the server from the previous training round. This process can cause client drift, especially with significant cross-client data heterogeneity, impacting model performance and convergence of the FL algorithm. To address these challenges, we introduce FedAF, a novel aggregation-free FL algorithm. In this framework, clients collaboratively learn condensed data by leveraging peer knowledge, the server subsequently trains the global model using the condensed data and soft labels received from the clients. FedAF inherently avoids the issue of client drift, enhances the quality of condensed data amid notable data heterogeneity, and improves the global model performance. Extensive numerical studies on several popular benchmark datasets show FedAF surpasses various state-of-the-art FL algorithms in handling label-skew and feature-skew data heterogeneity, leading to superior global model accuracy and faster convergence.

Table of Contents

  • 1. Introduction
  • 2. Background and Related Works
  • 3. Notations and Preliminaries
  • 4. The Proposed Method
  • 5. Experiments
  • 5.1. Results for Label-skew Data Heterogeneity
  • 5.2. Result for Feature-skew Data Heterogeneity
  • 5.3. Performance Analysis of FedAF
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — FedAF aggregation-free federated learning workflow

    algorithm

    FedAF is an aggregation-free federated learning algorithm designed for strongly heterogeneous client data. It does not upload or average client model parameters. Instead, each client learns a small class-balanced synthetic dataset, sends the condensed data and class-wise soft labels to the server, and the server trains the global model directly on these shared synthetic samples.

    Input: Private client datasets D1,…,DKD_1,\ldots,D_K; initial global model ww; number of classes CC; communication rounds TT; condensed images per class IPCIPC; local condensation steps; global training epochs; coefficients λloc\lambda_{loc}, λglob\lambda_{glob}, temperature τ\tau, and resampling coefficient γ\gamma.
    Initialize each client condensed dataset SkS_k with IPCIPC samples per class.
    for each communication round from 11 to TT do
        Server broadcasts the current global model ww and the current global class-wise mean logits to all clients.
        Each client interpolates ww with randomly sampled model parameters using w←γw+(1−γ)w~w \leftarrow \gamma w + (1-\gamma)\widetilde{w}.
        Each client computes class-wise mean logits from its real data and class-wise mean logits from its condensed data.
        Each client updates SkS_k for the prescribed number of condensation steps using feature-space distribution matching plus a sliced-Wasserstein logit-space regularizer against the global class-wise mean logits.
        Each client computes class-wise soft labels from its real-data mean logits and sends SkS_k, the soft labels, and the real-data mean logits to the server.
        Server averages the client mean logits and soft labels class by class.
        Server forms class-wise soft labels from the received condensed data and updates ww using cross-entropy on all received condensed data plus a symmetric KL-divergence matching loss between real-data soft labels and condensed-data soft labels.
    end for
    Output: The final global model ww.

    The reported default configuration used 50 condensed images per class, 1,000 local condensation steps, 500 global-model training epochs, and γ=0.9\gamma=0.9. The global model is therefore updated from client-condensed data and transferred soft knowledge rather than from aggregated local model updates, which removes the local fine-tuning step responsible for client drift in the conventional aggregate-then-adapt framework.

  2. Knowl 2 — Collaborative data condensation with cross-client logit matching

    model/method

    For client k∈{1,…,K}k\in\{1,\ldots,K\}, let DkD_k be its private real dataset and SkS_k its synthetic condensed dataset. Let CC be the number of classes, hw(⋅)h_w(\cdot) the feature extractor formed by all layers before the classifier of the current global model ww, and fw(⋅)f_w(\cdot) the corresponding pre-softmax logit function. For class cc, let Nk,cN_{k,c} and Mk,cM_{k,c} be the numbers of real and condensed samples, respectively. FedAF first matches real and condensed feature means through distribution matching:

    μk,creal=1Nk,c∑j=1Nk,chw(xk,cj),μk,csyn=1Mk,c∑j=1Mk,chw(x~k,cj),\mu^{real}_{k,c}=\frac{1}{N_{k,c}}\sum_{j=1}^{N_{k,c}}h_w(x^j_{k,c}),\qquad \mu^{syn}_{k,c}=\frac{1}{M_{k,c}}\sum_{j=1}^{M_{k,c}}h_w(\widetilde{x}^j_{k,c}), LDM(Sk,Dk)=∑c=0C−1∥μk,creal−μk,csyn∥22.\mathcal{L}_{DM}(S_k,D_k)=\sum_{c=0}^{C-1}\left\|\mu^{real}_{k,c}-\mu^{syn}_{k,c}\right\|_2^2.

    Here xk,cj∈Dkx^j_{k,c}\in D_k and x~k,cj∈Sk\widetilde{x}^j_{k,c}\in S_k are real and synthetic samples of class cc. The client also computes mean logits

    vk,c=1Nk,c∑j=1Nk,cfw(xk,cj),uk,c=1Mk,c∑j=1Mk,cfw(x~k,cj).\mathbf{v}_{k,c}=\frac{1}{N_{k,c}}\sum_{j=1}^{N_{k,c}}f_w(x^j_{k,c}),\qquad \mathbf{u}_{k,c}=\frac{1}{M_{k,c}}\sum_{j=1}^{M_{k,c}}f_w(\widetilde{x}^j_{k,c}).

    The server averages the real-data mean logits across clients, vc=K−1∑k=1Kvk,c\mathbf{v}_c=K^{-1}\sum_{k=1}^{K}\mathbf{v}_{k,c}, and broadcasts the resulting global class-wise logits to the clients. Each client then learns its condensed data with

    Lloc(Sk,Dk)=LDM(Sk,Dk)+λloc∑c=0C−1SWD⁡(uk,c,vc),\mathcal{L}_{loc}(S_k,D_k)=\mathcal{L}_{DM}(S_k,D_k)+\lambda_{loc}\sum_{c=0}^{C-1}\operatorname{SWD}(\mathbf{u}_{k,c},\mathbf{v}_c),

    where λloc≥0\lambda_{loc}\geq 0 is the local regularization weight and SWD⁡\operatorname{SWD} is the Sliced Wasserstein Distance used as the logit-space distance. The additional term makes each client’s condensed data reflect the knowledge distribution represented by the other clients, rather than matching only the client’s potentially biased local distribution.

  3. Knowl 3 — Local-global knowledge matching for global-model training

    model/method

    FedAF supplements the information loss caused by data condensation with soft labels computed from the original client data. For client kk and class cc, let vk,c\mathbf{v}_{k,c} be the mean pre-softmax logit of the real samples in class cc, and let σw(z,τ)\sigma_w(\mathbf{z},\tau) denote the final softmax of global model ww applied with temperature τ\tau. The client sends

    rk,c=σw(vk,c,τ)\mathbf{r}_{k,c}=\sigma_w(\mathbf{v}_{k,c},\tau)

    and the server averages these local soft labels class by class:

    rc=1K∑k=1Krk,c.\mathbf{r}_c=\frac{1}{K}\sum_{k=1}^{K}\mathbf{r}_{k,c}.

    Let ScS_c be the union of the received condensed samples belonging to class cc. The server computes a condensed-data soft label from the mean condensed logit,

    tc=σw(1∣Sc∣∑j=1∣Sc∣fw(x~cj),τ),\mathbf{t}_c=\sigma_w\left(\frac{1}{|S_c|}\sum_{j=1}^{|S_c|}f_w(\widetilde{x}^{j}_c),\tau\right),

    where fwf_w is the pre-softmax logit function and x~cj∈Sc\widetilde{x}^{j}_c\in S_c. With R=[r0,…,rC−1]\mathcal{R}=[\mathbf{r}_0,\ldots,\mathbf{r}_{C-1}] and T=[t0,…,tC−1]\mathcal{T}=[\mathbf{t}_0,\ldots,\mathbf{t}_{C-1}], the server trains the global model using

    Lglob(w,S)=LCE(w,S)+λglobLLGKM(w,S),\mathcal{L}_{glob}(w,S)=\mathcal{L}_{CE}(w,S)+\lambda_{glob}\mathcal{L}_{LGKM}(w,S), LLGKM(w,S)=12[DKL(R∥T)+DKL(T∥R)].\mathcal{L}_{LGKM}(w,S)=\frac{1}{2}\left[D_{KL}(\mathcal{R}\|\mathcal{T})+D_{KL}(\mathcal{T}\|\mathcal{R})\right].

    Here LCE\mathcal{L}_{CE} is cross-entropy on the received condensed data, DKLD_{KL} is Kullback–Leibler divergence, and λglob≥0\lambda_{glob}\geq0 controls the local-global matching term. The symmetric KL penalty transfers knowledge retained in the original client data to the server and stabilizes global training when the condensed samples are imperfect.

  4. Knowl 4 — Global-model resampling during condensation

    model/method

    During each client’s condensed-data optimization, FedAF perturbs the downloaded global model by interpolating its parameters with randomly sampled model parameters:

    w←γw+(1−γ)w~,w\leftarrow \gamma w+(1-\gamma)\widetilde{w},

    where ww is the current global-model parameter vector, w~\widetilde{w} is a randomly sampled parameter vector, and γ∈[0,1]\gamma\in[0,1] controls how much previously learned global knowledge is retained. The interpolation is used as the condensation backbone rather than as a replacement for the server’s final global model. In the reported experiments, γ=0.9\gamma=0.9 was used by default.

    On CIFAR10 with label-skew parameter α=0.1\alpha=0.1 over ten communication rounds, the resulting accuracies were 61.30 for γ=0.2\gamma=0.2, 63.52 for γ=0.5\gamma=0.5, 66.15 for γ=0.8\gamma=0.8, 66.97 for γ=0.9\gamma=0.9, and 64.92 for γ=1.0\gamma=1.0. Thus, retaining most of the global model knowledge while adding a small random perturbation performed better than either a larger perturbation or no perturbation. Using purely random parameters, γ=0\gamma=0, destabilized condensation and was excluded from further comparisons.

  5. Knowl 5 — Evaluation protocol under label and feature heterogeneity

    experimental setup

    FedAF was evaluated against FedAvg, FedProx, FedBN, MOON, FedDyn, FedGen, and the aggregation-free FedDM method. Label-skew experiments used Fashion-MNIST, CIFAR10, and CIFAR100 with K=10K=10 clients. Each training split was partitioned among clients with a Dirichlet distribution using α∈{0.02,0.05,0.1}\alpha\in\{0.02,0.05,0.1\}; smaller α\alpha represents stronger non-IID heterogeneity. The reported label-skew metric was the highest global-model accuracy within 20 communication rounds.

    Aggregate-then-adapt baselines used 10 local epochs, learning rate 0.01, and batch size 64. FedDM and FedAF used batch size 256, 1,000 local condensation steps, 50 images per class, and 500 global training epochs with batch size 256 and learning rate 0.001. Condensed classes were initialized from averages of randomly sampled local real images. FedAF used image learning rates 1.0 on CIFAR10, 0.1 on CIFAR100, and 0.2 on Fashion-MNIST; its global-model resampling coefficient was γ=0.9\gamma=0.9.

    Feature-skew experiments used a ten-class subset of DomainNet containing the six domains Clipart, Infograph, Painting, Quickdraw, Real, and Sketch. Six clients were assigned one distinct domain each, and all methods were trained for 10 communication rounds. The remaining FedAF and FedDM settings matched the label-skew experiments, with image learning rate 1.0.

  6. Knowl 6 — Label-skew accuracy results across three datasets

    data/table

    The following table reports the highest global-model accuracy within 20 communication rounds for three datasets and three Dirichlet heterogeneity levels. Each entry is the paper’s reported accuracy with its accompanying variability. FedAF is consistently the best method in all nine dataset–heterogeneity settings, and its advantage over FedDM is largest or nearly largest under the strongest heterogeneity.

    Method α=0.02\alpha=0.02 α=0.05\alpha=0.05 α=0.1\alpha=0.1
    FMNIST CIFAR10 CIFAR100 FMNIST CIFAR10 CIFAR100 FMNIST CIFAR10 CIFAR100
    FedAvg 56.50±\pm5.55 39.71±\pm1.15 30.80±\pm2.20 69.14±\pm5.84 46.51±\pm3.07 33.37±\pm0.75 82.19±\pm5.67 56.15±\pm4.62 39.97±\pm1.53
    FedProx 60.38±\pm5.00 36.46±\pm5.39 30.82±\pm0.80 69.33±\pm4.12 45.83±\pm2.23 36.61±\pm1.44 81.56±\pm4.52 58.54±\pm1.87 40.45±\pm1.53
    FedBN 58.26±\pm4.28 36.53±\pm2.52 29.73±\pm1.73 72.91±\pm4.69 45.13±\pm2.18 33.73±\pm2.15 77.33±\pm3.07 57.67±\pm3.21 39.84±\pm0.20
    MOON 51.33±\pm7.00 33.32±\pm1.13 33.41±\pm0.70 71.41±\pm4.08 47.41±\pm4.59 37.90±\pm0.80 81.61±\pm2.68 57.62±\pm4.99 40.24±\pm0.68
    FedDyn 69.79±\pm5.04 45.73±\pm3.98 35.01±\pm2.07 75.19±\pm5.49 57.68±\pm1.84 39.10±\pm0.34 84.73±\pm2.74 59.97±\pm2.20 41.81±\pm1.46
    FedGen 61.44±\pm2.07 36.61±\pm1.06 29.20±\pm2.09 75.48±\pm1.83 42.72±\pm2.11 33.56±\pm3.91 82.29±\pm2.53 58.17±\pm2.84 40.23±\pm1.06
    FedDM 85.36±\pm0.96 60.28±\pm0.82 44.15±\pm0.30 86.08±\pm0.68 62.97±\pm0.96 46.27±\pm0.98 86.65±\pm0.31 64.88±\pm0.35 47.05±\pm0.13
    FedAF 87.53±\pm0.32 65.15±\pm0.86 48.71±\pm0.33 87.29±\pm0.23 67.50±\pm0.76 49.49±\pm0.33 87.91±\pm0.41 69.11±\pm0.86 50.61±\pm0.26

    Relative to FedAvg, the largest reported FedAF gains were 25.44 percentage points on CIFAR10, 17.91 on CIFAR100, and 31.03 on Fashion-MNIST. Relative to FedDM, FedAF improved by as much as 4.87, 4.56, and 2.17 percentage points on CIFAR10, CIFAR100, and Fashion-MNIST, respectively.

  7. Knowl 7 — Faster convergence under severe label skew

    empirical result

    FedAF reached high global-model accuracy substantially earlier than both aggregate-then-adapt baselines and the aggregation-free FedDM baseline. Under the strongest label heterogeneity tested, CIFAR10 with α=0.02\alpha=0.02, FedAF reached the highest accuracy achieved by the aggregate-then-adapt methods within two communication rounds. FedDM required 15 rounds to reach a mean accuracy of 60%, whereas FedAF reached the same 60% level in three rounds, corresponding to the paper’s reported 80% improvement in convergence speed.

    The learning curves across CIFAR10, CIFAR100, and Fashion-MNIST showed the same qualitative pattern: FedAF rose rapidly during the first few rounds and then maintained the best or near-best accuracy, with the speed advantage especially visible on the harder datasets and under stronger heterogeneity.

  8. Knowl 8 — Feature-skew results on DomainNet

    data/table

    The table compares global-model accuracy after 10 communication rounds when six clients each hold one of six DomainNet domains. The domains are Clipart (C), Infograph (I), Painting (P), Quickdraw (Q), Real (R), and Sketch (S); Avg is the mean accuracy across the six domains. FedAF obtains the highest average accuracy and is either the best or second-best method in every individual domain.

    Method C I P Q R S Avg
    FedAvg 43.03 40.76 59.16 39.60 41.03 28.46 42.01
    FedProx 44.81 43.76 60.22 38.13 41.55 29.18 42.94
    FedBN 46.07 34.27 52.01 43.10 47.33 29.72 42.08
    MOON 48.80 37.97 56.26 48.07 42.02 29.72 43.81
    FedDyn 48.04 60.03 67.46 37.73 41.77 32.67 47.95
    FedGen 42.77 37.88 54.37 37.33 42.86 25.69 40.15
    FedDM 52.28 41.38 60.58 62.37 52.45 46.69 52.62
    FedAF 51.2 47.05 62.53 64.6 52.64 50.06 54.68

    FedAF’s mean accuracy of 54.68 exceeds FedDM’s 52.62 and all aggregate-then-adapt baselines. It also reached in two rounds the accuracy level that FedDM needed all ten rounds to reach, which the paper reports as an 80% convergence-speed improvement under feature skew.

  9. Knowl 9 — Ablation of collaborative condensation and local-global matching

    data/table

    The ablation isolates the two main FedAF additions on CIFAR10. The full method uses both collaborative data condensation (CDC) and local-global knowledge matching (LGKM). The variant without CDC retains LGKM but removes cross-client logit regularization during condensation; the variant without LGKM retains CDC but removes the server’s soft-label matching loss. FedDM supplies a comparison in which neither mechanism is present.

    Configuration α=0.02\alpha=0.02 α=0.05\alpha=0.05 α=0.1\alpha=0.1
    FedAF 65.15±\pm0.86 67.50±\pm0.76 69.11±\pm0.86
    w/o CDC 64.16±\pm0.83 65.88±\pm0.93 67.90±\pm0.53
    w/o LGKM 64.12±\pm0.85 66.27±\pm1.31 68.14±\pm0.81
    FedDM 60.28±\pm0.82 62.97±\pm0.96 64.88±\pm0.35

    Removing either mechanism lowers accuracy at every tested heterogeneity level, while retaining either one still improves over FedDM. The full method is best in all three settings, supporting the paper’s conclusion that higher-quality cross-client condensation and preservation of original-data knowledge in server training provide complementary gains.

  10. Knowl 10 — Effect of the number of condensed images per class

    data/table

    On CIFAR10, FedAF was evaluated with different numbers of condensed images per class (IPC) under three Dirichlet heterogeneity levels. Increasing IPC generally improves the global-model accuracy, but the improvement becomes less pronounced beyond IPC 50.

    Configuration α=0.02\alpha=0.02 α=0.05\alpha=0.05 α=0.1\alpha=0.1
    IPC=10 53.39±\pm2.09 55.33±\pm0.81 56.15±\pm0.42
    IPC=20 58.56±\pm0.55 60.89±\pm0.11 61.79±\pm0.59
    IPC=50 65.15±\pm0.86 67.50±\pm0.76 69.11±\pm0.86
    IPC=80 67.94±\pm1.18 70.07±\pm0.45 70.72±\pm0.37
    IPC=100 69.14±\pm0.56 71.27±\pm0.58 71.66±\pm0.37

    The paper uses IPC 50 as a practical compromise: increasing IPC raises communication cost approximately linearly, whereas the accuracy gains after IPC 50 are comparatively small. The authors further state that fewer condensed samples imply a lower condensed-to-original-data ratio and therefore better privacy retention according to the cited condensation-privacy analysis.

Coverage note — No substantial contributed material was omitted; the convergence plots are represented by their reported numerical comparisons, while background and related-work material was excluded.

References

  1. 1.Durmus Alp Emre Acar, Yue Zhao, Ramon Matas Navarro, Matthew Mattina, Paul N Whatmough, and Venkatesh Saligrama. Federated learning based on dynamic regularization. arXiv preprint arXiv:2111.04263, 2021. 1, 2, 5
  2. 2.Martin Arjovsky, Soumith Chintala, and Leon Bottou. Wasserstein generative adversarial networks. In Proceedings of the 34th International Conference on Machine Learning, pages 214–223. PMLR, 2017. 4
  3. 3.Nicolas Bonneel, Julien Rabin, Gabriel Peyre, and Hanspeter Pfister. Sliced and radon wasserstein barycenters of measures. Journal of Mathematical Imaging and Vision, 51:22–45, 2015. 4
  4. 4.George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A Efros, and Jun-Yan Zhu. Dataset distillation by matching training trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4750–4759, 2022. 2, 3
  5. 5.Hong-You Chen and Wei-Lun Chao. FedBE: Making bayesian model ensemble applicable to federated learning. In International Conference on Learning Representations, 2021. 1, 2
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. IEEE, 2009. 1
  7. 7.Ishan Deshpande, Ziyu Zhang, and Alexander G Schwing. Generative modeling using the sliced wasserstein distance. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3483–3491, 2018. 4
  8. 8.Tian Dong, Bo Zhao, and Lingjuan Lyu. Privacy for free: How does dataset condensation help privacy? In International Conference on Machine Learning, pages 5378–5396. PMLR, 2022. 3, 7
  9. 9.Chun-Mei Feng, Bangjun Li, Xinxing Xu, Yong Liu, Huazhu Fu, and Wangmeng Zuo. Learning federated visual prompt in null space for mri reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8064–8073, 2023. 2
  10. 10.Liang Gao, Huazhu Fu, Li Li, Yingwen Chen, Ming Xu, and Cheng-Zhong Xu. Feddc: Federated learning with non-iid data via local drift decoupling and correction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10112–10121, 2022. 1, 2
  11. 11.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014. 3
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 1
  13. 13.Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. Measuring the effects of non-identical data distribution for federated visual classification. arXiv preprint arXiv:1909.06335, 2019. 1, 2
  14. 14.Wenke Huang, Mang Ye, and Bo Du. Learn from others and be yourself in heterogeneous federated learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10143–10153, 2022. 2
  15. 15.Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian Stich, and Ananda Theertha Suresh. Scaffold: Stochastic controlled averaging for federated learning. In International conference on machine learning, pages 5132–5143. PMLR, 2020. 1, 2
  16. 16.Soheil Kolouri, Yang Zou, and Gustavo K Rohde. Sliced wasserstein kernels for probability distributions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5258–5267, 2016. 4
  17. 17.Soheil Kolouri, Phillip E. Pope, Charles E. Martin, and Gustavo K. Rohde. Sliced wasserstein auto-encoders. In International Conference on Learning Representations, 2019. 4
  18. 18.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. Technical report, Toronto, ON, Canada, 2009. 5
  19. 19.Qinbin Li, Bingsheng He, and Dawn Song. Model-contrastive federated learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10713–10722, 2021. 1, 2, 5
  20. 20.Qinbin Li, Yiqun Diao, Quan Chen, and Bingsheng He. Federated learning on non-iid data silos: An experimental study. In 2022 IEEE 38th International Conference on Data Engineering (ICDE), pages 965–978. IEEE, 2022. 1
  21. 21.Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. Proceedings of Machine learning and systems, 2:429–450, 2020. 2, 5
  22. 22.Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. On the convergence of fedavg on non-iid data. In International Conference on Learning Representations, 2020. 1
  23. 23.Xiaoxiao Li, Meirui JIANG, Xiaofei Zhang, Michael Kamp, and Qi Dou. FedBN: Federated learning on non-IID features via local batch normalization. In International Conference on Learning Representations, 2021. 2, 5, 7, 3
  24. 24.Tao Lin, Lingjing Kong, Sebastian U Stich, and Martin Jaggi. Ensemble distillation for robust model fusion in federated learning. Advances in Neural Information Processing Systems, 33:2351–2363, 2020. 1, 2
  25. 25.Ping Liu, Xin Yu, and Joey Tianyi Zhou. Meta knowledge condensation for federated learning. In The Eleventh International Conference on Learning Representations, 2023. 2, 3
  26. 26.Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Artificial intelligence and statistics, pages 1273–1282. PMLR, 2017. 1, 2, 5
  27. 27.Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for multi-source domain adaptation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1406–1415, 2019. 5
  28. 28.Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9 (86):2579–2605, 2008. 1
  29. 29.Jianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi, and H Vincent Poor. Tackling the objective inconsistency problem in heterogeneous federated optimization. Advances in neural information processing systems, 33:7611–7623, 2020. 1, 2
  30. 30.Jianyu Wang, Zachary Charles, Zheng Xu, Gauri Joshi, H Brendan McMahan, Maruan Al-Shedivat, Galen Andrew, Salman Avestimehr, Katharine Daly, Deepesh Data, et al. A field guide to federated optimization. arXiv preprint arXiv:2107.06917, 2021. 1
  31. 31.Kai Wang, Bo Zhao, Xiangyu Peng, Zheng Zhu, Shuo Yang, Shuo Wang, Guan Huang, Hakan Bilen, Xinchao Wang, and Yang You. Cafe: Learning to condense dataset by aligning features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12196–12205, 2022. 2, 3
  32. 32.Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros. Dataset distillation. arXiv preprint arXiv:1811.10959, 2018. 2, 3
  33. 33.Han Xiao, Kashif Rasul, and Roland Vollgraf. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017. 5
  34. 34.Yuanhao Xiong, Ruochen Wang, Minhao Cheng, Felix Yu, and Cho-Jui Hsieh. FedDM: Iterative distribution matching for communication-efficient federated learning. In Workshop on Federated Learning: Recent Advances and New Challenges (in Conjunction with NeurIPS 2022), 2022. 2, 3, 5
  35. 35.Lin Zhang, Li Shen, Liang Ding, Dacheng Tao, and Ling-Yu Duan. Fine-tuning global model via data-free knowledge distillation for non-iid federated learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10174–10183, 2022. 1
  36. 36.Bo Zhao and Hakan Bilen. Dataset condensation with differentiable siamese augmentation. In International Conference on Machine Learning, pages 12674–12685. PMLR, 2021. 2, 3
  37. 37.Bo Zhao and Hakan Bilen. Dataset condensation with distribution matching. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 6514–6523, 2023. 2, 3, 1
  38. 38.Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen. Dataset condensation with gradient matching. In International Conference on Learning Representations, 2021. 2
  39. 39.Yue Zhao, Meng Li, Liangzhen Lai, Naveen Suda, Damon Civin, and Vikas Chandra. Federated learning with non-iid data. arXiv preprint arXiv:1806.00582, 2018. 1
  40. 40.Hangyu Zhu, Jinjin Xu, Shiqing Liu, and Yaochu Jin. Federated learning on non-iid data: A survey. Neurocomputing, 465:371–390, 2021. 1
  41. 41.Zhuangdi Zhu, Junyuan Hong, and Jiayu Zhou. Data-free knowledge distillation for heterogeneous federated learning. In International conference on machine learning, pages 12878–12889. PMLR, 2021. 1, 2, 5

Citation

MLA
Wang, Y., et al. “An Aggregation-Free Federated Learning for Tackling Data Heterogeneity”. arXiv, 2024, http://arxiv.org/abs/2404.18962v1.
APA
Wang, Y., Fu, H., Kanagavelu, R., Wei, Q., Liu, Y., & Goh, R. S. M. (2024). An Aggregation-Free Federated Learning for Tackling Data Heterogeneity. arXiv. http://arxiv.org/abs/2404.18962v1
Chicago
Wang, Y., H. Fu, R. Kanagavelu, Q. Wei, Y. Liu, and R. S. M. Goh. 2024. “An Aggregation-Free Federated Learning for Tackling Data Heterogeneity”. arXiv. http://arxiv.org/abs/2404.18962v1.
Harvard
Wang, Y. et al. (2024) “An Aggregation-Free Federated Learning for Tackling Data Heterogeneity”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.18962v1.
Vancouver
1. Wang Y, Fu H, Kanagavelu R, Wei Q, Liu Y, Goh RSM (2024) An Aggregation-Free Federated Learning for Tackling Data Heterogeneity. arXiv

BibTeX

@article{wang2024aggregation,
  title = {An Aggregation-Free Federated Learning for Tackling Data Heterogeneity},
  author = {Wang, Yuan and Fu, Huazhu and Kanagavelu, Renuga and Wei, Qingsong and Liu, Yong and Goh, Rick Siow Mong},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.18962v1},
  eprint = {2404.18962}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE