Ditto: Fair and Robust Federated Learning Through Personalization

Tian LiShengyuan HuAhmad BeiramiVirginia Smith

article2020ICML1,491 citationsBest Paper Award at ICLR 2021 Secure ML Workshop

Introduces Ditto, a scalable personalized federated learning framework that resolves the competing constraints of device fairness and defense against data and model poisoning attacks in statistically heterogeneous networks.

Listen

Federated learning enables multiple remote devices to train a shared machine learning model collaboratively without centralizing their local data. In real-world enterprise deployments, these networks face significant statistical diversity because user data varies widely across devices. System operators must simultaneously satisfy multiple operational requirements: high model accuracy, fairness (ensuring uniform performance across all participating devices), and robustness (defending against malicious data or model poisoning attacks). Existing methods address fairness and defense in isolation, but these objectives directly conflict. Prior fair approaches overfit to corrupted participants by giving them extra weight, while traditional robust defenses often filter out rare, benign data distributions, causing severe performance disparities.

The article demonstrates how a personalized federated learning framework can inherently resolve the conflict between fairness and robustness. It proposes and evaluates Ditto, a multi-task learning approach that optimizes a local model for each device while regularizing it to stay close to an aggregate global model.

To evaluate the approach, the authors developed a mathematical analysis for linear models and carried out extensive empirical experiments across standard vision and language federated learning benchmarks, including FEMNIST, Fashion-MNIST, CelebA, and StackOverflow. The evaluation covered both convex and non-convex models under three common attack categories: label poisoning, random updates, and model replacement. The authors compared Ditto against state-of-the-art fair algorithms, robust aggregation baselines, and recent personalization techniques.

The investigation produced several key findings. First, personalization inherently improves resilience to attacks: Ditto achieved an average absolute test accuracy improvement of about 6 percentage points over the strongest robust defense baseline across all datasets and attack scenarios. Second, Ditto enhanced network fairness, reducing the variance in test accuracy across devices by approximately 10% while raising absolute test accuracy by 5 percentage points relative to state-of-the-art fair baselines on clean data. Third, under aggressive model poisoning where standard fair and global methods suffered catastrophic performance drops or failed to converge, Ditto maintained stable, high performance. Fourth, Ditto matched or exceeded alternative personalization methods without adding computational complexity, and it delivered further performance gains when combined directly with standard robust aggregation defenses.

These findings indicate that addressing data diversity through lightweight personalization allows organizations to resolve the trade-off between security and fair service delivery. In practical terms, Ditto protects system accuracy against adversarial devices without sacrificing performance on benign users who possess atypical data patterns. Because it shares only the standard updates required by base federated learning algorithms, it also avoids communication overhead and preserves existing privacy protections.

Organizations deploying federated learning systems should consider adopting personalized objectives like Ditto as a modular enhancement to their current pipelines. Engineering teams should implement local hyperparameter selection using on-device validation data, using smaller regularization values when attacks are suspected to allow benign devices to decouple from a compromised global model. When operating in high-threat environments, teams should combine personalized objectives with server-side robust aggregators for defense-in-depth.

These conclusions are supported by theoretical proofs on linear models and consistent empirical outcomes across diverse benchmarks. However, leaders should note that the evaluation assumes devices can reasonably distinguish attack intensity to tune regularization parameters, and the study focused primarily on training-time data and model poisoning. Further validation is recommended before deployment against other threat models, such as targeted backdoor attacks.

Cover for Ditto: Fair and Robust Federated Learning Through Personalization

Abstract

Fairness and robustness are two important concerns for federated learning systems. In this work, we identify that robustness to data and model poisoning attacks and fairness, measured as the uniformity of performance across devices, are competing constraints in statistically heterogeneous networks. To address these constraints, we propose employing a simple, general framework for personalized federated learning, Ditto, that can inherently provide fairness and robustness benefits, and develop a scalable solver for it. Theoretically, we analyze the ability of Ditto to achieve fairness and robustness simultaneously on a class of linear problems. Empirically, across a suite of federated datasets, we show that Ditto not only achieves competitive performance relative to recent personalization methods, but also enables more accurate, robust, and fair models relative to state-of-the-art fair or robust baselines.

Table of Contents

  • 1 Introduction
  • 2 Background & Related Work
  • 3 Ditto: Global-Regularized Federated Multi-Task Learning
  • 3.1 Ditto Objective
  • 3.2 Ditto Solver
  • 3.3 Analyzing the Fairness/Robustness Benefits of in Simplified Settings
  • 4 Experiments
  • 4.1 Robustness of Ditto
  • 4.2 Fairness of Ditto
  • 4.3 Addressing Competing Constraints
  • 4.4 Additional Properties of Ditto
  • 5 Conclusion and Future Work
  • References
  • A Analysis of the Federated Multi-Task Learning Objective Ditto
  • A.1 Properties of Ditto for Strongly Convex Functions
  • A.2 Federated Linear Regression
  • A.2.1 No Adversaries: Ditto for Accuracy and Fairness
  • A.2.2 With Adversaries: Ditto for Accuracy, Fairness, and Robustness
  • A.3 The Case of Federated Point Estimation
  • B Algorithm and Convergence Analysis
  • C Experimental Details
  • C.1 Datasets and Models
  • C.2 Personalization Baselines
  • D Additional and Complete Experiment Results
  • D.1 Comparing with Finetuning
  • D.2 Tuning λ\lambda
  • D.3 Ditto Augmented with Robust Baselines
  • D.4 Ditto Complete Results

Knowls

  1. Knowl 1 — Ditto Objective for Personalized Federated Learning

    model/method

    Ditto is a multi-task learning objective designed to learn personalized device-specific models while regularizing them toward an optimal global model. For a federated network of KK devices where each device k∈[K]k \in [K] has a local data distribution Dk\mathcal{D}_k and local loss objective Fk(vk):=Exk∼Dk[fk(vk;xk)]F_k(v_k) := \mathbb{E}_{x_k \sim \mathcal{D}_k}[f_k(v_k; x_k)], Ditto formulates the optimization for device kk as a bi-level problem:

    min⁡vkhk(vk;w∗):=Fk(vk)+λ2∥vk−w∗∥2s.t.w∗∈arg⁡min⁡wG(F1(w),…,FK(w))\min_{v_k} h_k(v_k; w^*) := F_k(v_k) + \frac{\lambda}{2} \|v_k - w^*\|^2 \quad \text{s.t.} \quad w^* \in \arg\min_w G(F_1(w), \dots, F_K(w))

    Here, w∗w^* is the optimal global model obtained by minimizing a network-wide aggregation function G(⋅)G(\cdot) (such as the standard federated averaging objective ∑k=1KpkFk(w)\sum_{k=1}^K p_k F_k(w) with weights pk≥0p_k \ge 0 and ∑k=1Kpk=1\sum_{k=1}^K p_k = 1), vkv_k is the personalized model for device kk, and λ≥0\lambda \ge 0 is a scalar hyperparameter that controls the trade-off between local and global learning. When λ=0\lambda = 0, the objective decouples into training independent local models on each device; as λ→∞\lambda \to \infty, the personalized models recover the optimal single global model w∗w^*.

  2. Knowl 2 — Alternating Optimization Algorithm for Ditto

    algorithm

    To solve the bi-level Ditto objective across a distributed network, optimization alternates between updating the global model ww across participating devices and solving the global-regularized local objective hk(vk;wt)h_k(v_k; w^t) on local devices.

    Input: Total devices KK, communication rounds TT, local personalization steps ss, global local steps rr, regularization parameter λ\lambda, global learning rate ηg\eta_g, local learning rate ηl\eta_l, initial global model w0w^0, initial personalized models {vk0}k=1K\{v_k^0\}_{k=1}^K, device weights {pk}k=1K\{p_k\}_{k=1}^K
    Output: Personalized models {vk}k=1K\{v_k\}_{k=1}^K, global model wTw^T
    for t=0,…,T−1t = 0, \dots, T-1 do
        Server selects a random subset of devices St⊆[K]S_t \subseteq [K]
        Server sends global model wtw^t to all devices k∈Stk \in S_t
        for device k∈Stk \in S_t in parallel do
            Initialize local global model copy: wkt←wtw_k^t \leftarrow w^t
            for step =1,…,r= 1, \dots, r do
                wkt←wkt−ηg∇Fk(wkt)w_k^t \leftarrow w_k^t - \eta_g \nabla F_k(w_k^t)
            for step =1,…,s= 1, \dots, s do
                vk←vk−ηl(∇Fk(vk)+λ(vk−wt))v_k \leftarrow v_k - \eta_l (\nabla F_k(v_k) + \lambda (v_k - w^t))
            Compute update: Δkt←wkt−wt\Delta_k^t \leftarrow w_k^t - w^t
            Send Δkt\Delta_k^t to server
        Server aggregates updates: wt+1←wt+1∣St∣∑k∈StΔktw^{t+1} \leftarrow w^t + \frac{1}{|S_t|} \sum_{k \in S_t} \Delta_k^t
    return {vk}k=1K,wT\{v_k\}_{k=1}^K, w^T

    Devices maintain their personalized model states vkv_k across communication rounds. Communication from client to server consists solely of global update vectors Δkt\Delta_k^t, preserving standard federated learning privacy and bandwidth characteristics.

  3. Knowl 3 — Convergence of Personalized Models Inherited from Global Model Convergence

    theoretical result

    Let FkF_k be μ\mu-strongly convex and smooth for all k∈[K]k \in [K], with uniformly bounded stochastic gradient variance E[∥∇Fk(w,ξ)∥2]≤G12\mathbb{E}[\|\nabla F_k(w, \xi)\|^2] \le G_1^2. Let uk∗:=arg⁡min⁡uFk(u)u_k^* := \arg\min_u F_k(u) be the local minimizer, w∗:=arg⁡min⁡wG(F1(w),…,FK(w))w^* := \arg\min_w G(F_1(w), \dots, F_K(w)) be the global minimizer, and assume bounded statistical heterogeneity across clients: ∥uk∗−w∗∥≤M\|u_k^* - w^*\| \le M for all k∈[K]k \in [K]. Let vk∗:=arg⁡min⁡v(Fk(v)+λ2∥v−w∗∥2)v_k^* := \arg\min_v (F_k(v) + \frac{\lambda}{2}\|v - w^*\|^2).

    If the sequence of global models wtw^t converges to w∗w^* at rate g(t)g(t) such that lim⁡t→∞g(t)=0\lim_{t \to \infty} g(t) = 0 and g(t+1)g(t)≥1−g(t)A\frac{g(t+1)}{g(t)} \ge 1 - \frac{g(t)}{A} for a positive constant AA, then when device kk is selected with probability pkp_k at round tt and uses local step-size η=2g(t)A(μ+λ)pk\eta = \frac{2g(t)}{A(\mu + \lambda)p_k}, there exists a constant C<∞C < \infty such that the personalized model vktv_k^t converges to vk∗v_k^* at rate:

    E[∥vkt−vk∗∥2]≤Cg(t)\mathbb{E}[\|v_k^t - v_k^*\|^2] \le C g(t)

    When the global objective solver is standard FedAvg, which achieves global convergence g(t)=O(1/t)g(t) = O(1/t), the personalized models satisfy E[∥vkt−vk∗∥2]=O(1/t)\mathbb{E}[\|v_k^t - v_k^*\|^2] = O(1/t).

  4. Knowl 4 — Optimality of Ditto for Distributed Linear Regression without Adversaries

    theoretical result

    Consider a Bayesian federated linear regression model across KK devices. Let θ∼Uniform(Rd)\theta \sim \text{Uniform}(\mathbb{R}^d) be drawn from an uninformative prior. Each device k∈[K]k \in [K] has a local ground-truth parameter wk=θ+ζkw_k = \theta + \zeta_k, where ζk∼N(0,τ2Id)\zeta_k \sim \mathcal{N}(0, \tau^2 I_d) are i.i.d., with τ\tau controlling task relatedness. Each device observes nn samples yk=Xkwk+zky_k = X_k w_k + z_k, where zk∼N(0,σ2In)z_k \sim \mathcal{N}(0, \sigma^2 I_n) and XkTXk=βIdX_k^T X_k = \beta I_d for all kk. The empirical local loss is Fk(w)=1n∥Xkw−yk∥2F_k(w) = \frac{1}{n}\|X_k w - y_k\|^2.

    The optimal regularization parameter λ∗\lambda^* that minimizes the expected test mean squared error on device kk is given by:

    λ∗=σ2nτ2\lambda^* = \frac{\sigma^2}{n \tau^2}

    At λ=λ∗\lambda = \lambda^*, the Ditto solution w^k(λ∗)\hat{w}_k(\lambda^*) achieves the Bayes optimal Minimum Mean Square Error (MMSE) estimator among all possible federated estimators. Furthermore, this exact same choice of λ∗\lambda^* simultaneously minimizes the variance of test error across devices, vark∈[K]{∥Xkw^k(λ)−yk∥2}\text{var}_{k \in [K]}\{\|X_k \hat{w}_k(\lambda) - y_k\|^2\}, thereby achieving optimal fairness across the network within Ditto's solution space.

  5. Knowl 5 — Simultaneous Robustness and Fairness Optimality of Ditto under Adversarial Data Poisoning

    theoretical result

    In the Bayesian linear regression framework with KK devices, let KaK_a devices be malicious (adversaries) and Kb=K−Ka≥1K_b = K - K_a \ge 1 be benign. For benign devices k∈[Kb]k \in [K_b], wk∼θ+N(0,τ2Id)w_k \sim \theta + \mathcal{N}(0, \tau^2 I_d); for malicious devices k∈[Ka]k \in [K_a], wk∼θ+N(0,τa2Id)w_k \sim \theta + \mathcal{N}(0, \tau_a^2 I_d) with τa>τ\tau_a > \tau. Each device has nn observations with observation noise variance σ2\sigma^2 and XkTXk=βIdX_k^T X_k = \beta I_d.

    The optimal regularization parameter λa∗\lambda_a^* that minimizes the expected test MSE across benign devices (maximizing Byzantine robustness) is:

    λa∗=σ2n(KKτ2+KaK−1(τa2−τ2))\lambda_a^* = \frac{\sigma^2}{n} \left( \frac{K}{K\tau^2 + \frac{K_a}{K-1}(\tau_a^2 - \tau^2)} \right)

    This same parameter λa∗\lambda_a^* simultaneously minimizes the variance of test error across benign devices (maximizing fairness). At λ=λa∗\lambda = \lambda_a^*, the expected benign test error is dσw,a2d \sigma_{w,a}^2 and the performance variance across benign devices is 2dσw,a42d \sigma_{w,a}^4, where:

    1σw,a2=nσ2+K−1Kτ2+σ2n+KaK−1(τa2−τ2)\frac{1}{\sigma_{w,a}^2} = \frac{n}{\sigma^2} + \frac{K-1}{K\tau^2 + \frac{\sigma^2}{n} + \frac{K_a}{K-1}(\tau_a^2 - \tau^2)}

    As the number of adversarial devices KaK_a or adversary capability τa\tau_a increases, λa∗\lambda_a^* strictly decreases, demonstrating that stronger attacks necessitate greater local personalization (smaller λ\lambda).

  6. Knowl 6 — Byzantine Robustness and Representation Fairness Metrics in Federated Learning

    definition

    Let [K][K] denote the set of all devices, and let [Kb]⊆[K][K_b] \subseteq [K] denote the subset of benign devices. Let Fk(w)F_k(w) denote the test loss (or test error) of model ww on device kk, and let Acck(w)\text{Acc}_k(w) denote the test accuracy.

    1. Byzantine Robustness: Model w1w_1 is defined to be more robust than model w2w_2 against a specific training-time attack if the mean test performance across the benign devices is higher for w1w_1 than for w2w_2 after training under that attack: 1∣Kb∣∑k∈KbAcck(w1)>1∣Kb∣∑k∈KbAcck(w2)\frac{1}{|K_b|} \sum_{k \in K_b} \text{Acc}_k(w_1) > \frac{1}{|K_b|} \sum_{k \in K_b} \text{Acc}_k(w_2)

    2. Representation Fairness: Model w1w_1 is defined to be fairer than model w2w_2 across the network if the distribution of test performance across benign devices is more uniform, measured by a smaller standard deviation of the local test performance: std({Fk(w1)}k∈Kb)<std({Fk(w2)}k∈Kb)\text{std}\left(\{F_k(w_1)\}_{k \in K_b}\right) < \text{std}\left(\{F_k(w_2)\}_{k \in K_b}\right)

  7. Knowl 7 — Tension Between Single-Model Fair FL and Byzantine Robustness

    empirical result

    When training a single global model in statistically heterogeneous networks, objectives designed to promote fairness directly compete with robustness against Byzantine data or model poisoning attacks. Fair FL methods (such as TERM, Agnostic FL, and qq-FFL) upweight devices or samples with large training losses to enforce uniform performance distributions. Under training-time data poisoning attacks (such as label flipping), malicious devices naturally generate large training errors. Consequently, fair objectives assign disproportionately large weights to corrupted devices, leading the global model to overfit heavily to adversarial inputs and causing severe degradation in benign device accuracy (e.g., in TERM, higher values of fairness parameter tt yield lower benign test accuracy under attack). Conversely, robust aggregation defenses (such as coordinate-wise median, Krum, Multi-Krum, gradient clipping, kk-norm, and kk-loss) filter out or trim statistical outliers, which discards rare but informative updates from legitimate minority devices, exacerbating performance variance and representation disparity.

  8. Knowl 8 — Accuracy and Fairness Benchmarking of Ditto against Global, Local, and Fair Baselines

    data/table

    Average test accuracy and standard deviation (in parentheses) across benign devices on Fashion MNIST and FEMNIST under clean conditions and three attack types: label poisoning (A1), random Gaussian updates (A2), and model replacement (A3). Ditto achieves higher mean test accuracy and lower or competitive performance variance across all scenarios, while fair global models (TERM, t=1t=1) degrade severely under adversarial attacks.

    Fashion MNIST Clean A1 (% adversaries) A2 (% adversaries) A3 (% adversaries)
    Method 20% 50% 80% 20% 50% 80% 10% 20% 50%
    Global (FedAvg) .911 (.08) .897 (.08) .855 (.10) .753 (.13) .900 (.08) .882 (.09) .857 (.10) .753 (.10) .551 (.13) .275 (.12)
    Local .876 (.10) .874 (.10) .876 (.11) .879 (.10) .874 (.10) .876 (.11) .879 (.10) .877 (.10) .874 (.10) .876 (.11)
    Fair (TERM, t=1t=1) .909 (.07) .751 (.12) .637 (.13) .547 (.11) .731 (.13) .637 (.14) .635 (.14) .653 (.13) .601 (.12) .131 (.16)
    Ditto .943 (.06) .944 (.07) .937 (.07) .907 (.10) .938 (.07) .930 (.08) .913 (.09) .921 (.09) .902 (.09) .873 (.11)
    FEMNIST Clean A1 (% adversaries) A2 (% adversaries) A3 (% adversaries)
    Method 20% 50% 80% 20% 50% 80% 10% 15% 20%
    Global (FedAvg) .804 (.11) .773 (.11) .727 (.12) .574 (.15) .774 (.11) .703 (.14) .636 (.15) .517 (.14) .487 (.14) .314 (.13)
    Local .628 (.15) .620 (.14) .627 (.14) .607 (.14) .620 (.14) .627 (.14) .607 (.14) .622 (.14) .621 (.14) .620 (.14)
    Fair (TERM, t=1t=1) .809 (.11) .636 (.15) .562 (.13) .478 (.12) .440 (.15) .336 (.12) .363 (.12) .353 (.12) .316 (.12) .299 (.11)
    Ditto .834 (.09) .802 (.10) .762 (.11) .672 (.13) .801 (.09) .700 (.15) .675 (.14) .685 (.15) .650 (.14) .613 (.13)
  9. Knowl 9 — Comparison of Ditto with Existing Personalized Federated Learning Methods

    data/table

    Average test accuracy and standard deviation (in parentheses) across all devices on FEMNIST (62-class CNN) and CelebA (binary CNN) benchmarks under clean conditions and a 50% label poisoning attack (A1). Ditto matches or outperforms competing personalization approaches—including mean-regularized multi-task learning (L2SGD), elastic weight consolidation (EWC), symmetrized KL divergence (SKL), meta-learning (Per-FedAvg HF), model interpolation (APFL, Mapper), and plain finetuning—without requiring second-order gradient calculations or Fisher matrix estimations.

    Clean 50% Adversaries (A1)
    Methods FEMNIST CelebA FEMNIST CelebA
    Global (FedAvg) .804 (.11) .911 (.19) .727 (.12) .538 (.28)
    Local .628 (.15) .692 (.27) .627 (.14) .682 (.27)
    Plain finetuning .815 (.09) .912 (.18) .734 (.12) .721 (.28)
    L2SGD .817 (.10) .899 (.18) .732 (.15) .725 (.25)
    EWC .810 (.11) .910 (.18) .756 (.12) .642 (.26)
    SKL .820 (.10) .915 (.16) .752 (.12) .708 (.27)
    Per-FedAvg (HF) .827 (.09) .907 (.17) .604 (.14) .756 (.26)
    Mapper .792 (.12) .773 (.25) .726 (.13) .704 (.27)
    APFL .811 (.11) .911 (.17) .750 (.11) .710 (.27)
    Ditto .836 (.10) .914 (.18) .767 (.10) .721 (.27)
  10. Knowl 10 — Augmenting Ditto with Byzantine-Robust Aggregation

    empirical result

    Because the global objective G(F1(w),…,FK(w))G(F_1(w), \dots, F_K(w)) in Ditto is modular and decoupled from local personalized models, existing Byzantine-robust aggregation rules (such as gradient clipping or Multi-Krum) can be integrated directly into the server aggregation step (Line 7 in Algorithm 1). Combining Ditto with robust aggregation yields performance superior to either standalone robust aggregation or vanilla Ditto, especially against strong model replacement attacks (A3). On the FEMNIST dataset:

    • Under 10% model replacement (A3): Global FedAvg achieves 0.517 accuracy, gradient clipping alone achieves 0.795, vanilla Ditto achieves 0.695, and Ditto + gradient clipping achieves 0.813.
    • Under 20% model replacement (A3): Global FedAvg achieves 0.364 accuracy, gradient clipping alone collapses to 0.061, vanilla Ditto achieves 0.650, and Ditto + gradient clipping achieves 0.672.
  11. Knowl 11 — Superiority of Joint Alternating Optimization over Post-Hoc Local Finetuning under Model Poisoning

    empirical result

    In non-convex neural network training under training-time model poisoning attacks (such as model replacement A3), solving Ditto via joint alternating optimization (Algorithm 1) significantly outperforms post-hoc local finetuning from the converged global model w∗w^*. Under iterative training attacks, the global model wtw^t is initially uncorrupted and gradually degrades as malicious updates accumulate. Joint alternating optimization allows local devices to incorporate useful global representations from earlier rounds before corruption dominates, steering personalized models into favorable optimization basins. In contrast, finetuning after global convergence begins from a heavily poisoned global checkpoint, yielding inferior test accuracy. While early stopping of global training before finetuning can theoretically help, establishing a reliable stopping criterion in practice is infeasible because network-wide loss and validation accuracy curves on heterogeneous data do not distinguish clean from corrupted states without oracle knowledge of device identities.

  12. Knowl 12 — Client-Side Validation Heuristic for Personalization Parameter Lambda

    model/method

    To tune the regularization hyperparameter λ\lambda without requiring the server to identify which devices are benign versus malicious, each device selects λ\lambda locally based on its local validation data. Devices evaluate whether an attack is strong (defined as model replacement attacks scaling updates by >10×>10\times or attacks where >50%>50\% of devices are corrupted):

    • For devices with fewer than 4 validation samples, a fixed value is used: λ=0.1\lambda = 0.1 for strong attacks, and λ=1.0\lambda = 1.0 for other attacks.
    • For devices with 5 or more validation samples, the parameter is selected from a candidate grid: λ∈{0.05,0.1,0.2}\lambda \in \{0.05, 0.1, 0.2\} for strong attacks, and λ∈{0.1,1.0,2.0}\lambda \in \{0.1, 1.0, 2.0\} for other attacks (or {0.01,0.05,0.1}\{0.01, 0.05, 0.1\} vs {0.05,0.1,0.3}\{0.05, 0.1, 0.3\} on StackOverflow).

    Because malicious devices' poisoned updates are independent of λ\lambda, local parameter selection on benign devices operates entirely on clean local validation signals, achieving test performance comparable to an oracle server-side tuning baseline that possesses ground-truth knowledge of malicious device identities.

Coverage note — The one-dimensional federated point estimation derivations (Section A.3) were omitted as standalone knowls because they are strictly special cases ($d=1$, $X_k = \mathbf{1}_{n \times 1}$) of the multi-dimensional federated linear regression theorems presented in Section A.2.

References

  1. 1.Tensorflow federated: Machine learning on decentralized data. URL https://www.tensorflow.org/federated.
  2. 2.Agarwal, A., Langford, J., and Wei, C.-Y. Federated residual learning. arXiv preprint arXiv:2003.12880, 2020.
  3. 3.Bagdasaryan, E., Veit, A., Hua, Y., Estrin, D., and Shmatikov, V. How to backdoor federated learning. In International Conference on Artificial Intelligence and Statistics, 2020.
  4. 4.Bhagoji, A. N., Chakraborty, S., Mittal, P., and Calo, S. Analyzing federated learning through an adversarial lens. In International Conference on Machine Learning, 2019.
  5. 5.Biggio, B., Nelson, B., and Laskov, P. Support vector machines under adversarial label noise. In Asian Conference on Machine Learning, 2011.
  6. 6.Biggio, B., Nelson, B., and Laskov, P. Poisoning attacks against support vector machines. In International Conference on Machine Learning, 2012.
  7. 7.Blanchard, P., Mhamdi, E. M. E., Guerraoui, R., and Stainer, J. Machine learning with adversaries: Byzantine tolerant gradient descent. In Advances in Neural Information Processing Systems, 2017.
  8. 8.Caldas, S., Wu, P., Li, T., Konečnỳ, J., McMahan, H. B., Smith, V., and Talwalkar, A. Leaf: A benchmark for federated settings. arXiv preprint arXiv:1812.01097, 2018.
  9. 9.Chang, H., Nguyen, T. D., Murakonda, S. K., Kazemi, E., and Shokri, R. On adversarial bias and the robustness of fair machine learning. arXiv preprint arXiv:2006.08669, 2020.
  10. 10.Chen, F., Luo, M., Dong, Z., Li, Z., and He, X. Federated meta-learning with fast convergence and efficient communication. arXiv preprint arXiv:1802.07876, 2018.
  11. 11.Chen, X., Liu, C., Li, B., Lu, K., and Song, D. Targeted backdoor attacks on deep learning systems using data poisoning. arXiv preprint arXiv:1712.05526, 2017.
  12. 12.Cohen, G., Afshar, S., Tapson, J., and van Schaik, A. Emnist: an extension of mnist to handwritten letters. arXiv preprint arXiv:1702.05373, 2017.
  13. 13.Deng, Y., Kamani, M. M., and Mahdavi, M. Distributionally robust federated averaging. Advances in Neural Information Processing Systems, 2020.
  14. 14.Deng, Y., Kamani, M. M., and Mahdavi, M. Adaptive personalized federated learning, 2021. URL https://openreview.net/forum?id=g0a-XYjpQ7r.
  15. 15.Dinh, C. T., Tran, N. H., and Nguyen, T. D. Personalized federated learning with moreau envelopes. In Advances in Neural Information Processing Systems, 2020.
  16. 16.Duarte, M. F. and Hu, Y. H. Vehicle classification in distributed sensor networks. Journal of Parallel and Distributed Computing, 2004.
  17. 17.Dumford, J. and Scheirer, W. Backdooring convolutional neural networks via targeted weight perturbations. arXiv preprint arXiv:1812.03128, 2018.
  18. 18.Evgeniou, T. and Pontil, M. Regularized multi–task learning. In International Conference on Knowledge Discovery and Data Mining, 2004.
  19. 19.Fallah, A., Mokhtari, A., and Ozdaglar, A. Personalized federated learning: A meta-learning approach. In Advances in Neural Information Processing Systems, 2020.
  20. 20.Fang, M., Cao, X., Jia, J., and Gong, N. Local model poisoning attacks to byzantine-robust federated learning. In USENIX Security Symposium, 2020.
  21. 21.Finn, C., Abbeel, P., and Levine, S. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning, 2017.
  22. 22.Ghosh, A., Chung, J., Yin, D., and Ramchandran, K. An efficient framework for clustered federated learning. In Advances in Neural Information Processing Systems, 2020.
  23. 23.Gu, T., Dolan-Gavitt, B., and Garg, S. Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733, 2017.
  24. 24.Hanzely, F. and Richtárik, P. Federated learning of a mixture of global and local models. arXiv preprint arXiv:2002.05516, 2020.
  25. 25.Hanzely, F., Hanzely, S., Horváth, S., and Richtárik, P. Lower bounds and optimal algorithms for personalized federated learning. Advances in Neural Information Processing Systems, 2020.
  26. 26.Hao, W., Mehta, N., Liang, K. J., Cheng, P., El-Khamy, M., and Carin, L. Waffle: Weight anonymized factorization for federated learning. arXiv preprint arXiv:2008.05687, 2020.
  27. 27.Hashimoto, T., Srivastava, M., Namkoong, H., and Liang, P. Fairness without demographics in repeated loss minimization. In International Conference on Machine Learning, 2018.
  28. 28.He, L., Karimireddy, S. P., and Jaggi, M. Byzantine-robust learning on heterogeneous datasets via resampling. In NeurIPS Workshop on Scalability, Privacy, and Security in Federated Learning, 2020.
  29. 29.Hu, Z., Shaloudegi, K., Zhang, G., and Yu, Y. FedMGDA+: Federated learning meets multi-objective optimization. arXiv preprint arXiv:2006.11489, 2020.
  30. 30.Huang, W. R., Geiping, J., Fowl, L., Taylor, G., and Goldstein, T. Metapoison: Practical general-purpose cleanlabel data poisoning. In Advances in Neural Information Processing Systems, 2020.
  31. 31.Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Zhao, T. SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020.
  32. 32.Jiang, Y., Konečnỳ, J., Rush, K., and Kannan, S. Improving federated learning personalization via model agnostic meta learning. arXiv preprint arXiv:1909.12488, 2019.
  33. 33.Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., et al. Advances and open problems in federated learning. arXiv preprint arXiv:1912.04977, 2019.
  34. 34.Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A. T. Scaffold: Stochastic controlled averaging for federated learning. In International Conference on Machine Learning, 2020.
  35. 35.Khodak, M., Balcan, M.-F. F., and Talwalkar, A. S. Adaptive gradient-based meta-learning methods. In Advances in Neural Information Processing Systems, 2019.
  36. 36.Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 2017.
  37. 37.Lamport, L., Shostak, R., and Pease, M. The byzantine generals problem. In Concurrency: the Works of Leslie Lamport. 2019.
  38. 38.Li, J., Khodak, M., Caldas, S., and Talwalkar, A. Differentially private meta-learning. In International Conference on Learning Representations, 2020a.
  39. 39.Li, L., Xu, W., Chen, T., Giannakis, G. B., and Ling, Q. Rsa: Byzantine-robust stochastic aggregation methods for distributed learning from heterogeneous datasets. In AAAI Conference on Artificial Intelligence, 2019.
  40. 40.Li, M., Soltanolkotabi, M., and Oymak, S. Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks. In International Conference on Artificial Intelligence and Statistics, 2020b.
  41. 41.Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V. Federated optimization in heterogeneous networks. In Conference on Machine Learning and Systems, 2020c.
  42. 42.Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V. Federated optimization in heterogeneous networks. Proceedings of Machine Learning and Systems, 2020d.
  43. 43.Li, T., Sanjabi, M., Beirami, A., and Smith, V. Fair resource allocation in federated learning. In International Conference on Learning Representations, 2020e.
  44. 44.Li, T., Beirami, A., Sanjabi, M., and Smith, V. Tilted empirical risk minimization. In International Conference on Learning Representations, 2021.
  45. 45.Li, X., Huang, K., Yang, W., Wang, S., and Zhang, Z. On the convergence of fedavg on non-iid data. In International Conference on Learning Representations, 2020f.
  46. 46.Liang, P. P., Liu, T., Ziyin, L., Salakhutdinov, R., and Morency, L.-P. Think locally, act globally: Federated learning with local and global representations. arXiv preprint arXiv:2001.01523, 2020.
  47. 47.Liu, Y., Ma, S., Aafer, Y., Lee, W., Zhai, J., Wang, W., and Zhang, X. Trojaning attack on neural networks. In Network and Distributed System Security Symposium, 2018.
  48. 48.Liu, Z., Luo, P., Wang, X., and Tang, X. Deep learning face attributes in the wild. In International Conference on Computer Vision, 2015.
  49. 49.London, B. PAC identifiability in federated personalization. In NeurIPS 2020 Workshop on Scalability, Privacy, and Security in Federated Learning, 2020.
  50. 50.Mahdavifar, H., Beirami, A., Touri, B., and Shamma, J. S. Global games with noisy information sharing. IEEE Transactions on Signal and Information Processing over Networks, 2018.
  51. 51.Mansour, Y., Mohri, M., Ro, J., and Suresh, A. T. Three approaches for personalization with applications to federated learning. arXiv preprint arXiv:2002.10619, 2020.
  52. 52.McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A. Communication-efficient learning of deep networks from decentralized data. In International Conference on Artificial Intelligence and Statistics, 2017.
  53. 53.Mohri, M., Sivek, G., and Suresh, A. T. Agnostic federated learning. In International Conference on Machine Learning, 2019.
  54. 54.Muhammad, K., Wang, Q., O’Reilly-Morgan, D., Tragos, E., Smyth, B., Hurley, N., Geraci, J., and Lawlor, A. Fedfast: Going beyond average for faster training of federated recommender systems. In International Conference on Knowledge Discovery & Data Mining, 2020.
  55. 55.Pillutla, K., Kakade, S. M., and Harchaoui, Z. Robust aggregation for federated learning. arXiv preprint arXiv:1912.13445, 2019.
  56. 56.Reddi, S., Charles, Z., Zaheer, M., Garrett, Z., Rush, K., Konečnỳ, J., Kumar, S., and McMahan, H. B. Adaptive federated optimization. In International Conference on Learning Representations, 2021.
  57. 57.Sattler, F., Müller, K.-R., and Samek, W. Clustered federated learning: Model-agnostic distributed multitask optimization under privacy constraints. IEEE Transactions on Neural Networks and Learning Systems, 2020.
  58. 58.Schwarz, J., Czarnecki, W., Luketina, J., Grabska-Barwinska, A., Teh, Y. W., Pascanu, R., and Hadsell, R. Progress & compress: A scalable framework for continual learning. In International Conference on Machine Learning, 2018.
  59. 59.Shafahi, A., Huang, W. R., Najibi, M., Suciu, O., Studer, C., Dumitras, T., and Goldstein, T. Poison frogs! targeted clean-label poisoning attacks on neural networks. In Advances in Neural Information Processing Systems, 2018.
  60. 60.Singhal, K., Sidahmed, H., Garrett, Z., Wu, S., Rush, K., and Prakash, S. Federated reconstruction: Partially local federated learning. arXiv preprint arXiv:2102.03448, 2021.
  61. 61.Smith, V., Chiang, C.-K., Sanjabi, M., and Talwalkar, A. S. Federated multi-task learning. In Advances in Neural Information Processing Systems, 2017.
  62. 62.Sun, G., Cong, Y., Dong, J., Wang, Q., and Liu, J. Data poisoning attacks on federated machine learning. arXiv preprint arXiv:2004.10020, 2020.
  63. 63.Sun, Z., Kairouz, P., Suresh, A. T., and McMahan, H. Can you really backdoor federated learning? arXiv preprint arXiv:1911.07963, 2019.
  64. 64.Wang, H., Sreenivasan, K., Rajput, S., Vishwakarma, H., Agarwal, S., Sohn, J.-y., Lee, K., and Papailiopoulos, D. Attack of the tails: Yes, you really can backdoor federated learning. In Advances in Neural Information Processing Systems, 2020.
  65. 65.Wang, K., Mathews, R., Kiddon, C., Eichner, H., Beaufays, F., and Ramage, D. Federated evaluation of on-device personalization. arXiv preprint arXiv:1910.10252, 2019.
  66. 66.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017.
  67. 67.Xie, C., Huang, K., Chen, P.-Y., and Li, B. DBA: Distributed backdoor attacks against federated learning. In International Conference on Learning Representations, 2020.
  68. 68.Xu, X. and Lyu, L. Towards building a robust and fair federated learning system. arXiv preprint arXiv:2011.10464, 2020.
  69. 69.Yin, D., Chen, Y., Kannan, R., and Bartlett, P. Byzantine-robust distributed learning: Towards optimal statistical rates. In International Conference on Machine Learning, 2018.
  70. 70.Yu, T., Bagdasaryan, E., and Shmatikov, V. Salvaging federated learning by local adaptation. arXiv preprint arXiv:2002.04758, 2020.
  71. 71.Zhang, M., Sapra, K., Fidler, S., Yeung, S., and Alvarez, J. M. Personalized federated learning with first order model optimization. In International Conference on Learning Representations, 2021.
  72. 72.Zhao, Y., Li, M., Lai, L., Suda, N., Civin, D., and Chandra, V. Federated learning with non-iid data. arXiv preprint arXiv:1806.00582, 2018.

Citation

MLA
Li, T., et al. “Ditto: Fair and Robust Federated Learning Through Personalization”. arXiv, 2020, http://arxiv.org/abs/2012.04221v3.
APA
Li, T., Hu, S., Beirami, A., & Smith, V. (2020). Ditto: Fair and Robust Federated Learning Through Personalization. arXiv. http://arxiv.org/abs/2012.04221v3
Chicago
Li, T., S. Hu, A. Beirami, and V. Smith. 2020. “Ditto: Fair and Robust Federated Learning Through Personalization”. arXiv. http://arxiv.org/abs/2012.04221v3.
Harvard
Li, T. et al. (2020) “Ditto: Fair and Robust Federated Learning Through Personalization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2012.04221v3.
Vancouver
1. Li T, Hu S, Beirami A, Smith V (2020) Ditto: Fair and Robust Federated Learning Through Personalization. arXiv

BibTeX

@article{li2020ditto,
  title = {Ditto: Fair and Robust Federated Learning Through Personalization},
  author = {Li, Tian and Hu, Shengyuan and Beirami, Ahmad and Smith, Virginia},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2012.04221v3},
  eprint = {2012.04221}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/