Exploiting Unintended Feature Leakage in Collaborative Learning

Luca MelisCongzheng SongEmiliano De CristofaroVitaly Shmatikov

article2018IEEE Symposium on Security and Privacy1,833 citations

Demonstrates that shared parameter updates in federated learning leak sensitive training data, enabling adversarial participants to extract specific records and infer private user attributes through both passive and active inference attacks.

Listen

Collaborative machine learning frameworks, such as federated learning, are widely adopted because they allow multiple organizations or user devices to jointly train models without directly sharing raw, sensitive data. This approach is frequently deployed in privacy-critical domains, such as healthcare collaborations and mobile applications. However, exchanging intermediate model updates creates an attack surface. The article sets out to evaluate whether an adversarial participant can exploit model updates during training to reconstruct exact data points or extract unintended sensitive properties that are entirely unrelated to the primary machine learning task.

To evaluate this threat, the authors developed and tested passive and active inference attacks across various collaborative configurations, including synchronized gradient updates and federated model averaging. The experiments encompassed multiple standard datasets covering vision, textual reviews, and location check-ins (such as LFW, FaceScrub, PIPA, Yelp reviews, and FourSquare). The attack framework involved training binary meta-classifiers on model snapshots and gradient updates to detect data membership and internal feature representations, testing groups ranging from two-party setups to multi-party configurations of up to 30 participants.

The findings show substantial information leakage across learning frameworks. First, the non-zero gradients in sparse embedding layers allowed the adversary to perform membership inference with perfect recall and up to 0.99 precision for location data and 0.92 for medical text batches. Second, the authors proved that deep learning models internally construct representations of unintended features; for example, an adversary inferred race or eyewear attributes with near-perfect accuracy (up to 1.0 area under the curve) during a gender-classification task, even when those attributes had zero correlation with the target task. Third, active adversaries using multi-task learning significantly separated target properties internally, allowing them to detect specific individuals even when target records made up only a small fraction of a batch. Fourth, the attacks successfully tracked temporal patterns, such as identifying the exact training round a specific user or property entered the pool. Although increasing the number of participants beyond 12 reduced vision-based attack accuracy, text-based author identification remained resilient even across 30 participants.

These results demonstrate that keeping raw data on local devices during collaborative learning does not guarantee privacy, creating compliance, regulatory, and confidentiality risks for participating organizations. Standard heuristic mitigations—including sharing fewer gradients, reducing vocabulary sizes, and adding dropout—failed to prevent leakage and frequently degraded joint model utility. Furthermore, rigorous defenses like participant-level differential privacy prevented model convergence in deployments with small participant counts.

Organizations deploying collaborative learning should recognize that gradient updates expose sensitive internal features and avoid assuming that distributed training is private by default. For decision-makers, further work is required before deploying collaborative models in sensitive domains with small participant pools. Next steps should focus on developing "least-privilege" feature-censoring techniques that constrain models to learn only task-relevant representations, alongside monitoring mechanisms to detect active multi-task manipulation.

Confidence in these findings is high for small to medium collaborative environments, but limitations apply. The attacks require auxiliary labeled data to train property classifiers, and their efficacy diminishes in very large-scale settings with thousands of participants where updates are heavily aggregated. Readers should therefore evaluate risks based on their specific participant scale and data sensitivity.

  • Paper: Deep Leakage from Gradients, Ligeng Zhu et al. (2019). This paper significantly deepens gradient-based privacy leakage in collaborative learning by demonstrating exact, pixel-level input reconstruction from shared gradients beyond the property and membership inferences shown in the source paper.
  • Paper: How To Backdoor Federated Learning, Eugene Bagdasaryan et al. (2018). This work explores active adversarial manipulation in federated learning by introducing backdoor injection via model replacement, building on the collaborative threat vectors identified in the source paper.
  • Paper: Federated Learning With Differential Privacy: Algorithms and Performance Analysis, Kang Wei et al. (2019). This study develops formal differential privacy mechanisms before model aggregation to theoretically analyze and mitigate the parameter-level information leakage highlighted by the source paper.
  • Paper: Advances and Open Problems in Federated Learning, P. Kairouz et al. (2019). This comprehensive survey synthesizes the broader privacy and security vulnerabilities of federated learning, contextualizing the update-leakage and property-inference risks proven in the source paper.
  • Paper: Federated Machine Learning, Qiang Yang et al. (2019). This survey provides a comprehensive architectural framework and privacy preservation taxonomy for federated learning in response to the data leakage risks demonstrated in collaborative learning setups.
Cover for Exploiting Unintended Feature Leakage in Collaborative Learning

Abstract

Collaborative machine learning and related techniques such as federated learning allow multiple participants, each with his own training dataset, to build a joint model by training locally and periodically exchanging model updates. We demonstrate that these updates leak unintended information about participants' training data and develop passive and active inference attacks to exploit this leakage. First, we show that an adversarial participant can infer the presence of exact data points -- for example, specific locations -- in others' training data (i.e., membership inference). Then, we show how this adversary can infer properties that hold only for a subset of the training data and are independent of the properties that the joint model aims to capture. For example, he can infer when a specific person first appears in the photos used to train a binary gender classifier. We evaluate our attacks on a variety of tasks, datasets, and learning configurations, analyze their limitations, and discuss possible defenses.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Machine learning (ML)
  • 2.2 Collaborative learning
  • 3 Reasoning about Privacy in Machine Learning
  • 3.1 Inferring class representatives
  • 3.2 Inferring membership in training data
  • 3.3 Inferring properties of training data
  • 4 Inference Attacks
  • 4.1 Threat model
  • 4.2 Overview of the attacks
  • 4.3 Membership inference
  • 4.4 Passive property inference
  • 4.5 Active property inference
  • 5 Datasets and model architectures
  • 6 Two-Party Experiments
  • 6.1 Membership inference
  • 6.2 Single-batch property inference
  • 6.3 Inferring when a property occurs
  • 6.4 Inference against well-generalized models
  • 6.5 Active property inference
  • 7 Multi-Party Experiments
  • 7.1 Synchronized SGD
  • 7.2 Model averaging
  • 8 Defenses
  • 8.1 Sharing fewer gradients
  • 8.2 Dimensionality reduction
  • 8.3 Dropout
  • 8.4 Participant-level differential privacy
  • 9 Limitations of the attacks
  • 9.1 Auxiliary data
  • 9.2 Number of participants
  • 9.3 Undetectable properties
  • 9.4 Attribution of inferred properties
  • 10 Related Work
  • 11 Conclusion
  • References

Knowls

  1. Knowl 1 — Passive Batch Property Classifier for Collaborative Learning

    algorithm

    An adversarial participant in collaborative learning can train a binary property classifier to infer whether other participants' training batches contain inputs exhibiting a specific property pp, using only observed parameter updates without deviating from the collaborative protocol.

    Input: Auxiliary datasets DpropadvD_{\text{prop}}^{\text{adv}} (data with property) and DnonpropadvD_{\text{nonprop}}^{\text{adv}} (data without property), number of collaborative training rounds TT, model parameter trajectory θ0,θ1,…,θT\theta_0, \theta_1, \dots, \theta_T
    Output: Trained binary batch property classifier fpropf_{\text{prop}}
    Gprop←∅G_{\text{prop}} \leftarrow \emptyset
    Gnonprop←∅G_{\text{nonprop}} \leftarrow \emptyset
    for t=1t = 1 to TT do
        Receive current joint model parameters θt\theta_t from parameter server
        Sample batch bpropadv⊂Dpropadvb_{\text{prop}}^{\text{adv}} \subset D_{\text{prop}}^{\text{adv}}
        Sample batch bnonpropadv⊂Dnonpropadvb_{\text{nonprop}}^{\text{adv}} \subset D_{\text{nonprop}}^{\text{adv}}
        Calculate gradient gprop←∇θL(bpropadv;θt)g_{\text{prop}} \leftarrow \nabla_\theta L(b_{\text{prop}}^{\text{adv}}; \theta_t)
        Calculate gradient gnonprop←∇θL(bnonpropadv;θt)g_{\text{nonprop}} \leftarrow \nabla_\theta L(b_{\text{nonprop}}^{\text{adv}}; \theta_t)
        Gprop←Gprop∪{gprop}G_{\text{prop}} \leftarrow G_{\text{prop}} \cup \{g_{\text{prop}}\}
        Gnonprop←Gnonprop∪{gnonprop}G_{\text{nonprop}} \leftarrow G_{\text{nonprop}} \cup \{g_{\text{nonprop}}\}
    end for
    Label all updates in GpropG_{\text{prop}} as positive (11) and in GnonpropG_{\text{nonprop}} as negative (00)
    Apply max pooling downsampling over gradient vectors to reduce feature dimensionality
    Train binary classifier fpropf_{\text{prop}} (e.g., Random Forest with 50 trees or Logistic Regression) on labeled gradient pairs (Gprop,Gnonprop)(G_{\text{prop}}, G_{\text{nonprop}})
    return fpropf_{\text{prop}}

    To execute the attack during training, at step tt the adversary computes the observed update contributed by other participants, gobs=(θt−θt−1)−Δθtadvg_{\text{obs}} = (\theta_t - \theta_{t-1}) - \Delta\theta_t^{\text{adv}}, where Δθtadv\Delta\theta_t^{\text{adv}} is the adversary's own submitted update. The observed gradient gobsg_{\text{obs}} is passed to fpropf_{\text{prop}}, which outputs a probability score in [0,1][0, 1] indicating whether the underlying batch contains the target property. For entire-dataset inference, scores are averaged across all training rounds.

  2. Knowl 2 — Active Property Inference Attack via Multi-Task Learning

    model/method

    An active adversarial participant can enhance property leakage from honest participants by manipulating the shared model representation using multi-task learning during collaborative training.

    The adversary attaches an auxiliary classification head to the final layer of their local copy of the joint model. The auxiliary head is designed to predict a sensitive property label pp, while the primary head predicts the main task label yy. Given a training example (x,y,p)(x, y, p), the adversary optimizes a joint loss function:

    Lmt(x,y,p;θ)=α⋅Lmain(x,y;θ)+(1−α)⋅Lprop(x,p;θ)L_{\text{mt}}(x, y, p; \theta) = \alpha \cdot L_{\text{main}}(x, y; \theta) + (1 - \alpha) \cdot L_{\text{prop}}(x, p; \theta)

    where LmainL_{\text{main}} is the loss on the main classification task, LpropL_{\text{prop}} is the loss on the property classification task, θ\theta denotes the model parameters, and α∈[0,1]\alpha \in [0, 1] is a hyperparameter weighting the two objectives.

    During each round of collaborative training, the adversary uploads gradient updates ∇θLmt\nabla_\theta L_{\text{mt}} computed with respect to the joint objective. This forces the underlying neural network representations to learn features that separate data with and without property pp. Consequently, the gradient updates computed by honest participants on batches containing the property become distinctly separable from updates on batches without the property, substantially increasing the accuracy of subsequent property inference attacks without degrading main-task accuracy.

  3. Knowl 3 — Membership Inference Attack via Sparse Embedding Layer Gradients

    algorithm

    In collaborative learning models processing discrete token sequences (such as natural language or location check-ins), inputs are mapped to continuous vectors via an embedding matrix Wemb∈R∣V∣×dW_{\text{emb}} \in \mathbb{R}^{|V| \times d}, where ∣V∣|V| is the vocabulary size and dd is embedding dimensionality. During backward propagation, the gradient ∇WembL\nabla_{W_{\text{emb}}} L is sparse: rows corresponding to tokens absent from the training batch are strictly zero, whereas rows corresponding to present tokens are non-zero.

    An adversary can exploit this sparsity to infer whether a specific record rr containing a set of tokens Vr⊂VV_r \subset V was part of any participant's training batch:

    Input: Candidate record rr with token set Vr⊆VV_r \subseteq V, sequence of observed parameter updates Δθ1,Δθ2,…,ΔθT\Delta\theta_1, \Delta\theta_2, \dots, \Delta\theta_T across TT training iterations
    Output: Binary decision M∈{True,False}M \in \{\text{True}, \text{False}\} indicating membership of rr
    for t=1t = 1 to TT do
        Extract non-zero rows from the embedding parameter update in Δθt\Delta\theta_t to form batch vocabulary VtV_t
        if Vr⊆VtV_r \subseteq V_t then
            return True
        end if
    end for
    return False

    Because any true training record must appear in at least one training batch, the adversary's observed batch vocabulary VtV_t will contain VrV_r, yielding a recall of 1.01.0 (no false negatives). The precision of the attack decreases as the batch size increases, since larger batches contain more unique tokens and elevate the likelihood of accidental superset containment.

  4. Knowl 4 — Internal Feature Representation and Gradient Leakage in Deep Neural Networks

    model/method

    During backpropagation in deep neural networks, parameter gradients at intermediate layers directly incorporate the activations (features) computed at earlier layers. For a fully connected layer with linear mapping hl+1=Wlhlh_{l+1} = W_l h_l, the gradient of the loss EE with respect to weight matrix WlW_l is:

    ∂E∂Wl=∂E∂hl+1⋅hlT\frac{\partial E}{\partial W_l} = \frac{\partial E}{\partial h_{l+1}} \cdot h_l^T

    where hlh_l is the activation vector from layer ll and ∂E∂hl+1\frac{\partial E}{\partial h_{l+1}} is the backpropagated error vector from layer l+1l+1. For convolutional layers, weight gradients are convolutions between backpropagated error maps and feature maps hlh_l.

    Deep neural networks naturally learn intermediate representations in early and intermediate layers (e.g., convolutional and pooling layers) that cluster data according to auxiliary features and demographic attributes, even when those features are completely independent of the target classification label. Because the weight gradients ∇WlE\nabla_{W_l} E depend directly on hlh_l, the parameter updates transmitted in collaborative learning inherit these cluster separations and leak feature values to an observer.

  5. Knowl 5 — Single-Batch Property Inference on Uncorrelated Features in LFW

    data/table

    In two-party collaborative learning on the Labeled Faces in the Wild (LFW) dataset, an adversary can infer demographic and visual attributes from honest participants' gradient updates, even when those attributes have zero or near-zero correlation with the model's main classification objective.

    Main Task Inferred Property Pearson Correlation Property AUC
    Gender Race: Black -0.005 1.00
    Gender Race: Asian -0.018 0.93
    Gender Eyewear: Sunglasses -0.025 1.00
    Gender Eyewear: Eyeglasses 0.157 0.94
    Smile Race: Black 0.062 1.00
    Smile Race: Asian 0.047 0.93
    Smile Eyewear: Sunglasses -0.016 1.00
    Smile Eyewear: Eyeglasses -0.083 0.97
    Age Race: Black -0.084 1.00
    Age Race: Asian -0.078 0.97
    Race Eyewear: Sunglasses 0.026 1.00
    Race Eyewear: Eyeglasses -0.116 0.96
    Eyewear Race: Black 0.034 1.00
    Eyewear Race: Asian -0.119 0.91
    Hair Eyewear: Sunglasses -0.013 1.00
    Hair Eyewear: Eyeglasses 0.139 0.96

    The data shows that despite negligible Pearson correlation between the primary classification label and the target attribute (e.g., Pearson r=−0.005r = -0.005 between Gender and Race: Black, or r=−0.025r = -0.025 between Gender and Sunglasses), the binary batch property classifier achieves area under the ROC curve (AUC) scores between 0.910.91 and 1.001.00. This confirms that collaborative gradient updates disclose unintended features that are distinct from class characteristics.

  6. Knowl 6 — Property Inference Under Fractional Presence and Dynamic Batch Occurrences

    empirical result

    Passive property inference attacks succeed when target features are present in only a fraction of batch samples and can accurately track when properties enter or exit the training stream over time:

    1. Fractional Batch Presence: On the FaceScrub dataset (binary gender classification main task), inferring whether a specific individual's face is present in a batch achieves an AUC score exceeding 0.800.80 when only 50%50\% of the images in the batch belong to that individual. On the Yelp-author dataset (review score prediction main task), identifying the author of reviews achieves an AUC score above 0.950.95 when as few as 30%30\% of the reviews in the batch belong to the target author.
    2. Temporal Tracking of Property Occurrence: When training on the PIPA dataset to classify whether an image contains a young adult, the adversary's property classifier tracks the presence of a "same gender" attribute, outputting elevated probability scores specifically during iterations 500500 to 15001500 where target batches are injected, and near-zero scores elsewhere. On FaceScrub, property classifiers for specific identities accurately output high scores during designated active intervals (e.g., ID 0 during iterations 0–5000\text{--}500 and 1500–20001500\text{--}2000, ID 1 during iterations 500–1000500\text{--}1000 and 2000–25002000\text{--}2500) while assigning near-zero probabilities across all iterations to an identity (ID 2) absent from the target dataset.
  7. Knowl 7 — Property Inference in Multi-Party Synchronized SGD and Federated Model Averaging

    empirical result

    Property inference attacks extend to multi-party collaborative learning and federated learning with model averaging, though efficacy depends on participant count and aggregation:

    1. Multi-Party Synchronized SGD: When KK participants train concurrently with synchronized gradient updates, the adversary observes the sum of updates ∑k≠advgtk\sum_{k \ne \text{adv}} g_t^k. On LFW (gender classification main task, race property inference), attack AUC decreases as KK grows, remaining around 0.800.80 at K=12K = 12 participants before degrading further at K=24K = 24. On Yelp-author, the adversary observes the union of token vocabularies across all honest participants; however, author identification AUC remains above 0.800.80–0.950.95 for several prolific authors even at K=30K = 30 participants due to distinctive word usage patterns.
    2. Federated Model Averaging: In federated learning where clients submit locally aggregated models θtk\theta_t^k computed over 1010 local batches per round, the adversary emulates local training to distinguish aggregated parameter updates. On FaceScrub, the adversary successfully separates trials where a target face constitutes 80%80\% of one client's data from control trials (0%0\%) at K=3K = 3 and K=5K = 5 participants for select identities (IDs 1 and 3). Furthermore, when a target participant holding a specific face property joins training dynamically at round t=250t = 250, the adversary's detection score jumps from near 00 to approximately 1.01.0 immediately upon entry.
  8. Knowl 8 — Empirical Evaluation of Defenses Against Feature Leakage in Collaborative Learning

    data/table

    Common heuristics and privacy mechanisms fail to provide robust protection against gradient-based feature and membership inference in collaborative learning without heavily degrading model performance:

    1. Sharing Fewer Gradients: Restricting participants to share only a random fraction of their parameter gradients per round reduces communication but preserves property leakage on the CLiPS Stylometry Investigation (CSI) Corpus:
    Target Property 10% Gradients Shared 50% Gradients Shared 100% Gradients Shared
    Top region (Antwerpen) 0.84 0.86 0.93
    Gender 0.90 0.91 0.93
    Veracity 0.94 0.99 0.99
    1. Vocabulary Restriction: Restricting vocabulary size for embedding layers reduces membership inference precision but severely degrades main task accuracy. On FourSquare, shrinking top location vocabulary from 30,000 to 1,000 reduces membership attack precision from 0.910.91 to 0.520.52, but drops model AUC from 0.640.64 to random guessing (0.500.50). On CSI, shrinking vocabulary from 4,000 to 500 lowers attack precision from 0.940.94 to 0.820.82 while reducing model AUC from 0.910.91 to 0.840.84.
    2. Dropout Regularization: Adding dropout (with probability pdropp_{\text{drop}}) after pooling layers increases attack AUC from 0.940.94 at pdrop=0.1p_{\text{drop}} = 0.1 to 0.990.99 at pdrop=0.9p_{\text{drop}} = 0.9 on CSI, because stochastic feature masking acts like feature bagging and increases the variance between participants' updates.
    3. Participant-Level Differential Privacy: Applying participant-level differentially private federated learning algorithms to small cohorts (K≤30K \le 30) injects sufficient Gaussian noise (via the moments accountant) to prevent model convergence entirely.
  9. Knowl 9 — Limitations of Collaborative Feature Leakage Attacks

    limitation

    Inference attacks exploiting unintended feature leakage in collaborative learning face several practical constraints:

    1. Auxiliary Data Requirements: Passive property inference requires auxiliary data labeled with the target property sampled from the same distribution or domain. Active property inference additionally requires auxiliary samples labeled for both the primary task and the target property.
    2. Participant Scaling: In synchronized SGD and federated averaging, observing only aggregated updates from KK participants causes inference accuracy to drop as KK increases, as individual contributions become obscured within the sum.
    3. Internal Feature Separability: The attack relies on the underlying neural network establishing separate internal representations for the target property. For properties that the model architecture does not internally separate during training (e.g., certain face identities under model averaging), property classifiers fail.
    4. Attribution Ambiguity: In multi-party settings where updates from multiple honest participants are aggregated, detecting that a property is present in the joint update does not identify which specific participant provided the data, unless complemented by external auxiliary metadata or temporal entry events.

Coverage note — None was omitted; all key attack methodologies, formal models, empirical results across datasets, defenses, and limitations are fully covered.

References

  1. 1.M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang. Deep learning with differential privacy. In CCS, 2016.
  2. 2.G. Ateniese, L. V. Mancini, A. Spognardi, A. Villani, D. Vitali, and G. Felici. Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers. IJSN, 10(3):137–150, 2015.
  3. 3.M. Backes, P. Berrang, M. Humbert, and P. Manoharan. Membership Privacy in MicroRNA-based Studies. In CCS, 2016.
  4. 4.BBC. Google DeepMind NHS app test broke UK privacy law. https://bbc.in/2A1MftK, 2017.
  5. 5.K. Bonawitz, V. Ivanov, B. Kreuter, A. Marcedone, H. B. McMahan, S. Patel, D. Ramage, A. Segal, and K. Seth. Practical secure aggregation for privacy-preserving machine learning. In CCS, 2017.
  6. 6.N. Carlini, C. Liu, J. Kos, Ú. Erlingsson, and D. Song. The Secret Sharer: Measuring unintended neural network memorization & extracting secrets. arXiv:1802.08232, 2018.
  7. 7.C.-H. Chang, L. Rampasek, and A. Goldenberg. Dropout feature ranking for deep learning models. arXiv:1712.08645, 2017.
  8. 8.T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang. MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv:1512.01274, 2015.
  9. 9.T. M. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman. Project Adam: Building an efficient and scalable deep learning training system. In OSDI, 2014.
  10. 10.K. Cho, B. Van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio. Learning phrase representations using RNN encoder-decoder for statistical machine translation. In EMNLP, 2014.
  11. 11.J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. Le, et al. Large scale distributed deep networks. In NIPS, 2012.
  12. 12.S. Dieleman, J. Schlüter, C. Raffel, et al. Lasagne: First release. http://dx.doi.org/10.5281/zenodo.27878, 2015.
  13. 13.C. Dwork and M. Naor. On the difficulties of disclosure prevention in statistical databases or the case for differential privacy. Journal of Privacy and Confidentiality, 2(1):93–107, 2010.
  14. 14.C. Dwork, A. Smith, T. Steinke, J. Ullman, and S. Vadhan. Robust traceability from trace amounts. In FOCS, 2015.
  15. 15.H. Edwards and A. Storkey. Censoring representations with an adversary. In ICLR, 2016.
  16. 16.M. Fredrikson, S. Jha, and T. Ristenpart. Model inversion attacks that exploit confidence information and basic countermeasures. In CCS, 2015.
  17. 17.M. Fredrikson, E. Lantz, S. Jha, S. Lin, D. Page, and T. Ristenpart. Privacy in pharmacogenetics: An end-to-end case study of personalized Warfarin dosing. In USENIX Security, 2014.
  18. 18.K. Ganju, Q. Wang, W. Yang, C. A. Gunter, and N. Borisov. Property inference attacks on fully connected neural networks using permutation invariant representations. In CCS, 2018.
  19. 19.General Data Protection Regulation. https://en.wikipedia.org/wiki/General_Data_Protection_Regulation, 2018.
  20. 20.R. C. Geyer, T. Klein, and M. Nabi. Differentially private federated learning: A client level perspective. arXiv:1712.07557, 2017.
  21. 21.I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio. Deep Learning. MIT Press, 2016.
  22. 22.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In NIPS, 2014.
  23. 23.J. Hamm, Y. Cao, and M. Belkin. Learning privately from multiparty data. In ICML, 2016.
  24. 24.J. Hayes, L. Melis, G. Danezis, and E. De Cristofaro. LOGAN: Membership inference attacks against generative models. In PETS, 2019.
  25. 25.B. Hitaj, G. Ateniese, and F. Pérez-Cruz. Deep models under the GAN: Information leakage from collaborative deep learning. In CCS, 2017.
  26. 26.T. K. Ho. Random decision forests. In DAR, 1995.
  27. 27.N. Homer, S. Szelinger, M. Redman, D. Duggan, W. Tembe, J. Muehling, J. V. Pearson, D. A. Stephan, S. F. Nelson, and D. W. Craig. Resolving individuals contributing trace amounts of DNA to highly complex mixtures using high-density SNP genotyping microarrays. PLoS Genetics, 4(8), 2008.
  28. 28.G. B. Huang, M. Ramesh, T. Berg, and E. Learned-Miller. Labeled faces in the wild: A database for studying face recognition in unconstrained environments. Technical Report 07–49, University of Massachusetts, Amherst, 2007.
  29. 29.A. Jochems, T. M. Deist, I. El Naqa, M. Kessler, C. Mayo, J. Reeves, S. Jolly, M. Matuszak, R. Ten Haken, J. van Soest, et al. Developing and validating a survival prediction model for NSCLC patients through distributed learning across 3 countries. Int J Radiat Oncol Biol Phys, 99(2):344–352, 2017.
  30. 30.A. Jochems, T. M. Deist, J. Van Soest, M. Eble, P. Bulens, P. Coucke, W. Dries, P. Lambin, and A. Dekker. Distributed learning: Developing a predictive model based on data from multiple hospitals without data leaving the hospital–a real life proof of concept. Radiother Oncol, 121(3):459–467, 2016.
  31. 31.Y. Kim. Convolutional neural networks for sentence classification. arXiv:1408.5882, 2014.
  32. 32.Y. LeCun, Y. Bengio, and G. Hinton. Deep learning. Nature, 2015.
  33. 33.Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally. Deep gradient compression: Reducing the communication bandwidth for distributed training. In ICLR, 2018.
  34. 34.Y. Long, V. Bindschaedler, L. Wang, D. Bu, X. Wang, H. Tang, C. A. Gunter, and K. Chen. Understanding membership inferences on well-generalized learning models. arXiv:1802.04889, 2018.
  35. 35.H. B. McMahan, E. Moore, D. Ramage, S. Hampson, et al. Communication-efficient learning of deep networks from decentralized data. In AISTATS, 2017.
  36. 36.H. B. McMahan, D. Ramage, K. Talwar, and L. Zhang. Learning differentially private language models without losing accuracy. In ICLR, 2018.
  37. 37.P. Mohassel and Y. Zhang. SecureML: A system for scalable privacy-preserving machine learning. In S&P, 2017.
  38. 38.P. Moritz, R. Nishihara, I. Stoica, and M. I. Jordan. SparkNet: Training deep networks in Spark. arXiv:1511.06051, 2015.
  39. 39.M. Nasr, R. Shokri, and A. Houmansadr. Machine learning with membership privacy using adversarial regularization. In CCS, 2018.
  40. 40.H.-W. Ng and S. Winkler. A data-driven approach to cleaning large face datasets. In ICIP, 2014.
  41. 41.S. J. Oh, M. Augustin, M. Fritz, and B. Schiele. Towards reverse-engineering black-box neural networks. In ICLR, 2018.
  42. 42.S. A. Osia, A. S. Shamsabadi, A. Taheri, K. Katevas, S. Sajadmanesh, H. R. Rabiee, N. D. Lane, and H. Haddadi. A hybrid deep learning architecture for privacy-preserving mobile analytics. arXiv:1703.02952, 2017.
  43. 43.S. A. Osia, A. Taheri, A. S. Shamsabadi, K. Katevas, H. Haddadi, and H. R. Rabiee. Deep private-feature extraction. TKDE, 2019.
  44. 44.J. Pang and Y. Zhang. DeepCity: A feature learning framework for mining location check-ins. In ICWSM (Poster Papers), 2017.
  45. 45.N. Papernot, M. Abadi, Ú. Erlingsson, I. Goodfellow, and K. Talwar. Semi-supervised knowledge transfer for deep learning from private training data. In ICLR, 2017.
  46. 46.N. Papernot, S. Song, I. Mironov, A. Raghunathan, K. Talwar, and Ú. Erlingsson. Scalable private learning with PATE. In ICLR, 2018.
  47. 47.M. Pathak, S. Rane, and B. Raj. Multiparty differential privacy via aggregation of locally trained classifiers. In NIPS, 2010.
  48. 48.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. JMLR, 12, 2011.
  49. 49.L. T. Phong, Y. Aono, T. Hayashi, L. Wang, and S. Moriai. Privacy-preserving deep learning: Revisited and enhanced. In ATIS, 2017.
  50. 50.A. Pyrgelis, C. Troncoso, and E. De Cristofaro. Knock knock, who’s there? membership inference on aggregate location data. In NDSS, 2018.
  51. 51.J. Schmidhuber. Deep learning in neural networks: An overview. Neural Networks, 2015.
  52. 52.R. Shokri and V. Shmatikov. Privacy-preserving deep learning. In CCS, 2015.
  53. 53.R. Shokri, M. Stronati, C. Song, and V. Shmatikov. Membership inference attacks against machine learning models. In S&P, 2017.
  54. 54.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015.
  55. 55.C. Song, T. Ristenpart, and V. Shmatikov. Machine learning models that remember too much. In CCS, 2017.
  56. 56.N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. JMLR, 2014.
  57. 57.F. Tramèr, F. Zhang, A. Juels, M. K. Reiter, and T. Ristenpart. Stealing machine learning models via prediction APIs. In USENIX Security, 2016.
  58. 58.S. Truex, L. Liu, M. E. Gursoy, L. Yu, and W. Wei. Towards demystifying membership inference attacks. arXiv:1807.09173, 2018.
  59. 59.L. van der Maaten and G. Hinton. Visualizing data using t-SNE. JMLR, 2008.
  60. 60.B. Verhoeven and W. Daelemans. CLiPS Stylometry Investigation (CSI) Corpus: A Dutch corpus for the detection of age, gender, personality, sentiment and deception in text. In LREC, 2014.
  61. 61.B. Wang and N. Z. Gong. Stealing hyperparameters in machine learning. In S&P, 2018.
  62. 62.E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Yu. Petuum: A new platform for distributed machine learning on big data. IEEE Transactions on Big Data, 2015.
  63. 63.D. Yang, D. Zhang, L. Chen, and B. Qu. NationTelescope: Monitoring and visualizing large-scale collective behavior in LBSNs. JNCA, 55:170–180, 2015.
  64. 64.D. Yang, D. Zhang, and B. Qu. Participatory cultural mapping based on collective behavior in location based social networks. ACM TIST, 2015.
  65. 65.D. Yang, D. Zhang, B. Qu, and P. Cudre-Mauroux. PrivCheck: Privacy-preserving check-in data publishing for personalized location based services. In UbiComp, 2016.
  66. 66.S. Yeom, I. Giacomelli, M. Fredrikson, and S. Jha. Privacy risk in machine learning: Analyzing the connection to overfitting. In CSF, 2018.
  67. 67.N. Zhang, M. Paluri, Y. Taigman, R. Fergus, and L. Bourdev. Beyond frontal faces: Improving person recognition using multiple cues. In CVPR, 2015.
  68. 68.M. Zinkevich, M. Weimer, L. Li, and A. J. Smola. Parallelized stochastic gradient descent. In NIPS, 2010.

Citation

MLA
Melis, L., et al. “Exploiting Unintended Feature Leakage in Collaborative Learning”. arXiv, 2018, http://arxiv.org/abs/1805.04049v3.
APA
Melis, L., Song, C., Cristofaro, E. D., & Shmatikov, V. (2018). Exploiting Unintended Feature Leakage in Collaborative Learning. arXiv. http://arxiv.org/abs/1805.04049v3
Chicago
Melis, L., C. Song, E. D. Cristofaro, and V. Shmatikov. 2018. “Exploiting Unintended Feature Leakage in Collaborative Learning”. arXiv. http://arxiv.org/abs/1805.04049v3.
Harvard
Melis, L. et al. (2018) “Exploiting Unintended Feature Leakage in Collaborative Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1805.04049v3.
Vancouver
1. Melis L, Song C, Cristofaro ED, Shmatikov V (2018) Exploiting Unintended Feature Leakage in Collaborative Learning. arXiv

BibTeX

@article{melis2018exploiting,
  title = {Exploiting Unintended Feature Leakage in Collaborative Learning},
  author = {Melis, Luca and Song, Congzheng and Cristofaro, Emiliano De and Shmatikov, Vitaly},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1805.04049v3},
  eprint = {1805.04049}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF