Label Leakage and Protection in Two-party Split Learning
Oscar LiJiankai SunXin YangWeihao GaoHongyi ZhangJunyuan XieVirginia SmithChong Wang
Exposes how adversaries can reconstruct private labels in two-party split learning and develops Marvell, an optimized perturbation defense that minimizes worst-case label leakage while preserving model utility.
Modern privacy regulations and competitive concerns make it critical for organizations to collaborate on machine learning without sharing proprietary raw data or sensitive target labels, such as user purchase conversions or medical diagnoses. Two-party split learning is widely viewed as a privacy-preserving framework because organizations split a deep neural network across intermediate model layers and exchange only intermediate outputs and gradients rather than raw records. However, because each transmitted gradient corresponds to an aligned, specific data record, standard aggregate privacy defenses like differential privacy cannot be applied directly. This raises an urgent question regarding whether communicating cut-layer gradients unintentionally reveals the proprietary target labels.
The article evaluates the vulnerability of two-party split learning to private label theft during model training, quantifies this privacy loss, and develops principled noise-perturbation defenses to safeguard sensitive labels without degrading model utility.
To investigate this issue, the authors formalize a threat model involving an honest-but-curious non-label party seeking to reconstruct the other party's binary labels from cut-layer gradients. They introduce a privacy loss metric termed leak AUC, which measures an adversary's classification performance across all decision thresholds, where a value of 1.0 indicates total label leakage and 0.5 denotes random guessing. The authors demonstrate two realistic attack methods based on gradient length and direction. To counter these threats, they formulate a defense named Marvell, which mathematically optimizes class-specific Gaussian noise perturbations to minimize worst-case label leakage under strict optimization power constraints. The authors evaluate these attacks and defenses across three large-scale benchmark datasets: online advertising datasets Criteo and Avazu, and the SIIM-ISIC skin lesion image dataset.
The findings demonstrate that split learning without active defenses suffers from severe privacy leakage. In unperturbed training across all datasets, both the gradient norm and gradient direction scoring attacks consistently achieve leak AUC values near 1.0, enabling the non-label party to fully reconstruct private ground-truth labels at both the cut layer and internal network layers. Standard isotropic Gaussian noise fails to defend against directional attacks, leaving leak AUC values above 0.9 on the image dataset even when high noise is added. In contrast, Marvell reduces worst-case leak AUC close to the baseline level of 0.5 across both cut and earlier network layers while preserving predictive accuracy. An alternative heuristic defense, max_norm, also effectively mitigates the identified attacks but lacks the mathematical flexibility to tune privacy-utility tradeoffs.
These results demonstrate that standard split learning provides an illusion of label privacy. Gradient exchanges create substantial compliance, security, and competitive risks for institutions handling proprietary outcomes. Marvell demonstrates that mathematically structured noise—aligning perturbations with the difference between positive and negative class gradient distributions—effectively eliminates label reconstruction risks while maintaining viable model utility and generalization performance on real-world tasks.
Organizations deploying split learning in production should not rely on raw gradient exchanges and should implement structured perturbation defenses. Marvell is recommended when organizations require explicit, tunable control over the balance between privacy protection and predictive performance. For simpler implementations where hyperparameter tuning is impractical, the max_norm heuristic offers an effective out-of-the-box alternative against standard geometric attacks. Future development should evaluate defenses in multi-class classification, multi-party federated architectures, and against potential multi-step attacks where adversaries track historical gradients over multiple update iterations.
The conclusions are supported by theoretical bounds and extensive empirical validations across tabular and image domains. However, decision-makers should recognize that the optimization framework assumes class-conditional gradients approximate Gaussian distributions and focuses on binary classification. Confidence in the defense is high for standard single-iteration attacks, but additional testing is warranted before deploying in non-binary or multi-party collaborative environments.
- Paper: Deep Leakage from Gradients, Ligeng Zhu et al. (2019). This seminal paper demonstrates how private training inputs and labels can be mathematically inverted and reconstructed from shared gradients in collaborative learning architectures.
- Paper: Inverting Gradients - How easy is it to break privacy in federated learning?, Jonas Geiping et al. (2020). This work establishes theoretical bounds and numerical methods showing that deep network layer inputs can be reconstructed strictly from parameter gradient updates.
- Paper: Exploiting Unintended Feature Leakage in Collaborative Learning, Luca Melis et al. (2018). This study analyzes passive and active feature and property leakage stemming from intermediate gradient exchanges in collaborative multi-party learning.
- Paper: Deep Learning with Differential Privacy, Martín Abadi et al. (2016). This foundational work introduces gradient perturbation mechanisms and privacy budgeting via differentially private stochastic gradient descent.
- Paper: Federated Machine Learning, Qiang Yang et al. (2019). This paper establishes the taxonomy and architecture of vertical and split federated learning across organizations with partitioned features and labels.
- Paper: Differentially Private Empirical Risk Minimization, Kamalika Chaudhuri et al. (2009). This study formalizes objective and output perturbation strategies for differential privacy in risk minimization, grounding noise optimization under utility constraints.
- Paper: Hyperparameter Tuning with Renyi Differential Privacy, Nicolas Papernot et al. (2022). This work extends rigorous differential privacy protection frameworks into the hyperparameter tuning phase of machine learning training.
- Paper: Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models, Haonan Duan et al. (2023). This paper extends differential privacy and gradient perturbation principles to prevent private prompt data leakage in large language models.
