Hyperparameter Tuning with Renyi Differential Privacy
Nicolas PapernotThomas Steinke
Establishes Rényi Differential Privacy bounds for hyperparameter search across multiple training runs, proving that cumulative privacy leakage remains modest when each candidate model is individually private.
Machine learning systems frequently retain sensitive training data, making differential privacy—a mathematical standard that prevents the leakage of individual records—critical for deployment in domains such as healthcare. Standard practice involves tuning hyperparameters (such as learning rates or regularization weights) across multiple training runs to optimize accuracy. However, practitioners often tune these configurations without privacy protections and only apply differential privacy to the final training run, assuming that hyperparameter settings reveal minimal information.
The article evaluates the privacy risks inherent in hyperparameter tuning and demonstrates how to rigorously preserve privacy across multiple training repetitions within the Rényi differential privacy framework.
To establish this, the article first demonstrates how non-private hyperparameter tuning leaks information using a synthetic classification experiment where outliers systematically distort optimal settings. It then develops a theoretical framework for private random hyperparameter search, where the number of training repetitions is randomized rather than fixed. The authors evaluate candidate probability distributions for these run counts—specifically truncated negative binomial, logarithmic, and Poisson distributions—and validate the approach through 500 trials fine-tuning a convolutional neural network on the benchmark MNIST dataset using differentially private stochastic gradient descent.
The investigation yields four key findings. First, tuning hyperparameters non-privately creates an exploitable channel for membership inference attacks, allowing adversaries to detect whether specific outlier records exist in the training set. Second, repeating a private base algorithm a fixed number of times causes privacy loss to accumulate linearly with the number of runs, offering no theoretical improvement over standard composition. Third, drawing the number of runs randomly reduces this privacy penalty dramatically, yielding bounds that grow logarithmically or remain independent of the repetition count. Fourth, the Poisson distribution delivers the superior privacy-utility balance in practical settings, avoiding the severe performance pitfalls of heavy-tailed distributions while maintaining tightly concentrated runtimes.
These findings demonstrate that organizations cannot treat hyperparameter selection as a separate, privacy-free step without undermining the compliance and confidentiality guarantees of the final model. Although private tuning increases the overall privacy parameter by a factor of roughly two to three relative to a single run, randomizing the repetition count prevents the severe linear degradation that would otherwise render extensive model search computationally or mathematically impermissible.
Organizations training machine learning models on sensitive data should integrate differential privacy into the hyperparameter tuning phase rather than applying it solely to final models. Teams should employ random search with Poisson-distributed run counts when operating under constrained privacy budgets. For broader reporting, practitioners should transparently disclose both the baseline privacy parameter of individual training runs and the cumulative privacy parameter of the overall hyperparameter search procedure.
The analysis carries several boundary conditions. The current framework applies strictly to non-adaptive hyperparameter selection (such as random search) and does not cover adaptive optimization techniques like Bayesian optimization. Additionally, empirical validation remains focused on synthetic data and benchmark image classification tasks. Confidence in the underlying theoretical guarantees remains exceptionally high, as the authors provide formal proofs showing that the derived privacy bounds are mathematically tight up to low-order terms.
- Paper: Deep Learning with Differential Privacy, Martín Abadi et al. (2016). It introduces the foundational DP-SGD framework and moment-accounting privacy analysis that the source uses as the baseline private training mechanism to be tuned.
- Paper: Differentially Private Empirical Risk Minimization, Kamalika Chaudhuri et al. (2009). It formalizes differential privacy within empirical risk minimization and examines early procedures for private model and parameter selection.
- Paper: On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation, Gavin C. Cawley et al. (2010). It details how hyperparameter selection over finite validation sets leads to overfitting and information leakage during model selection.
- Paper: Practical Bayesian Optimization of Machine Learning Algorithms, Jasper Snoek et al. (2012). It provides standard principles and algorithms for iterative hyperparameter tuning in machine learning models.
- Paper: What Can We Learn Privately?, Shiva Prasad Kasiviswanathan et al. (2008). It establishes fundamental sample complexity and theoretical feasibility bounds for learning statistical models under differential privacy constraints.
No sufficiently relevant recommendations were found.
