Why Is Public Pretraining Necessary for Private Model Training?
Arun GaneshMahdi HaghifamMilad NasrSewoong OhThomas SteinkeOm ThakkarAbhradeep Guha ThakurtaLun Wang
Explains why public pretraining is essential for differentially private learning by theoretically proving and empirically demonstrating that noiseless early optimization is required to select viable basins before fine-tuning on sensitive data.
Differential privacy provides mathematical guarantees that machine learning models will not leak sensitive user training data. However, training models strictly from scratch on private data causes a severe drop in accuracy and utility compared to standard non-private training. While pretraining models on non-sensitive, public data before privately fine-tuning them dramatically restores performance, this gain is substantially larger than what traditional transfer learning explains. The article investigates why public pretraining is so critical to private model training and demonstrates what mechanisms drive this performance gap.
The article evaluates the hypothesis that non-convex machine learning optimization operates in two distinct phases: first, navigating the complex loss landscape to select a favorable low-loss region or "basin," and second, fine-tuning within that basin to reach the minimum. Choosing the initial basin requires strong, clear gradient signals that are easily corrupted by the noise added for differential privacy, making private optimization from scratch fail. In contrast, local optimization within an established basin requires much less noise-sensitive exploration, allowing private data to be used effectively once the basin is found.
The authors analyze this problem through both formal mathematical constructions and empirical experiments on standard vision and speech benchmarks. Theoretically, they construct synthetic data distributions and loss functions to analyze sample efficiency with and without public pretraining under both in-distribution and out-of-distribution settings. Empirically, they run private gradient descent experiments on CIFAR-10 image classification (comparing in-distribution and out-of-distribution public pretraining against public post-training) and evaluate the geometry of loss landscapes on a Conformer speech recognition model trained on Librispeech.
The key findings show that early optimization steps are the most critical for model utility and the most sensitive to privacy noise. In image classification experiments, allocating a fixed public data budget entirely to pretraining yielded substantially higher accuracy than allocating it to post-training or distributing privacy budgets evenly. Furthermore, loss landscape interpolations on speech models revealed that a publicly pretrained and privately fine-tuned model settled in the exact same favorable basin as a model trained fully without privacy constraints. Conversely, models trained purely on private data from scratch ended up trapped in an inferior basin separated by high-loss barriers. Theoretically, the authors prove that a mixed strategy using public pretraining followed by private fine-tuning achieves low error, whereas using either dataset alone fails to achieve the target performance under standard sample limits.
These findings indicate that organizations training privacy-preserving models must prioritize reducing noise during the earliest phase of training rather than throughout the entire training lifecycle. Public data, even when limited in volume or drawn from an out-of-distribution source, acts as an essential compass to guide models past chaotic initial loss landscapes. Practitioners should leverage public pretraining wherever possible rather than relying on architectural workarounds or post-processing adjustments during private fine-tuning. However, practitioners must also be judicious, as public data must be audited to prevent unintended memorization or copyright and privacy issues.
The article notes that while its empirical findings hold across standard deep learning architectures, its formal mathematical impossibility bounds are established on synthetic constructions. Confidence in the underlying two-phase mechanism is high due to consistent experimental and geometric evidence across image and speech domains. Future work should focus on closing the gap between synthetic theoretical bounds and general deep learning landscapes, as well as establishing formal optimization frameworks when public datasets exhibit extreme distribution shifts from target private tasks.
- Paper: Deep Learning with Differential Privacy, Martín Abadi et al. (2016). Its differentially private SGD and privacy-accounting framework provide the core training and privacy machinery that the source investigates under pretraining.
- Paper: Why Does Unsupervised Pre-training Help Deep Learning?, Dumitru Erhan et al. (2010). Its analysis of how unsupervised pretraining steers deep networks through non-convex optimization landscapes prepares you for the source’s basin-selection hypothesis.
- Paper: What Can We Learn Privately?, Shiva Prasad Kasiviswanathan et al. (2008). Its private-learning sample-complexity framework supplies theoretical context for the source’s comparison of private-only and hybrid public-private training.
No sufficiently relevant recommendations were found.
