Why Is Public Pretraining Necessary for Private Model Training?

Arun GaneshMahdi HaghifamMilad NasrSewoong OhThomas SteinkeOm ThakkarAbhradeep Guha ThakurtaLun Wang

article2023ICML46 citations

Explains why public pretraining is essential for differentially private learning by theoretically proving and empirically demonstrating that noiseless early optimization is required to select viable basins before fine-tuning on sensitive data.

Listen

Differential privacy provides mathematical guarantees that machine learning models will not leak sensitive user training data. However, training models strictly from scratch on private data causes a severe drop in accuracy and utility compared to standard non-private training. While pretraining models on non-sensitive, public data before privately fine-tuning them dramatically restores performance, this gain is substantially larger than what traditional transfer learning explains. The article investigates why public pretraining is so critical to private model training and demonstrates what mechanisms drive this performance gap.

The article evaluates the hypothesis that non-convex machine learning optimization operates in two distinct phases: first, navigating the complex loss landscape to select a favorable low-loss region or "basin," and second, fine-tuning within that basin to reach the minimum. Choosing the initial basin requires strong, clear gradient signals that are easily corrupted by the noise added for differential privacy, making private optimization from scratch fail. In contrast, local optimization within an established basin requires much less noise-sensitive exploration, allowing private data to be used effectively once the basin is found.

The authors analyze this problem through both formal mathematical constructions and empirical experiments on standard vision and speech benchmarks. Theoretically, they construct synthetic data distributions and loss functions to analyze sample efficiency with and without public pretraining under both in-distribution and out-of-distribution settings. Empirically, they run private gradient descent experiments on CIFAR-10 image classification (comparing in-distribution and out-of-distribution public pretraining against public post-training) and evaluate the geometry of loss landscapes on a Conformer speech recognition model trained on Librispeech.

The key findings show that early optimization steps are the most critical for model utility and the most sensitive to privacy noise. In image classification experiments, allocating a fixed public data budget entirely to pretraining yielded substantially higher accuracy than allocating it to post-training or distributing privacy budgets evenly. Furthermore, loss landscape interpolations on speech models revealed that a publicly pretrained and privately fine-tuned model settled in the exact same favorable basin as a model trained fully without privacy constraints. Conversely, models trained purely on private data from scratch ended up trapped in an inferior basin separated by high-loss barriers. Theoretically, the authors prove that a mixed strategy using public pretraining followed by private fine-tuning achieves low error, whereas using either dataset alone fails to achieve the target performance under standard sample limits.

These findings indicate that organizations training privacy-preserving models must prioritize reducing noise during the earliest phase of training rather than throughout the entire training lifecycle. Public data, even when limited in volume or drawn from an out-of-distribution source, acts as an essential compass to guide models past chaotic initial loss landscapes. Practitioners should leverage public pretraining wherever possible rather than relying on architectural workarounds or post-processing adjustments during private fine-tuning. However, practitioners must also be judicious, as public data must be audited to prevent unintended memorization or copyright and privacy issues.

The article notes that while its empirical findings hold across standard deep learning architectures, its formal mathematical impossibility bounds are established on synthetic constructions. Confidence in the underlying two-phase mechanism is high due to consistent experimental and geometric evidence across image and speech domains. Future work should focus on closing the gap between synthetic theoretical bounds and general deep learning landscapes, as well as establishing formal optimization frameworks when public datasets exhibit extreme distribution shifts from target private tasks.

arXiv: 2302.09483
  • Paper: Deep Learning with Differential Privacy, Martín Abadi et al. (2016). Its differentially private SGD and privacy-accounting framework provide the core training and privacy machinery that the source investigates under pretraining.
  • Paper: Why Does Unsupervised Pre-training Help Deep Learning?, Dumitru Erhan et al. (2010). Its analysis of how unsupervised pretraining steers deep networks through non-convex optimization landscapes prepares you for the source’s basin-selection hypothesis.
  • Paper: What Can We Learn Privately?, Shiva Prasad Kasiviswanathan et al. (2008). Its private-learning sample-complexity framework supplies theoretical context for the source’s comparison of private-only and hybrid public-private training.

No sufficiently relevant recommendations were found.

Cover for Why Is Public Pretraining Necessary for Private Model Training?

Abstract

In the privacy-utility tradeoff of a model trained on benchmark language and vision tasks, remarkable improvements have been widely reported when the model is pretrained on public data. Some gain is expected as these models inherit the benefits of transfer learning, which is the standard motivation in non-private settings. However, the stark contrast in the gain of pretraining between non-private and private machine learning suggests that the gain in the latter is rooted in a fundamentally different cause. To explain this phenomenon, we hypothesize that the non-convex loss landscape of a model training necessitates the optimization algorithm to go through two phases. In the first, the algorithm needs to select a good “basin” in the loss landscape. In the second, the algorithm solves an easy optimization within that basin. The former is a harder problem to solve with private data, while the latter is harder to solve with public data due to a distribution shift or data scarcity. Guided by this intuition, we provide theoretical constructions that provably demonstrate the separation between private training with and without public pretraining. Further, systematic experiments on CIFAR10 and Librispeech provide supporting evidence for our hypothesis.

Table of Contents

  • 1 Introduction
  • 1.1 Other Related Work
  • 1.2 Background on differential privacy and DP-SCO
  • 2 Necessity of public pretraining
  • 2.1 Pretraining on in-distribution public data
  • 2.2 Construction
  • 2.3 Analysis
  • 2.4 Pretraining on out-of-distribution public data
  • 3 Experiments
  • 3.1 CIFAR10 Experiments
  • 3.2 Manifold on Large Speech Model
  • 4 Discussion
  • References
  • A Notation Reference
  • B Survey of the gain of pretraining
  • C Missing Details from Section
  • C.1 Proof of Theorem
  • D Quadratic Example

Knowls

  1. Knowl 1 — Public pretraining and private fine-tuning can outperform either data source alone

    theoretical result

    For every integer p≥1p\geq 1, the paper constructs an unconstrained stochastic optimization problem over Rp4×Rp\mathbb{R}^{p^4}\times\mathbb{R}^{p} and a family of data distributions with the following separation. There are pp public samples and p2p^2 private samples, both drawn from the same distribution. For δ=o(p−2)\delta=o(p^{-2}), any (1,δ)(1,\delta)-differentially private algorithm using only the private samples and any non-private algorithm using only the public samples each have worst-case expected excess population loss Ω(1)\Omega(1) over the constructed distribution family. In contrast, for δ≥2−p\delta\geq 2^{-p}, an algorithm that first takes a gradient-descent step using the public samples and then runs (1,δ)(1,\delta)-DP-SGD on the private samples achieves expected excess population loss O(1/p)O(1/p) for every distribution in the family. Here excess population loss is the expected population loss of the algorithm’s output minus the population minimum. The result is an existence theorem for a synthetic task, not a general lower bound for deep-learning problems.

  2. Knowl 2 — The separation persists with arbitrarily abundant out-of-distribution public data

    theoretical result

    For every integer p≥1p\geq 1 and every public sample count npub≥pn_{\mathrm{pub}}\geq p, the paper constructs pairs of public and private distributions on Rp4×Rp\mathbb{R}^{p^4}\times\mathbb{R}^{p}. The public distribution has the form τ1×δ0\tau_1\times\delta_0, where δ0\delta_0 places the second data component at zero; the private distribution is τ1×τ2\tau_1\times\tau_2. Thus the public and private data share the first component but the public data reveal nothing about the private distribution of the second component. Evaluating all algorithms on the private population loss, the construction makes the worst-case expected excess loss Ω(1)\Omega(1) both for a private-only (1,δ)(1,\delta)-DP algorithm using p2p^2 private samples when δ=o(p−2)\delta=o(p^{-2}) and for a non-private algorithm using only the npubn_{\mathrm{pub}} public samples. Yet, for δ≥2−p\delta\geq 2^{-p}, public gradient descent followed by (1,δ)(1,\delta)-DP-SGD on p2p^2 private samples achieves expected excess loss O(1/p)O(1/p). The lower bound therefore applies even when the amount of public data is arbitrarily large, because its distribution does not identify the private task’s second component.

  3. Knowl 3 — A gated two-component loss creates a basin-selection phase followed by local optimization

    model/method

    The paper constructs a non-convex loss on parameters θ=(θ1,θ2)∈Rp4×Rp\theta=(\theta_1,\theta_2)\in\mathbb{R}^{p^4}\times\mathbb{R}^{p} and data d=(d1,d2)d=(d_1,d_2), where d1∈Bp4(0,1)d_1\in B_{p^4}(0,1) and d2∈Bp(0,r)d_2\in B_p(0,r). Its per-example loss is

    ℓ((θ1,θ2);(d1,d2))=ℓ1(θ1;d1)+p q(θ1)ℓ2(θ2;d2),\ell((\theta_1,\theta_2);(d_1,d_2))=\ell_1(\theta_1;d_1)+p\,q(\theta_1)\ell_2(\theta_2;d_2),

    with ℓ1(θ1;d1)=min⁡{12∥θ1−d1∥22,92}\ell_1(\theta_1;d_1)=\min\{\tfrac12\|\theta_1-d_1\|_2^2,\tfrac92\} and ℓ2(θ2;d2)=min⁡{0,∥θ2−d2∥22/(2r2)−92}\ell_2(\theta_2;d_2)=\min\{0,\|\theta_2-d_2\|_2^2/(2r^2)-\tfrac92\}. The second term is non-positive and is activated near a shell-shaped basin in the θ1\theta_1 coordinates. Specifically, let S=Bp4(0,M+R1)∖Bp4(0,M−R1)S=B_{p^4}(0,M+R_1)\setminus B_{p^4}(0,M-R_1), where MM is the radius of the hard private-mean-estimation instances used in the construction and R1R_1 is chosen so the relevant minimizers lie inside the shell. Let Δ(θ1)=∥θ1−ΠS(θ1)∥2\Delta(\theta_1)=\|\theta_1-\Pi_S(\theta_1)\|_2, with ΠS\Pi_S the Euclidean projection onto SS, and choose R2<M−R1R_2<M-R_1. Set q(θ1)=0q(\theta_1)=0 when Δ(θ1)>R2\Delta(\theta_1)>R_2, q(θ1)=1−Δ(θ1)/R2q(\theta_1)=1-\Delta(\theta_1)/R_2 when 0<Δ(θ1)≤R20<\Delta(\theta_1)\leq R_2, and q(θ1)=1q(\theta_1)=1 when Δ(θ1)=0\Delta(\theta_1)=0. Outside the shell, optimization acts only on the high-dimensional θ1\theta_1 problem; inside it, the low-dimensional θ2\theta_2 problem becomes active. The construction uses the first stage to make basin selection difficult privately and the second to make accurate local optimization difficult with the small public sample.

  4. Knowl 4 — Two-phase optimization is the proposed explanation for pretraining’s privacy benefit

    model/method

    The paper’s hypothesis is that gradient-based training of typical non-convex machine-learning losses has two functionally different phases. First, training must use data to select a favorable basin among multiple local minima; second, it must optimize within the selected basin. The authors propose that private noise is especially damaging during basin selection: when the useful gradient directing training toward a good basin is small, noise can send the optimization toward a worse basin, so this phase may require substantially more samples privately than non-privately. Once a good basin has been selected, the remaining optimization is treated as effectively locally convex and may be more amenable to private fine-tuning. Conversely, public data may be inadequate for this later phase because of scarcity or distribution shift. This is a hypothesis motivated by the paper’s constructions and experiments, not a claim proved for all deep-learning loss landscapes.

  5. Knowl 5 — A constrained quadratic construction also exhibits the separation

    theoretical result

    The paper extends its separation to a quadratic loss when optimization is constrained to a product of Euclidean balls. In a construction with parameter and data components (θ1,θ2)(\theta_1,\theta_2) and (d1,d2)(d_1,d_2), the loss has the form

    ℓ((θ1,θ2);(d1,d2))=12∥θ1−d1∥22+p2r2∥θ2−d2∥22,\ell((\theta_1,\theta_2);(d_1,d_2))=\tfrac12\|\theta_1-d_1\|_2^2+\frac{p}{2r^2}\|\theta_2-d_2\|_2^2,

    where θ1,d1∈Rp4\theta_1,d_1\in\mathbb{R}^{p^4}, θ2,d2∈Rp\theta_2,d_2\in\mathbb{R}^{p}, and rr is chosen small. With pp public examples and p2p^2 private examples, the theorem gives worst-case excess loss Ω(1)\Omega(1) for public-only non-private training and for private-only (1,δ)(1,\delta)-DP training when δ=o(p−2)\delta=o(p^{-2}), while projected gradient descent on the public examples followed by projected DP-SGD on the private examples achieves O(1/p)O(1/p) excess loss for δ≥2−p\delta\geq 2^{-p}. The constrained setting is important: the construction controls parameter movement during public optimization. The paper notes that an unconstrained adaptation is possible but may be inefficient, because the small scale rr can require Ω(1/r2)\Omega(1/r^2) public gradient steps.

  6. Knowl 6 — CIFAR-10 experiments favor allocating public training to the beginning

    empirical result

    In CIFAR-10 experiments, a ConvNet was trained with DP-SGD for 60 epochs using clipping norm 1, learning rate 0.001, batch size 256, and Adam. For the in-distribution public-data condition, the 60,000-image dataset was split into 2,000 public and 58,000 private images; public training used Adam with learning rate 0.002. The experiment allotted a total of 200 public-training epochs and varied how many were used before rather than after private training. Accuracy was highest when the public epochs were used for pretraining, with the gains from additional pretraining becoming small after about 100 epochs. For the out-of-distribution condition, 20,000 training images from CINIC-10 served as public data, with a 30-epoch public-training budget; this condition also favored pretraining over post-training. These comparisons support the claim that the early training rounds are particularly sensitive to privacy noise.

  7. Knowl 7 — Reducing privacy noise in the first CIFAR-10 iteration improves accuracy

    data/table

    The CIFAR-10 experiment compared spending more privacy budget (using a lower noise multiplier) on the first private-training iteration with spending it on the last iteration, at the same end-to-end privacy level. The reported test accuracies were:

    ε\varepsilon First epoch Last epoch
    1 46.7%±0.346.7\%\pm 0.3 46.3%±0.346.3\%\pm 0.3
    3 49.6%±0.649.6\%\pm 0.6 48.0%±0.548.0\%\pm 0.5
    8 54.0%±0.854.0\%\pm 0.8 52.0%±0.952.0\%\pm 0.9

    The comparison used σ2=0.6\sigma^2=0.6 for the iteration receiving the higher budget. Accuracy was higher when the first epoch received that budget for each tested ε\varepsilon, supporting the experimental conclusion that early optimization is more sensitive to privacy noise than late optimization.

  8. Knowl 8 — Publicly pretrained Librispeech model occupies the oracle model’s basin

    empirical result

    The paper compared three ConformerM models trained on Librispeech: an oracle reference trained on the complete dataset for 100,000 steps; a private model trained with DP-Adam on a uniform 90% sample for 20,000 steps; and a privately fine-tuned model pretrained with Adam for 10,000 steps on the other 10% and then privately trained on the 90% for 1,000 steps. Hyperparameters for the latter two models were tuned for test word error rate under the same privacy budget, ε=9.8\varepsilon=9.8 and δ=10−6\delta=10^{-6}. The authors linearly interpolated the three sets of weights and evaluated RNNT loss on a 128-example subset of Librispeech’s test-other set. The interpolated landscape placed the publicly pretrained and privately fine-tuned model in the same basin as the oracle reference, while the fully private model lay in a different basin separated by a loss hill. The ℓ2\ell_2 weight distance from the oracle was 671.22 for the pretrained-and-fine-tuned model and 1738.27 for the fully private model. Changing pretraining to 1% of the training samples or to 5,000 pretraining steps produced similar qualitative landscapes.

  9. Knowl 9 — The theoretical necessity result is limited to constructed tasks

    limitation

    The paper proves the need for public pretraining only on carefully constructed synthetic optimization problems. Its experiments provide evidence for the proposed basin-selection explanation on CIFAR-10 and Librispeech, but do not establish that the theoretical separation holds for deep-learning problems generally. The authors identify closing the gap between these constructions and standard benchmark loss landscapes as an open research direction.

Coverage note — No substantial contributed material was deliberately omitted; proof details and the survey of prior public-pretraining results were excluded because they are not standalone contributions of this paper.

References

  1. 1.Abadi, M., Chu, A., Goodfellow, I. J., McMahan, H. B., Mironov, I., Talwar, K., and Zhang, L. Deep learning with differential privacy. In Proc. of the 2016 ACM SIGSAC Conf. on Computer and Communications Security (CCS’16), pp. 308–318, 2016.
  2. 2.Alon, N., Bassily, R., and Moran, S. Limits of private learning with access to public data. Advances in neural information processing systems, 32, 2019.
  3. 3.Amid, E., Ganesh, A., Mathews, R., Ramaswamy, S., Song, S., Steinke, T., Suriyakumar, V. M., Thakkar, O., and Thakurta, A. Public data-assisted mirror descent for private model training. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pp. 517–535. PMLR, 2022. URL https://proceedings.mlr.press/v162/amid22a.html.
  4. 4.Asi, H., Duchi, J., Fallah, A., Javidbakht, O., and Talwar, K. Private adaptive gradient methods for convex optimization. In International Conference on Machine Learning, pp. 383–392. PMLR, 2021a.
  5. 5.Asi, H., Levy, D. A. N., and Duchi, J. Adapting to function difficulty and growth conditions in private optimization. In Advances in Neural Information Processing Systems, 2021b.
  6. 6.Avent, B., Dubey, Y., and Korolova, A. The power of the hybrid model for mean estimation. Proceedings on Privacy Enhancing Technologies, 2020:48–68, 10 2020.
  7. 7.Bassily, R., Smith, A., and Thakurta, A. Private empirical risk minimization: Efficient algorithms and tight error bounds. In Proc. of the 2014 IEEE 55th Annual Symp. on Foundations of Computer Science (FOCS), pp. 464–473, 2014.
  8. 8.Bassily, R., Feldman, V., Talwar, K., and Thakurta, A. G. Private stochastic convex optimization with optimal rates. In Advances in Neural Information Processing Systems 32, pp. 11279–11288, 2019.
  9. 9.Bassily, R., Cheu, A., Moran, S., Nikolov, A., Ullman, J., and Wu, S. Private query release assisted by public data. In III, H. D. and Singh, A. (eds.), Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp. 695–703. PMLR, 13–18 Jul 2020a. URL https://proceedings.mlr.press/v119/bassily20a.html.
  10. 10.Bassily, R., Feldman, V., Guzman, C., and Talwar, K. Stability of stochastic gradient descent on nonsmooth convex losses. Advances in Neural Information Processing Systems, 33:4381–4391, 2020b.
  11. 11.Bie, A., Kamath, G., and Singhal, V. Private estimation with public data. Advances in neural information processing systems 35, 2022.
  12. 12.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  13. 13.Bozinovski, S. and Fulgosi, A. The influence of pattern similarity and transfer learning upon training of a base perceptron b2. In Proceedings of Symposium Informatica, volume 3, pp. 121–126, 1976.
  14. 14.Bu, Z., Wang, Y.-X., Zha, S., and Karypis, G. Differentially private bias-term only fine-tuning of foundation models. In Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS 2022.
  15. 15.Bu, Z., Wang, Y.-X., Zha, S., and Karypis, G. Differentially private optimization on large model at small cost. arXiv preprint arXiv:2210.00038, 2022.
  16. 16.De, S., Berrada, L., Hayes, J., Smith, S. L., and Balle, B. Unlocking high-accuracy differentially private image classification through scale. arXiv preprint arXiv:2204.13650, 2022.
  17. 17.Dwork, C., McSherry, F., Nissim, K., and Smith, A. Calibrating noise to sensitivity in private data analysis. In Proc. of the Third Conf. on Theory of Cryptography (TCC), pp. 265–284, 2006. URL http://dx.doi.org/10.1007/11681878_14.
  18. 18.Feldman, V., Koren, T., and Talwar, K. Private stochastic convex optimization: Optimal rates in linear time. In Proc. of the Fifty-Second ACM Symp. on Theory of Computing (STOC’20), 2020.
  19. 19.Ganesh, A., Thakurta, A., and Upadhyay, J. On the universality of langevin diffusion for private euclidean (convex) optimization, 2022. URL https://arxiv.org/abs/2204.01585. https://openreview.net/forum?id=ZrJPdY5k6sg.
  20. 20.Ginart, A., van der Maaten, L., Zou, J., and Guo, C. Submix: Practical private prediction for large-scale language models. arXiv preprint arXiv:2201.00971, 2022.
  21. 21.Golatkar, A., Achille, A., Wang, Y.-X., Roth, A., Kearns, M., and Soatto, S. Mixed differential privacy in computer vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8376–8386, 2022.
  22. 22.Gopi, S., Lee, Y. T., and Liu, D. Private convex optimization via exponential mechanism. In Conference on Learning Theory, pp. 1948–1989. PMLR, 2022.
  23. 23.Graves, A. Sequence transduction with recurrent neural networks. arXiv preprint arXiv:1211.3711, 2012.
  24. 24.Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., et al. Conformer: Convolution-augmented transformer for speech recognition. Proc. Interspeech 2020, pp. 5036–5040, 2020.
  25. 25.Hardt, M., Recht, B., and Singer, Y. Train faster, generalize better: Stability of stochastic gradient descent. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML’16, pp. 1225–1234. JMLR.org, 2016.
  26. 26.Hayes, T. P. A large-deviation inequality for vector-valued martingales. 2003. URL http://agl.cs.unm.edu/~hayes/papers/VectorAzuma/VectorAzuma20030207.pdf.
  27. 27.He, J., Li, X., Yu, D., Zhang, H., Kulkarni, J., Lee, Y. T., Backurs, A., Yu, N., and Bian, J. Exploring the limits of differentially private deep learning with group-wise clipping. arXiv preprint arXiv:2212.01539, 2022.
  28. 28.Jiang, D., Zhang, G., Karami, M., Chen, X., Shao, Y., and Yu, Y. Dp2-vae: Differentially private pre-trained variational autoencoders. arXiv preprint arXiv:2208.03409, 2022.
  29. 29.Kairouz, P., Ribero, M., Rush, K., and Thakurta, A. Fast dimension independent private adagrad on publicly estimated subspaces. arXiv preprint arXiv:2008.06570, 2020.
  30. 30.Kerrigan, G., Slack, D., and Tuyls, J. Differentially private language models benefit from public pre-training. In Proceedings of the Second Workshop on Privacy in NLP, pp. 39–45, 2020.
  31. 31.Kulkarni, J., Lee, Y. T., and Liu, D. Private non-smooth erm and sco in subquadratic steps. Advances in Neural Information Processing Systems, 34, 2021.
  32. 32.Kurakin, A., Chien, S., Song, S., Geambasu, R., Terzis, A., and Thakurta, A. Toward training at imagenet scale with differential privacy. arXiv preprint arXiv:2201.12328, 2022.
  33. 33.Li, X., Liu, D., Hashimoto, T., Inan, H. A., Kulkarni, J., Lee, Y., and Thakurta, A. G. When does differentially private learning not suffer in high dimensions? In Advances in Neural Information Processing Systems, 2022a.
  34. 34.Li, X., Tramer, F., Liang, P., and Hashimoto, T. Large language models can be strong differentially private learners. In International Conference on Learning Representations, 2022b.
  35. 35.Liu, T., Vietri, G., Steinke, T., Ullman, J., and Wu, S. Leveraging public data for practical private query release. In International Conference on Machine Learning, pp. 6968–6977. PMLR, 2021.
  36. 36.Luo, Z., Wu, D. J., Adeli, E., and Fei-Fei, L. Scalable differential privacy with sparse network finetuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5059–5068, 2021.
  37. 37.Nan, L., Radev, D., Zhang, R., Rau, A., Sivaprasad, A., Hsieh, C., Tang, X., Vyas, A., Verma, N., Krishna, P., et al. Dart: Open-domain structured data record to text generation. arXiv preprint arXiv:2007.02871, 2020.
  38. 38.Nasr, M., Mahloujifar, S., Tang, X., Mittal, P., and Houmansadr, A. Effectively using public data in privacy preserving machine learning, 2022. https://openreview.net/pdf?id=5R96mIU85IW.
  39. 39.Novikova, J., Dusek, O., and Rieser, V. The e2e dataset: New challenges for end-to-end generation. arXiv preprint arXiv:1706.09254, 2017.
  40. 40.Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 5206–5210. IEEE, 2015.
  41. 41.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  42. 42.Sharif Razavian, A., Azizpour, H., Sullivan, J., and Carlsson, S. Cnn features off-the-shelf: an astounding baseline for recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp. 806–813, 2014.
  43. 43.Song, S., Chaudhuri, K., and Sarwate, A. D. Stochastic gradient descent with differentially private updates. In 2013 IEEE Global Conference on Signal and Information Processing, pp. 245–248. IEEE, 2013.
  44. 44.Steinke, T. and Ullman, J. Tight lower bounds for differentially private selection. In 2017 IEEE 58th Annual Symposium on Foundations of Computer Science (FOCS), pp. 552–563. IEEE, 2017.
  45. 45.Tramer, F. and Boneh, D. Differentially private learning needs better features (or much more data). In International Conference on Learning Representations, 2020.
  46. 46.Tramer, F., Kamath, G., and Carlini, N. Considerations for differentially private learning with large-scale public pretraining. arXiv preprint arXiv:2212.06470, 2022.
  47. 47.Yu, D., Zhang, H., Chen, W., and Liu, T.-Y. Do not let privacy overbill utility: Gradient embedding perturbation for private learning. In International Conference on Learning Representations, 2021.
  48. 48.Yu, D., Naik, S., Backurs, A., Gopi, S., Inan, H. A., Kamath, G., Kulkarni, J., Lee, Y. T., Manoel, A., Wutschitz, L., et al. Differentially private fine-tuning of language models. In International Conference on Learning Representations, 2022.
  49. 49.Zhou, Y., Wu, S., and Banerjee, A. Bypassing the ambient dimension: Private sgd with gradient subspace identification. In International Conference on Learning Representations, 2021.

Citation

MLA
Ganesh, A., et al. “Why Is Public Pretraining Necessary for Private Model Training?”. International Conference on Machine Learning, vol. 202, 2023, pp. 10611–27, https://proceedings.mlr.press/v202/ganesh23a.html.
APA
Ganesh, A., Haghifam, M., Nasr, M., Oh, S., Steinke, T., Thakkar, O., Thakurta, A. G., & Wang, L. (2023). Why Is Public Pretraining Necessary for Private Model Training?. International Conference on Machine Learning, 202, 10611–10627. https://proceedings.mlr.press/v202/ganesh23a.html
Chicago
Ganesh, A., M. Haghifam, M. Nasr, et al. 2023. “Why Is Public Pretraining Necessary for Private Model Training?”. International Conference on Machine Learning 202: 10611–27. https://proceedings.mlr.press/v202/ganesh23a.html.
Harvard
Ganesh, A. et al. (2023) “Why Is Public Pretraining Necessary for Private Model Training?”, International Conference on Machine Learning. PMLR, pp. 10611–10627. Available at: https://proceedings.mlr.press/v202/ganesh23a.html.
Vancouver
1. Ganesh A, Haghifam M, Nasr M, Oh S, Steinke T, Thakkar O, Thakurta AG, Wang L (2023) Why Is Public Pretraining Necessary for Private Model Training?. In: International Conference on Machine Learning. PMLR, pp 10611–10627

BibTeX

@InProceedings{pmlr-v202-ganesh23a,
  title = 	 {Why Is Public Pretraining Necessary for Private Model Training?},
  author =       {Ganesh, Arun and Haghifam, Mahdi and Nasr, Milad and Oh, Sewoong and Steinke, Thomas and Thakkar, Om and Guha Thakurta, Abhradeep and Wang, Lun},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {10611--10627},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/ganesh23a/ganesh23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/ganesh23a.html},
  abstract = 	 {In the privacy-utility tradeoff of a model trained on benchmark language and vision tasks, remarkable improvements have been widely reported when the model is pretrained on public data. Some gain is expected as these models inherit the benefits of transfer learning, which is the standard motivation in non-private settings. However, the stark contrast in the gain of pretraining between non-private and private machine learning suggests that the gain in the latter is rooted in a fundamentally different cause. To explain this phenomenon, we hypothesize that the non-convex loss landscape of a model training necessitates the optimization algorithm to go through two phases. In the first, the algorithm needs to select a good “basin” in the loss landscape. In the second, the algorithm solves an easy optimization within that basin. The former is a harder problem to solve with private data, while the latter is harder to solve with public data due to a distribution shift or data scarcity. Guided by this intuition, we provide theoretical constructions that provably demonstrate the separation between private training with and without public pretraining. Further, systematic experiments on CIFAR10 and Librispeech provide supporting evidence for our hypothesis.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/