Built independently by an author, for readers. Read the story and support ChapterPal

keyword

private model training

Private model training is the process of developing machine learning models on sensitive datasets using privacy-preserving techniques that prevent the exposure, leakage, or reconstruction of individual training records. Most commonly implemented through frameworks such as differential privacy, this approach modifies optimization algorithms by clipping gradients and injecting calibrated statistical noise during the learning process. By mathematically bounding the influence of any single data point on the final model parameters, private model training defends against data extraction and membership inference attacks while seeking an optimal balance between formal privacy guarantees and predictive utility.

1 item

Why Is Public Pretraining Necessary for Private Model Training?

Why Is Public Pretraining Necessary for Private Model Training?

Arun Ganesh, Mahdi Haghifam, Milad Nasr, Sewoong Oh, Thomas Steinke, Om Thakkar, Abhradeep Guha Thakurta, Lun Wang

OrganizationsGoogleUniversity of TorontoUniversity of Washington

Why you should read this

Explains why public pretraining is essential for differentially private learning by theoretically proving and empirically demonstrating that noiseless early optimization is required to select viable basins before fine-tuning on sensitive data.

In the privacy-utility tradeoff of a model trained on benchmark language and vision tasks, remarkable improvements have been widely reported when the model is pretrained on public data. Some gain is expected as these models inherit the benefits of transfer learning, which is the standard motivation in non-private settings. However, the stark contrast in the gain of pretraining between non-private and private machine learning suggests that the gain in the latter is rooted in a fundamentally different cause. To explain this phenomenon, we hypothesize that the non-convex loss landscape of a model training necessitates the optimization algorithm to go through two phases. In the first, the algorithm needs to select a good “basin” in the loss landscape. In the second, the algorithm solves an easy optimization within that basin. The former is a harder problem to solve with private data, while the latter is harder to solve with public data due to a distribution shift or data scarcity. Guided by this intuition, we provide theoretical constructions that provably demonstrate the separation between private training with and without public pretraining. Further, systematic experiments on CIFAR10 and Librispeech provide supporting evidence for our hypothesis.

Added

2026-10-04