Differentially Private Bias-Term Fine-tuning of Foundation Models

Zhiqi BuYu-Xiang WangSheng ZhaGeorge Karypis

article2024ICML65 citationsBest Paper Award (NeurIPS TSRML Workshop, 2022)

Proposes a differentially private bias-term fine-tuning method that updates only 0.1% of parameters to drastically cut memory and computational overhead, matching state-of-the-art accuracy while enabling practical private fine-tuning on high-resolution vision and long-sequence language tasks.

Listen

Adapting large pre-trained foundation models to specialized downstream tasks using sensitive user data poses severe privacy risks, including data leakage and reconstruction attacks. Differential privacy provides a rigorous mathematical safeguard against these vulnerabilities, but existing private fine-tuning techniques suffer from severe computational drawbacks. Conventional private methods require caching massive activation tensors in memory and processing per-sample weight gradients, which creates bottlenecks that make fine-tuning large models or high-dimensional data (such as long documents and high-resolution images) prohibitively expensive or technically infeasible.

The article evaluates differentially private bias-term fine-tuning (DP-BiTFiT), demonstrating that updating only the bias terms of neural networks eliminates these computational bottlenecks while matching state-of-the-art accuracy under strict privacy guarantees. The authors systematically evaluated the method across multiple language and vision benchmarks (such as GLUE, E2E, ImageNet, CIFAR, and CelebA) using foundation architectures including BERT, RoBERTa, GPT-2, Vision Transformers, and ResNet.

The investigation produced several key findings regarding efficiency and task performance. First, updating only the bias terms restricts optimization to approximately 0.1% of the total network parameters while eliminating the need to store activation tensors during the forward pass. Second, DP-BiTFiT is 2 to 30 times faster and requires 2 to 8 times less memory than standard private full fine-tuning, operating about 1.5 times faster even than non-private full fine-tuning. Third, in terms of task utility, DP-BiTFiT performs on par with private full fine-tuning and parameter-efficient alternatives like LoRA and Adapters; in larger models such as GPT-2 Large, it slightly outperformed private full fine-tuning. Finally, because its private computation overhead does not scale with feature dimension, the method scales efficiently to long sequence lengths and large image resolutions that typically crash other privacy-preserving frameworks.

These findings have direct strategic implications for technical infrastructure, operational costs, and regulatory compliance. Organizations handling sensitive data can deploy differential privacy without purchasing specialized, ultra-high-memory hardware or enduring substantial training slowdowns. In distributed learning scenarios, updating only 0.1% of parameters reduces network communication volume by up to 1,000 times, significantly lowering multi-node bandwidth requirements and cloud hosting expenses.

Engineering teams are recommended to adopt DP-BiTFiT as a default, lightweight baseline for privacy-preserving transfer learning on transformer architectures. For architectures that lack native bias parameters in certain layers (such as certain convolutional networks or LLaMA), practitioners should use the proposed extension that introduces trainable bias terms, or implement a hybrid two-phase training protocol where full private training is run for one or two initial epochs before switching to bias-only updates. Future work should pilot hybrid frameworks combining bias tuning with low-rank adapters and evaluate performance on multi-billion-parameter models and document-scale text tasks.

No sufficiently relevant recommendations were found.

Cover for Differentially Private Bias-Term Fine-tuning of Foundation Models

Abstract

We study the problem of differentially private (DP) fine-tuning of large pre-trained models – a recent privacy-preserving approach suitable for solving downstream tasks with sensitive data. Existing work has demonstrated that high accuracy is possible under strong privacy constraint, yet requires significant computational overhead or modifications to the network architecture. We propose differentially private bias-term fine-tuning (DP-BiTFiT), which matches the state-of-the-art accuracy for DP algorithms and the efficiency of the standard BiTFiT. DP-BiTFiT is model agnostic (not modifying the network architecture), parameter efficient (only training about 0.1% of the parameters), and computation efficient (almost removing the overhead caused by DP, in both the time and space complexity). On a wide range of tasks, DP-BiTFiT is 2 ∼ 30× faster and uses 2 ∼ 8× less memory than DP full fine-tuning, even faster than the standard full fine-tuning. This amazing efficiency enables us to conduct DP fine-tuning on language and vision tasks with long-sequence texts and high-resolution images, which were computationally difficult using existing methods. We open-source our code at FastDP (https://github.com/awslabs/fast-differential-privacy).

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Differentially private Bias-Term Fine-Tuning
  • 3.1 Parameter efficiency
  • 3.2 Complexity of weight and bias training
  • 3.3 Scalability of DP algorithms
  • 3.3.1 EFFICIENCY V.S. FEATURE DIMENSION
  • 3.3.2 EFFICIENCY V.S. MODEL SIZE
  • 3.4 Applicability of DP-BiTFiT
  • 4 Experiments
  • 4.1 Text classification
  • 4.2 Natural Language Generation
  • 4.3 Image classification
  • 5 Discussion
  • Impact Statement
  • References
  • A Detailed analysis
  • A.1 Back-propagation
  • A.2 Making BiTFiT work with convolutional neural networks
  • A.2.1 WALK-AROUND 1
  • A.2.2 WALK-AROUND 2
  • B Implementation of DP-BiTFiT
  • C Complexity analysis
  • D Experiment details
  • D.1 Language tasks
  • D.2 Image tasks
  • E Additional tables and figures
  • E.1 Parameter efficiency of DP-BiTFiT
  • E.2 More results on DP-BiTFiT and language tasks
  • E.3 More results on two-phase training
  • E.4 Hyperparameter tuning for DP-BiTFiT

Knowls

  1. Knowl 1 — DP-BiTFiT privately fine-tunes only existing bias parameters

    algorithm

    Differentially private bias-term fine-tuning (DP-BiTFiT) applies DP-SGD to the concatenation bb of a network’s existing bias parameters, while leaving its weights fixed. For training example ii, let gi=∇bℓig_i=\nabla_b\ell_i be the gradient of its loss and let dbd_b be the number of bias parameters. A minibatch is formed by independently including each of the nn examples with probability pp. Let C(x;R)C(x;R) be a clipping multiplier satisfying C(x;R)≤R/xC(x;R)\le R/x, where RR is the clipping threshold, and let σ\sigma be the noise multiplier. DP-BiTFiT releases the summed gradient

    G=∑i∈IC(∥gi∥2;R)gi+σRz,z∼N(0,Idb),G=\sum_{i\in I} C(\|g_i\|_2;R)g_i+\sigma Rz,\qquad z\sim\mathcal{N}(0,I_{d_b}),

    where II is the sampled minibatch and IdbI_{d_b} is the dbd_b-dimensional identity matrix. Privacy is obtained using the sampled-Gaussian mechanism and accounting across the TT iterations; the resulting guarantee is expressed as (ε,δ)(\varepsilon,\delta)-DP.

    Input: Training set of n examples; bias parameters b; sampling probability p; iterations T; clipping threshold R; noise multiplier sigma; clipping function C; optimizer
    For iteration t = 1,...,T:
        Include each training example independently with probability p to form I
        For each example i in I:
            Backpropagate its loss and collect the per-example gradients g_i for all bias parameters
            Compute the single global norm ||g_i||_2 across all layers
            Set c_i = C(||g_i||_2; R)
        Sum the clipped gradients G = sum over i in I of c_i g_i
        Add Gaussian noise G = G + sigma R z, where z is standard normal with one coordinate per bias parameter
        Update b using G and the selected optimizer
    Output: The fine-tuned bias parameters b

    The paper implements the per-example bias gradients with backward hooks, without forward hooks that cache layer activations. The optimizer can be SGD, Adam, or another gradient-based optimizer; sampling, noise, and iteration counts are selected for the target privacy budget and task.

  2. Knowl 2 — Bias gradients can be computed without layer input activations

    equation

    For an affine layer with input ai,t∈Rda_{i,t}\in\mathbb{R}^{d} and output si,t=ai,tW+b∈Rps_{i,t}=a_{i,t}W+b\in\mathbb{R}^{p} at position tt, the bias b∈Rpb\in\mathbb{R}^{p} is shared across the TT positions of example ii. Define hi,t=∇si,tℓi∈Rph_{i,t}=\nabla_{s_{i,t}}\ell_i\in\mathbb{R}^{p} as the output gradient at position tt. The per-example bias gradient is

    gi=∇bℓi=∑t=1Thi,t.g_i=\nabla_b\ell_i=\sum_{t=1}^{T}h_{i,t}.

    Thus the bias gradient is obtained by summing the output gradient over positions; it does not require the input activation ai,ta_{i,t}. By contrast, the per-example weight gradient depends on the layer inputs, through terms of the form hi,tai,t⊤h_{i,t}a_{i,t}^{\top}. DP-BiTFiT uses the bias-gradient property to avoid caching activations and to avoid computing per-example gradients for weights. The paper states that the same input-independent bias-gradient computation applies to convolutional and normalization layers when their affine biases are trained.

  3. Knowl 3 — DP-BiTFiT reduces training cost by eliminating per-example weight-gradient work

    theoretical result

    In the paper’s FLOP accounting, consider a network whose layer ll maps B×Tl×dlB\times T_l\times d_l inputs to B×Tl×plB\times T_l\times p_l outputs. Here BB is minibatch size, TlT_l is the number of positions, and dl,pld_l,p_l are the input and output feature dimensions. Summed over layers, the reported leading time costs are approximately

    \text{DP full fine-tuning}: >8B\sum_l T_l p_l d_l,\qquad \text{DP-BiTFiT}: \approx4B\sum_l T_l p_l d_l.$$ The paper therefore estimates DP-BiTFiT to be $1.5\times$ faster than standard full fine-tuning and more than $2\times$ faster than DP full fine-tuning under this accounting. For bias training, the per-layer gradient-and-clipping work is $BT_lp_l+3Bp_l$; the additional DP term is $3Bp_l$, independent of $T_l$. The reported per-layer training-space accounting for DP-BiTFiT is $p_l+Bp_l$, and it does not require storing activation tensors. These comparisons concern the paper’s stated layer-wise operation and storage accounting, not a hardware-independent runtime guarantee.
  4. Knowl 4 — DP-BiTFiT matches full DP fine-tuning closely on RoBERTa classification

    empirical result

    On four text-classification tasks, DP-BiTFiT was evaluated against full fine-tuning using RoBERTa-base and RoBERTa-large at privacy budget ε=8\varepsilon=8. The entries below are test accuracies in percent; “standard” denotes non-private training. DP-BiTFiT trains only bias parameters, while the reported full method trains all parameters.

    Model Method SST2 QNLI QQP MNLI-m
    RoBERTa-base Full, standard 94.5 91.4 87.3 85.9
    RoBERTa-base Full, DP 92.1 87.9 86.1 83.2
    RoBERTa-base BiTFiT, standard 93.5 87.3 86.1 83.4
    RoBERTa-base BiTFiT, DP 92.4 86.9 85.6 82.9
    RoBERTa-large Full, standard 96.2 93.6 87.9 90.3
    RoBERTa-large Full, DP 93.8 91.1 87.5 87.0
    RoBERTa-large BiTFiT, standard 95.5 92.2 87.9 89.3
    RoBERTa-large BiTFiT, DP 94.5 91.1 86.9 88.3

    At ε=8\varepsilon=8, DP-BiTFiT is close to DP full fine-tuning across these model/task pairs: its accuracy is sometimes slightly lower and, for RoBERTa-large on SST2 and MNLI-m, higher. The reported classification experiments used a maximum sequence length of 256.

  5. Knowl 5 — On E2E generation, DP-BiTFiT approaches DP full fine-tuning as GPT-2 scales

    empirical result

    The paper evaluates GPT-2 on the E2E restaurant-description generation task at ε=8\varepsilon=8. The metrics are perplexity (lower is better), BLEU, ROUGE-L, NIST, METEOR, and CIDEr (higher is better). The table reports DP full fine-tuning and DP-BiTFiT results; values are reproduced as reported.

    Model Method Perplexity BLEU ROUGE-L NIST METEOR CIDEr
    GPT2-small Full, DP 2.33 63.60 67.07 7.71 0.40 1.94
    GPT2-small BiTFiT, DP 2.89 60.56 64.96 6.14 0.37 1.62
    GPT2-medium Full, DP 2.25 64.22 67.53 8.17 0.42 2.08
    GPT2-medium BiTFiT, DP 2.67 61.02 66.13 7.18 0.39 1.80
    GPT2-large Full, DP 2.26 64.64 68.97 8.30 0.42 2.16
    GPT2-large BiTFiT, DP 2.59 65.21 67.88 8.43 0.42 2.15

    The gap in BLEU narrows with model size: DP-BiTFiT scores 3.04 and 3.20 points below DP full fine-tuning on GPT2-small and GPT2-medium, respectively, but 0.57 points above it on GPT2-large. On GPT2-large, DP-BiTFiT also exceeds DP full fine-tuning on NIST, while the other reported metrics are close but not uniformly better.

  6. Knowl 6 — ViT-large DP-BiTFiT achieves high CIFAR accuracy across privacy budgets

    empirical result

    A pretrained ViT-large was fine-tuned for three epochs on CIFAR10 and CIFAR100 at the listed privacy budgets. CIFAR images were resized to 224×224224\times224; the downstream classification head was replaced for the new class set. Accuracies are percentages. DP-BiTFiT trains biases together with the newly initialized last-layer weights, while the DP full baseline trains the full model.

    Dataset Method ε=1\varepsilon=1 ε=2\varepsilon=2 ε=4\varepsilon=4 ε=8\varepsilon=8
    CIFAR10 DP last-layer 98.4 98.6 98.6 98.7
    CIFAR10 DP-BiTFiT 98.9 99.0 99.0 99.0
    CIFAR10 DP full 98.9 98.9 99.0 99.0
    CIFAR100 DP last-layer 86.2 87.3 88.1 88.8
    CIFAR100 DP-BiTFiT 90.2 91.2 91.8 92.3
    CIFAR100 DP full 87.7 90.1 91.0 91.3

    DP-BiTFiT matches DP full fine-tuning on CIFAR10 to within 0.1 percentage points at each reported budget. On CIFAR100, it exceeds the DP full result at all four budgets, including 91.2% versus 90.1% at ε=2\varepsilon=2.

  7. Knowl 7 — Measured runtime and memory favor DP-BiTFiT over DP full fine-tuning

    empirical result

    The paper reports that, across its range of tasks, DP-BiTFiT is 22–30×30\times faster and uses 22–8×8\times less memory than DP full fine-tuning; it also reports that DP-BiTFiT can be faster than standard full fine-tuning. A QQP text-classification ablation with RoBERTa-large on one 40-GB A100 GPU gives a concrete runtime comparison: with batch size 20, the fastest DP full-fine-tuning implementation (GhostClip) took 119 minutes per epoch. Disabling weight gradients while leaving forward hooks active took 80 minutes; additionally removing forward hooks reduced this to 63 minutes; using the larger batch size enabled by DP-BiTFiT’s memory savings reduced the epoch time to 43 minutes. These measurements illustrate that avoiding weight-gradient computation and activation caching both contribute to the observed speedup.

  8. Knowl 8 — Bias parameters are a small trainable fraction across diverse architectures

    empirical result

    The measured share of parameters trained by BiTFiT, and therefore by DP-BiTFiT, is small across the tested vision and language models. The reported percentages are percentages of total model parameters, not fractions of a percent.

    • Vision models: ResNet18 (11.7M parameters), 0.043%; ResNet50 (25.6M), 0.113%; ViT-small-patch16 (21.7M), 0.238%; ViT-base-patch16 (85.8M), 0.120%; ViT-large-patch16 (303M), 0.090%.
    • Language models: GPT2-small (124M), 0.082%; GPT2-medium (355M), 0.076%; GPT2-large (774M), 0.066%; RoBERTa-base (125M), 0.083%; RoBERTa-large (355M), 0.077%.

    These measurements support the paper’s claim that bias-only fine-tuning is model-agnostic and typically trains about 0.1% of the parameters. The proportions vary by architecture; for example, the tested ViT-small has 0.238% bias parameters.

  9. Knowl 9 — Models without useful bias terms limit direct application of DP-BiTFiT

    limitation

    DP-BiTFiT depends on trainable bias parameters, so it may be inapplicable or less effective in architectures that omit them. The paper identifies bias-free models such as LLAMA and convolutional layers followed by batch normalization as examples. Batch normalization is also unsuitable for the paper’s DP training setup because its statistics depend on samples; replacing it with group normalization changes the architecture, and biases that were ineffective with batch normalization can become relevant with group normalization. In a ResNet18 CelebA multi-label experiment, DP-BiTFiT reached 86.9% accuracy, compared with 88.4% for DP full fine-tuning. The proposed DP-BiTFiT-Add workaround adds zero- or randomly initialized biases and then trains them privately. On that task it reached 87.3%; it trained less than 0.1% of ResNet18 parameters, and the paper reports 0.03% for LLAMA2-7B. The paper states that the added biases preserve pretrained utility at initialization, while enlarging the trainable parameter space.

  10. Knowl 10 — A short DP full-fine-tuning phase can improve bias-only fine-tuning

    model/method

    The paper defines XX+BiTFiT as a two-phase private fine-tuning method: first perform DP full fine-tuning for XX epochs, then perform DP-BiTFiT for the remaining epochs. Setting X=0X=0 gives DP-BiTFiT throughout; setting XX to the total number of epochs gives DP full fine-tuning throughout. In the reported CIFAR100 experiment with ViT-large-patch16-224 at ε=2\varepsilon=2, accuracy was 19.4% for 00+BiTFiT, 82.4% for 11+BiTFiT, 89.0% for 22+BiTFiT, and 90.1% for DP full fine-tuning. The paper reports that using at most two initial full-fine-tuning epochs can yield accuracy comparable to full fine-tuning while retaining some of the later-stage efficiency of bias-only training.

Coverage note — Supplementary per-attribute CelebA scores and detailed hyperparameter grids are omitted because they support, rather than constitute, distinct contributions.

References

  1. 1.Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., and Zhang, L. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security, pp. 308–318, 2016.
  2. 2.Banerjee, S. and Lavie, A. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pp. 65–72, Ann Arbor, Michigan, June 2005. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/W05-0909.
  3. 3.Beltagy, I., Peters, M. E., and Cohan, A. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  4. 4.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  5. 5.Bu, Z., Dong, J., Long, Q., and Su, W. J. Deep learning with gaussian differential privacy. Harvard data science review, 2020(23), 2020.
  6. 6.Bu, Z., Mao, J., and Xu, S. Scalable and efficient training of large convolutional neural networks with differential privacy. arXiv preprint arXiv:2205.10683, 2022a.
  7. 7.Bu, Z., Wang, Y.-X., Zha, S., and Karypis, G. Automatic clipping: Differentially private deep learning made easier and stronger. arXiv preprint arXiv:2206.07136, 2022b.
  8. 8.Bu, Z., Wang, Y.-X., Zha, S., and Karypis, G. Differentially private optimization on large model at small cost. In International Conference on Machine Learning, pp. 3192–3218. PMLR, 2023.
  9. 9.Cai, H., Gan, C., Zhu, L., and Han, S. Tinytl: Reduce memory, not parameters for efficient on-device learning. Advances in Neural Information Processing Systems, 33:11285–11297, 2020.
  10. 10.Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pp. 2633–2650, 2021.
  11. 11.Chen, T., Xu, B., Zhang, C., and Guestrin, C. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174, 2016.
  12. 12.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
  13. 13.De, S., Berrada, L., Hayes, J., Smith, S. L., and Balle, B. Unlocking high-accuracy differentially private image classification through scale. arXiv preprint arXiv:2204.13650, 2022.
  14. 14.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020.
  15. 15.Dusek, O., Novikova, J., and Rieser, V. Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge. Computer Speech & Language, 59:123–156, January 2020. doi: 10.1016/j.csl.2019.06.009.
  16. 16.Dwork, C., McSherry, F., Nissim, K., and Smith, A. Calibrating noise to sensitivity in private data analysis. In Theory of cryptography conference, pp. 265–284. Springer, 2006.
  17. 17.Goodfellow, I. Efficient per-example gradient computations. arXiv preprint arXiv:1510.01799, 2015.
  18. 18.Gopi, S., Lee, Y. T., and Wutschitz, L. Numerical composition of differential privacy. Advances in Neural Information Processing Systems, 34, 2021.
  19. 19.Goyal, P., Dollar, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677, 2017.
  20. 20.Haim, N., Vardi, G., Yehudai, G., Shamir, O., and Irani, M. Reconstructing training data from trained neural networks. arXiv preprint arXiv:2206.07758, 2022.
  21. 21.He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G. Towards a unified view of parameter-efficient transfer learning. In International Conference on Learning Representations, 2021.
  22. 22.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  23. 23.Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pp. 2790–2799. PMLR, 2019.
  24. 24.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  25. 25.Iyer, S., Dandekar, N., and Csernai, K. First quora dataset release: Question pairs, 2017. URL https://data.quora.com/First-Quora-Dataset-Release-Question-Pairs.
  26. 26.Jain, P., Jain, A., Nrusimha, A., Gholami, A., Abbeel, P., Gonzalez, J., Keutzer, K., and Stoica, I. Checkmate: Breaking the memory wall with optimal tensor rematerialization. Proceedings of Machine Learning and Systems, 2:497–511, 2020.
  27. 27.Karras, T., Aila, T., Laine, S., and Lehtinen, J. Progressive growing of gans for improved quality, stability, and variation. In International Conference on Learning Representations, 2018.
  28. 28.Karras, T., Laine, S., and Aila, T. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4401–4410, 2019.
  29. 29.Kenton, J. D. M.-W. C. and Toutanova, L. K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pp. 4171–4186, 2019.
  30. 30.Kornblith, S., Shlens, J., and Le, Q. V. Do better imagenet models transfer better? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2661–2671, 2019.
  31. 31.Koskela, A., Jalko, J., and Honkela, A. Computing tight differential privacy guarantees using fft. In International Conference on Artificial Intelligence and Statistics, pp. 2560–2569. PMLR, 2020.
  32. 32.Kurakin, A., Chien, S., Song, S., Geambasu, R., Terzis, A., and Thakurta, A. Toward training at imagenet scale with differential privacy. arXiv preprint arXiv:2201.12328, 2022.
  33. 33.Lee, J. and Kifer, D. Scaling up differentially private deep learning with fast per-example gradient clipping. arXiv preprint arXiv:2009.03106, 2020.
  34. 34.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 3045–3059, 2021.
  35. 35.Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7871–7880, 2020.
  36. 36.Lhoest, Q., Villanova del Moral, A., Jernite, Y., Thakur, A., von Platen, P., Patil, S., Chaumond, J., Drame, M., Plu, J., Tunstall, D., Davison, J., Sasko, M., Chhablani, G., Malik, B., Brandeis, S., Le Scao, T., Sanh, V., Xu, C., Patry, N., McMillan-Major, A., Schmid, P., Gugger, S., Delangue, C., Matussiere, T., Debut, L., Bekman, S., Cistac, P., Goehringer, T., Mustar, V., Lagunas, F., Rush, A., and Wolf, T. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 175–184, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.emnlp-demo.21.
  37. 37.Li, X., Tramer, F., Liang, P., and Hashimoto, T. Large language models can be strong differentially private learners. arXiv preprint arXiv:2110.05679, 2021.
  38. 38.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190, 2021.
  39. 39.Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pp. 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/W04-1013.
  40. 40.Lin, Z., Madotto, A., and Fung, P. Exploring versatile generative language model via parameter-efficient transfer learning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 441–459, 2020.
  41. 41.Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J. Gpt understands, too. arXiv preprint arXiv:2103.10385, 2021.
  42. 42.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  43. 43.Mahabadi, R. K., Henderson, J., and Ruder, S. Compacter: Efficient low-rank hypercomplex adapter layers. arXiv preprint arXiv:2106.04647, 2021.
  44. 44.Mehta, H., Thakurta, A., Kurakin, A., and Cutkosky, A. Large scale transfer learning for differentially private image classification. arXiv preprint arXiv:2205.02973, 2022.
  45. 45.Mironov, I., Talwar, K., and Zhang, L. Renyi differential privacy of the sampled gaussian mechanism. arXiv preprint arXiv:1908.10530, 2019. URL http://arxiv.org/abs/1908.10530.
  46. 46.Pan, S. J. and Yang, Q. A survey on transfer learning. IEEE Transactions on knowledge and data engineering, 22(10):1345–1359, 2009.
  47. 47.Papineni, K., Roukos, S., Ward, T., and jing Zhu, W. Bleu: a method for automatic evaluation of machine translation. pp. 311–318, 2002.
  48. 48.Pfeiffer, J., Kamath, A., Ruckle, A., Cho, K., and Gurevych, I. Adapterfusion: Non-destructive task composition for transfer learning. In 16th Conference of the European Chapter of the Associationfor Computational Linguistics, EACL 2021, pp. 487–503. Association for Computational Linguistics (ACL), 2021.
  49. 49.Polyak, B. T. and Juditsky, A. B. Acceleration of stochastic approximation by averaging. SIAM journal on control and optimization, 30(4):838–855, 1992.
  50. 50.Qiao, S., Wang, H., Liu, C., Shen, W., and Yuille, A. Micro-batch training with batch-channel normalization and weight standardization. arXiv preprint arXiv:1903.10520, 2019.
  51. 51.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners.
  52. 52.Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–16. IEEE, 2020.
  53. 53.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250, 2016.
  54. 54.Rebuffi, S.-A., Bilen, H., and Vedaldi, A. Learning multiple visual domains with residual adapters. Advances in neural information processing systems, 30, 2017.
  55. 55.Ruckle, A., Geigle, G., Glockner, M., Beck, T., Pfeiffer, J., Reimers, N., and Gurevych, I. Adapterdrop: On the efficiency of adapters in transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7930–7946, 2021.
  56. 56.Sadjadi, S. O., Kheyrkhah, T., Tong, A., Greenberg, C. S., Reynolds, D. A., Singer, E., Mason, L. P., Hernandez-Cordero, J., et al. The 2017 nist language recognition evaluation. In Odyssey, pp. 82–89, 2018.
  57. 57.Shokri, R., Stronati, M., Song, C., and Shmatikov, V. Membership inference attacks against machine learning models. In 2017 IEEE symposium on security and privacy (SP), pp. 3–18. IEEE, 2017.
  58. 58.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pp. 1631–1642, 2013.
  59. 59.Subramani, P., Vadivelu, N., and Kamath, G. Enabling fast differentially private sgd via just-in-time compilation and vectorization. Advances in Neural Information Processing Systems, 34, 2021.
  60. 60.Sun, C., Shrivastava, A., Singh, S., and Gupta, A. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE international conference on computer vision, pp. 843–852, 2017.
  61. 61.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  62. 62.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  63. 63.Tramer, F. and Boneh, D. Differentially private learning needs better features (or much more data). arXiv preprint arXiv:2011.11660, 2020.
  64. 64.Vedantam, R., Lawrence Zitnick, C., and Parikh, D. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4566–4575, 2015.
  65. 65.Wang, Y.-X., Balle, B., and Kasiviswanathan, S. P. Subsampled renyi differential privacy and analytical moments accountant. In International Conference on Artificial Intelligence and Statistics, pp. 1226–1235. PMLR, 2019.
  66. 66.Williams, A., Nangia, N., and Bowman, S. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 1112–1122. Association for Computational Linguistics, 2018. URL http://aclweb.org/anthology/N18-1101.
  67. 67.Yang, X., Zhang, H., Chen, W., and Liu, T.-Y. Normalized/clipped sgd with perturbation for differentially private non-convex optimization. arXiv preprint arXiv:2206.13033, 2022.
  68. 68.Yousefpour, A., Shilov, I., Sablayrolles, A., Testuggine, D., Prasad, K., Malek, M., Nguyen, J., Ghosh, S., Bharadwaj, A., Zhao, J., Cormode, G., and Mironov, I. Opacus: User-friendly differential privacy library in PyTorch. arXiv preprint arXiv:2109.12298, 2021.
  69. 69.Yu, D., Naik, S., Backurs, A., Gopi, S., Inan, H. A., Kamath, G., Kulkarni, J., Lee, Y. T., Manoel, A., Wutschitz, L., et al. Differentially private fine-tuning of language models. arXiv preprint arXiv:2110.06500, 2021a.
  70. 70.Yu, D., Zhang, H., Chen, W., and Liu, T.-Y. Do not let privacy overbill utility: Gradient embedding perturbation for private learning. In International Conference on Learning Representations, 2021b. URL https://openreview.net/forum?id=7aogOj_VYO0.
  71. 71.Yu, D., Zhang, H., Chen, W., Yin, J., and Liu, T.-Y. Large scale private learning via low-rank reparametrization. In International Conference on Machine Learning, pp. 12208–12218. PMLR, 2021c.
  72. 72.Zaken, E. B., Goldberg, Y., and Ravfogel, S. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 1–9, 2022.

Citation

MLA
Bu, Z., et al. “Differentially Private Bias-Term Fine-tuning of Foundation Models”. arXiv, 2022, http://arxiv.org/abs/2210.00036v3.
APA
Bu, Z., Wang, Y.-X., Zha, S., & Karypis, G. (2022). Differentially Private Bias-Term Fine-tuning of Foundation Models. arXiv. http://arxiv.org/abs/2210.00036v3
Chicago
Bu, Z., Y.-X. Wang, S. Zha, and G. Karypis. 2022. “Differentially Private Bias-Term Fine-tuning of Foundation Models”. arXiv. http://arxiv.org/abs/2210.00036v3.
Harvard
Bu, Z. et al. (2022) “Differentially Private Bias-Term Fine-tuning of Foundation Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.00036v3.
Vancouver
1. Bu Z, Wang Y-X, Zha S, Karypis G (2022) Differentially Private Bias-Term Fine-tuning of Foundation Models. arXiv

BibTeX

@article{bu2022differentially,
  title = {Differentially Private Bias-Term Fine-tuning of Foundation Models},
  author = {Bu, Zhiqi and Wang, Yu-Xiang and Zha, Sheng and Karypis, George},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.00036v3},
  eprint = {2210.00036}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/