Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models

Haonan DuanAdam DziedzicNicolas PapernotFranziska Boenisch

article2023NeurIPS116 citations

Presents practical differentially private prompt learning methods for large language models, demonstrating that private soft prompt tuning and ensemble-based discrete prompt generation protect sensitive context data from membership inference attacks while matching non-private task accuracy.

Listen

Large language models deliver strong performance across business tasks using in-context prompting, where sample demonstrations guide model behavior without modifying model parameters. However, incorporating proprietary text or sensitive personal information into prompts creates severe confidentiality and compliance risks. Prior defenses have relied on differentially private fine-tuning, but this approach demands substantial computational power, requires specialized access to model parameters, and fails with standard commercial model application programming interfaces (APIs). The article evaluates whether prompt data leaks sensitive information and introduces practical methods to train both soft and discrete prompts under rigorous mathematical differential privacy guarantees.

The authors first establish an attack method to measure privacy vulnerabilities in prompted models, showing that standard prompting leaks private data. They then develop two privacy-preserving frameworks: PromptDPSGD, which uses private gradient descent on soft prompt embeddings while keeping the base language model frozen, and PromptPATE, which creates an ensemble of private discrete prompts ("a flock of stochastic parrots") and transfers their collective knowledge into a clean, public prompt via a differentially private noisy voting mechanism. The evaluation covers standard natural language classification benchmarks across proprietary black-box APIs, including GPT-3 and Claude, as well as open architectures like RoBERTa.

The article demonstrates four primary findings. First, existing prompted models are highly vulnerable to privacy attacks; a simple membership inference attack achieves an average area under the curve (AUC) of 0.84 on GPT-3, reliably identifying sensitive prompt data. Second, PromptPATE effectively neutralizes this risk—reducing attack success to roughly 0.50 (equivalent to random guessing)—while matching non-private utility; for example, on the SST-2 benchmark with GPT-3, it achieves 92.7% accuracy under a strict privacy budget, closely trailing the non-private baseline of 95.2% and heavily outperforming the 82.0% zero-shot baseline. Third, PromptPATE remains effective even when the public transfer data comes from a different domain or task than the private data, such as using news articles to protect encyclopedia extracts (reaching 74.6% accuracy versus a 44.2% zero-shot baseline). Fourth, PromptDPSGD achieves accuracy within 3% to 7% of non-private baselines across various tasks while adjusting orders of magnitude fewer parameters (under 10,000 parameters) than full model fine-tuning (125 million parameters).

These findings prove that organizations can deploy sensitive, proprietary workflows on public or commercial LLMs without trading off compliance, cost, or performance. Private prompting dramatically cuts storage requirements by avoiding the need to host separate model weights per task and enables concurrent batch processing of multiple distinct tasks. Because PromptPATE requires only black-box text outputs, enterprises can enforce formal differential privacy on existing commercial cloud APIs immediately.

Organizations handling confidential downstream tasks should adopt private prompt learning over full model fine-tuning. Engineering teams using black-box cloud APIs should implement PromptPATE for discrete prompt generation, while teams hosting internal models with gradient access can use PromptDPSGD to minimize parameter storage. Decision-makers should note that these methods protect downstream prompt data rather than the underlying pretraining data of the base LLM, and PromptPATE currently requires trusting the API provider during intermediate queries unless paired with cryptographic safeguards.

arXiv: 2305.15594
  • Paper: Deep Learning with Differential Privacy, Martín Abadi et al. (2016). This seminal work introduces differentially private stochastic gradient descent (DP-SGD) with gradient clipping and calibrated noise, which forms the core optimization algorithm adapted by PromptDPSGD.
  • Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). This paper establishes parameter-efficient soft prompt tuning with frozen base models, providing the architectural foundation for the soft prompt embedding optimization in the source.
  • Paper: GPT Understands, Too, Xiao Liu et al. (2021). This paper introduces continuous prompt tuning (P-Tuning) to adapt frozen language models via continuous embeddings, serving as a key predecessor to soft prompt learning techniques.
  • Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). This study demonstrates how large language models memorize and leak training data under query attacks, motivating the privacy defenses developed in the source.
  • Paper: The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks, Nicholas Carlini et al. (2018). This foundational work establishes methodologies and exposure metrics for unintended data memorization in generative sequence models, framing the leakage risks evaluated in the source.
  • Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This paper establishes in-context few-shot prompting for large language models, the primary operational paradigm whose privacy vulnerabilities the source examines.
  • Paper: Differentially Private Empirical Risk Minimization, Kamalika Chaudhuri et al. (2009). This foundational paper provides the theoretical principles for empirical risk minimization under differential privacy guarantees that underlie privacy-preserving machine learning.
  • Paper: What Can We Learn Privately?, Shiva Prasad Kasiviswanathan et al. (2008). This foundational paper formalizes the theoretical sample complexity and learnability bounds of machine learning under differential privacy constraints.
Cover for Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models

Abstract

Large language models (LLMs) are excellent in-context learners. However, the sensitivity of data contained in prompts raises privacy concerns. Our work first shows that these concerns are valid: we instantiate a simple but highly effective membership inference attack against the data used to prompt LLMs. To address this vulnerability, one could forego prompting and resort to fine-tuning LLMs with known algorithms for private gradient descent. However, this comes at the expense of the practicality and efficiency offered by prompting. Therefore, we propose to privately learn to prompt. We first show that soft prompts can be obtained privately through gradient descent on downstream data. However, this is not the case for discrete prompts. Thus, we orchestrate a noisy vote among an ensemble of LLMs presented with different prompts, i.e., a flock of stochastic parrots. The vote privately transfers the flock’s knowledge into a single public prompt. We show that LLMs prompted with our private algorithms closely match the non-private baselines. For example, using GPT3 as the base model, we achieve a downstream accuracy of 92.7% on the sst2 dataset with (ε = 0.147, δ = 10⁻⁶)-differential privacy vs. 95.2% for the non-private baseline. Through our experiments, we also show that our prompt-based approach is easily deployed with existing commercial APIs.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 3 Private Information about Prompt Data Leaks from Prompted LLMs
  • 4 Methods for Privacy Preserving Prompts
  • 4.1 PromptDPSGD: DPSGD for Private Soft Prompt Learning
  • 4.2 PromptPATE: PATE for Privacy Preserving Discrete Prompts
  • 4.3 Advantages of (Private) Prompting over (Private) Fine-Tuning
  • 5 Experimental Evaluation
  • 5.1 PromptDPSGD
  • 5.2 PromptPATE
  • 6 Conclusions and Outlook
  • Acknowledgments
  • References
  • A Broader Impacts
  • B Limitations
  • C Additional Insights into our Methods
  • C.1 PromptDPSGD
  • C.2 PromptPATE
  • C.3 Privacy Analysis
  • D Additional Results
  • D.1 Membership Inference Attacks
  • D.2 PromptPATE on Claude
  • D.3 More results for PromptDPSGD
  • E Additional Setup
  • E.1 PromptDPSGD
  • E.2 PromptPATE
  • E.2.1 Hyperparameters for Confident-GNMax
  • E.2.2 Dataset Preprocessing

Knowls

  1. Knowl 1 — PromptPATE for Differentially Private Discrete Prompt Learning

    model/method

    PromptPATE enables differentially private downstream adaptation of black-box large language models (LLMs) without accessing model parameters or computing internal gradients, requiring only black-box query access to next-token predictions.

    The framework operates across three distinct stages:

    1. Flock of Teachers (Private Discrete Prompts): A private downstream dataset Dpriv={(px,py)}D_{\text{priv}} = \{(p_x, p_y)\} is split into disjoint subsets. Each subset is used to construct a distinct discrete demonstration prompt Pi=[Instruction,(px,1,py,1),… ]P_i = [\text{Instruction}, (p_{x,1}, p_{y,1}), \dots] that is prepended to queries sent to an LLM LL, creating an ensemble of EE private teacher models LPiL_{P_i}.

    2. Private Knowledge Transfer via Confident-GNMAX: An unlabeled public dataset Dpub={xk}D_{\text{pub}} = \{x_k\} is presented to all EE teacher prompts. For each candidate sequence xkx_k, teacher ii outputs a predicted class label token yi,ky_{i,k}. The aggregated vote count for class label jj across the ensemble is denoted nj(xk)=∑i=1EI(yi,k=j)n_j(x_k) = \sum_{i=1}^E \mathbb{I}(y_{i,k} = j). The Confident-GNMAX aggregator checks consensus with threshold TT and Gaussian noise σ1\sigma_1: max⁡j{nj(xk)}+N(0,σ12)≥T\max_j \{n_j(x_k)\} + \mathcal{N}(0, \sigma_1^2) \ge T If this condition fails, the query is rejected (returning ⊥\bot) to prevent leaking boundary information. If the threshold is met, the assigned pseudo-label y~k\tilde{y}_k is computed via noisy argmax: y~k=arg⁡max⁡j(nj(xk)+N(0,σ22))\tilde{y}_k = \arg\max_j \left( n_j(x_k) + \mathcal{N}(0, \sigma_2^2) \right)

    3. Student Prompt Selection: The privately labeled pairs (xk,y~k)(x_k, \tilde{y}_k) form candidate discrete student prompts. To avoid privacy leakage from evaluating prompts on private validation data, the candidate prompts are evaluated on a held-out portion of the newly labeled public validation data. The prompt achieving the highest public validation accuracy is selected as the public student prompt PstudentP_{\text{student}} and deployed with the frozen base LLM for end users.

  2. Knowl 2 — PromptDPSGD Algorithm for Private Soft Prompt Learning

    algorithm

    PromptDPSGD adapts a frozen large language model LL to a private downstream classification task by learning continuous soft prompt embeddings P∈Rs×eP \in \mathbb{R}^{s \times e} (where ss is the prompt token length and ee is the embedding dimensionality) using differentially private stochastic gradient descent (DP-SGD).

    Input: Private downstream dataset D={(xi,yi)}i=1ND = \{(x_i, y_i)\}_{i=1}^N, prompt length ss, embedding dimension ee, frozen LLM LL, loss function ℓ(LP,x,y)\ell(L_P, x, y), learning rate ηt\eta_t, noise multiplier σ\sigma, sampling rate qq, gradient clipping norm cc, total iterations TT
    Output: Privatized prompt embedding PT∈Rs×eP_T \in \mathbb{R}^{s \times e}, privacy parameters (ε,δ)(\varepsilon, \delta)
    Initialize P0∈Rs×eP_0 \in \mathbb{R}^{s \times e} at random
    for t=0t = 0 to T−1T-1 do
        Sample mini-batch Bt⊆DB_t \subseteq D via Poisson subsampling with probability qq
        for each (xi,yi)∈Bt(x_i, y_i) \in B_t do
            Compute per-sample gradient gt(xi)←∇Ptℓ(LPt(xi),yi)g_t(x_i) \leftarrow \nabla_{P_t} \ell(L_{P_t}(x_i), y_i)
            Clip per-sample gradient gˉt(xi)←gt(xi)/max⁡(1,∥gt(xi)∥2c)\bar{g}_t(x_i) \leftarrow g_t(x_i) / \max\left(1, \frac{\|g_t(x_i)\|_2}{c}\right)
        end for
        Sample Gaussian noise vector Zt∼N(0,σ2c2I)Z_t \sim \mathcal{N}(0, \sigma^2 c^2 I)
        Compute privatized batch gradient g~t←1∣Bt∣(∑xi∈Btgˉt(xi)+Zt)\tilde{g}_t \leftarrow \frac{1}{|B_t|} \left( \sum_{x_i \in B_t} \bar{g}_t(x_i) + Z_t \right)
        Update soft prompt parameters Pt+1←Pt−ηtg~tP_{t+1} \leftarrow P_t - \eta_t \tilde{g}_t
    end for
    Compute overall privacy expenditure (ε,δ)(\varepsilon, \delta)
    return PT,(ε,δ)P_T, (\varepsilon, \delta)

    The optimization updates solely the continuous prompt parameters PtP_t, keeping all internal LLM weights frozen. Privacy guarantees are evaluated using the standard DP-SGD moments accountant over the TT subsampled Gaussian mechanism iterations.

  3. Knowl 3 — PromptDPSGD for Parameter-Efficient Differentially Private Adaptation

    model/method

    PromptDPSGD is a differentially private parameter-efficient adaptation technique for transformer-based language models. Instead of applying DP-SGD to all internal model parameters (full DP fine-tuning) or adapter layers (such as LoRA), PromptDPSGD freezes the underlying language model entirely and optimizes continuous prompt embeddings prepended to the input layer (soft prompts) or continuous prefix embeddings prepended to every transformer layer (prefix tuning).

    Compared to DP fine-tuning, PromptDPSGD provides several operational advantages:

    • Parameter efficiency: Learning soft prompts of length s=10s=10 requires optimizing under 10K10\text{K} parameters (e.g., 2,306 parameters on RoBERTa-Base for SST-2), and prefix tuning requires under 100K100\text{K} parameters (e.g., 19,970 parameters on RoBERTa-Base), compared to 125M125\text{M} parameters for full fine-tuning and 1.2M1.2\text{M} parameters for LoRA.
    • Storage efficiency: Instead of storing a full model checkpoint (approximately 500 MB for RoBERTa-Base) per downstream task, only a task-specific prompt vector (<40 KB< 40\text{ KB} for soft prompts, <400 KB< 400\text{ KB} for prefix tuning) needs to be stored.
    • Mixed-task inference: Because the base model remains unmodified, prompts from multiple distinct downstream tasks can be combined and executed within the same forward inference batch without switching checkpoints.
  4. Knowl 4 — Theoretical Differential Privacy Guarantee for PromptDPSGD

    theoretical result

    Let TT denote the total number of training iterations, q∈(0,1]q \in (0, 1] the Poisson subsampling rate per iteration, and c>0c > 0 the ℓ2\ell_2-norm clipping bound on per-example gradients in PromptDPSGD. Under the sampled Gaussian mechanism with noise scale σ\sigma, there exist fundamental constants c1c_1 and c2c_2 such that for any target ε<c1q2T\varepsilon < c_1 q^2 T and any δ>0\delta > 0, PromptDPSGD satisfies (ε,δ)(\varepsilon, \delta)-differential privacy with respect to the private training dataset if the noise scale satisfies: σ≥c2qcTlog⁡(1/δ)ε\sigma \ge c_2 \frac{q c \sqrt{T \log(1/\delta)}}{\varepsilon}

    Because the gradient clipping threshold cc bounds the ℓ2\ell_2 sensitivity of the sum of clipped gradients in each mini-batch, the privacy accountant tracks privacy loss over the sequence of updates to the prompt embedding P∈Rs×eP \in \mathbb{R}^{s \times e}. By the post-processing property of differential privacy, any subsequent query or public deployment of the privatized prompt PTP_T incurs no additional privacy leakage regarding the private training dataset.

  5. Knowl 5 — Membership Inference Attack Formulation on Prompted Large Language Models

    model/method

    Membership inference attacks (MIAs) against in-context prompted large language models evaluate whether a specific demonstration tuple (px,py)(p_x, p_y) was included in the discrete prompt P=[Instruction,(px,py),… ]P = [\text{Instruction}, (p_x, p_y), \dots] prepended to inputs during black-box model inference.

    Attack Formulation:

    1. The adversary targets a prompted language model LPL_P accessible only via black-box queries that return output probability vectors over vocabulary tokens.
    2. The adversary possesses a candidate set of nn text sequences with true class labels {(xi,li)}i=1n\{(x_i, l_i)\}_{i=1}^n, where a small subset are true members used in PP and the majority are non-members (e.g., a realistic 1:50 member-to-non-member ratio).
    3. The adversary queries each sequence xix_i to LPL_P, obtaining the output probability vector yi=LP(xi)∈[0,1]My_i = L_P(x_i) \in [0, 1]^M, where MM is the vocabulary size.
    4. The adversary extracts the predicted probability yi,liy_{i, l_i} assigned to the correct class label token lil_i.

    Because LLMs assign substantially higher probability to the correct label token when the demonstration (xi,li)(x_i, l_i) was included in the in-context prompt compared to when it was not, candidate points are classified as prompt members by thresholding yi,liy_{i, l_i}.

  6. Knowl 6 — Empirical Vulnerability of Prompted Large Language Models to Membership Inference

    empirical result

    Standard in-context prompting of large language models without privacy defenses leaks membership information about the demonstration data contained in the prompt.

    Evaluating black-box membership inference attacks across 100 independent trials (1 member demonstration versus 50 non-member private training samples per trial) yields the following empirical attack performance:

    • GPT-3 (Babbage, 1-shot prompts):

      • DBpedia: average AUC-ROC of 0.840.84
      • AG News: average AUC-ROC of 0.710.71
      • SST-2: average AUC-ROC of 0.580.58
      • TREC: average AUC-ROC of 0.510.51
    • GPT-2-XL (4-shot prompts):

      • AG News: average AUC-ROC of 0.860.86
      • CB: average AUC-ROC of 0.840.84
      • SST-2: average AUC-ROC of 0.720.72
      • TREC: average AUC-ROC of 0.690.69

    In all cases with strong performance, output probabilities for member demonstrations at the target label token are substantially higher than those for non-member samples, confirming that prompt data is vulnerable to privacy extraction.

  7. Knowl 7 — Downstream Accuracy and Differential Privacy Guarantees of PromptPATE

    empirical result

    PromptPATE provides strong differential privacy guarantees while closely matching non-private prompting performance across multiple commercial LLM backends (GPT-3 Babbage, GPT-3 Curie, and Claude-v1) and NLP classification benchmarks (SST-2, AG News, TREC, DBpedia) under both in-distribution (IID) and out-of-distribution (OOD) public transfer datasets (at δ=10−6\delta = 10^{-6}):

    • GPT-3 Babbage (200 teacher prompts, 1-shot):

      • SST-2: Non-private baseline 93.8%93.8\%, zero-shot 76.3%76.3\%. PromptPATE achieves 88.8±2.3%88.8 \pm 2.3\% (IID SST-2, ε=0.178\varepsilon = 0.178) and 87.2±1.9%87.2 \pm 1.9\% (OOD IMDB, ε=0.187\varepsilon = 0.187).
      • AG News: Non-private 78.2%78.2\%, zero-shot 62.0%62.0\%. PromptPATE achieves 71.7±0.8%71.7 \pm 0.8\% (IID AG News, ε=0.248\varepsilon = 0.248) and 67.9±1.7%67.9 \pm 1.7\% (OOD AriseTV, ε=0.258\varepsilon = 0.258).
      • TREC: Non-private 58.7%58.7\%, zero-shot 40.7%.PromptPATEachieves40.7\%. PromptPATE achieves 52.8 \pm 1.5%(IIDTREC,(IID TREC,\varepsilon = 0.281)and) and 50.9 \pm 3.5%(OODQQP,(OOD QQP,\varepsilon = 0.293$).
      • DBpedia: Non-private 85.6%85.6\%, zero-shot 44.2%44.2\%. PromptPATE achieves 80.3±1.3%80.3 \pm 1.3\% (IID DBpedia, ε=0.194\varepsilon = 0.194) and 74.6±1.4%74.6 \pm 1.4\% (OOD AG News, ε=0.203\varepsilon = 0.203).
    • GPT-3 Curie (SST-2): Non-private 95.2%95.2\%, zero-shot 82.0%82.0\%. PromptPATE achieves 92.3±1.1%92.3 \pm 1.1\% (IID, ε=0.147\varepsilon = 0.147) and 92.7±0.8%92.7 \pm 0.8\% (OOD IMDB, ε=0.154\varepsilon = 0.154).

    • Claude-v1 (400 teacher prompts, full black-box output):

      • SST-2 (2-shot): Non-private 98.0%98.0\%, zero-shot 92.7%92.7\%; PromptPATE achieves 95.7±1.4%95.7 \pm 1.4\% at ε=0.048\varepsilon = 0.048.
      • AG News (2-shot): Non-private 82.7%82.7\%, zero-shot 72.4%72.4\%; PromptPATE achieves 74.6±1.5%74.6 \pm 1.5\% at ε=0.056\varepsilon = 0.056.
      • TREC (4-shot): Non-private 82.2%82.2\%, zero-shot 69.0%69.0\%; PromptPATE achieves 79.3±1.2%79.3 \pm 1.2\% at ε=0.068\varepsilon = 0.068.
      • DBpedia (1-shot): Non-private 93.5%93.5\%, zero-shot 88.0%88.0\%; PromptPATE achieves 90.9±0.6%90.9 \pm 0.6\% at ε=0.042\varepsilon = 0.042.
  8. Knowl 8 — Comparative Performance and Parameter Efficiency of PromptDPSGD versus Fine-Tuning

    data/table

    The downstream classification accuracy (%) of PromptDPSGD (Soft-Prompt and Prefix tuning) evaluated on RoBERTa-Base (125M125\text{M} base parameters) is compared against full DP fine-tuning and DP LoRA tuning across GLUE benchmark tasks at DP budget ε=8\varepsilon = 8 (with δ=1/N\delta = 1/N, where NN is dataset size) and non-private baseline ε=∞\varepsilon = \infty:

    Dataset Soft-Prompt (Our) Prefix (Our) Full-Tuning LoRA-Tuning
    Tuned Params <10K<10\text{K} <100K<100\text{K} 125M125\text{M} 1.2M1.2\text{M}
    Guarantee ε=8\varepsilon = 8 ε=∞\varepsilon = \infty ε=8\varepsilon = 8 ε=∞\varepsilon = \infty ε=8\varepsilon = 8 ε=∞\varepsilon = \infty ε=8\varepsilon = 8 ε=∞\varepsilon = \infty
    sst2 92.31% 95.64% 91.97% 96.33% 85.89% 96.40% 92.97% 96.60%
    qnli 84.11% 89.48% 87.17% 94.84% 84.81% 94.70% 88.59% 94.70%
    qqp 81.52% 86.56% 82.58% 91.42% 86.15% 92.20% 86.26% 92.20%
    mnli 75.15% 82.49% 80.57% 90.34% 83.30% 90.20% 82.92% 90.20%

    On simpler tasks (SST-2, QNLI), soft prompts and prefix tuning match or exceed full DP fine-tuning despite tuning orders of magnitude fewer parameters (2,3062,306 parameters for SST-2 soft prompt vs. 125M125\text{M} for full tuning). On more complex classification tasks (QQP, MNLI), prefix tuning remains within 2%–3%2\%\text{--}3\% of LoRA fine-tuning while leaving the base model unmodified.

  9. Knowl 9 — Data Efficiency and Privacy-Utility Dynamics in PromptPATE

    empirical result

    Ablation studies of PromptPATE on GPT-3 Babbage (using DBpedia as private data and AG News as public transfer data) demonstrate two key operational dynamics:

    1. High Teacher Consensus Reduces Privacy Expenditure: When 200 teacher prompts vote on 500 public input queries, the proportion of teachers agreeing on the modal class is consistently high across the majority of queries. In the Confident-GNMAX mechanism, higher consensus among teachers permits lower noise addition and minimal query rejections, keeping total accumulated privacy expenditure low (ε<0.3\varepsilon < 0.3 at δ=10−6\delta = 10^{-6}).

    2. Rapid Utility Saturation with Few Public Queries: Student test accuracy grows rapidly as the public dataset size increases from 0 to 100 queries, reaching approximately 73%–74%73\%\text{--}74\% test accuracy at ε≈0.146\varepsilon \approx 0.146. Beyond 100 queries up to 500 queries (where ε\varepsilon reaches 0.2030.203), student accuracy plateaus. Thus, discrete student prompts require as few as 100 labeled public demonstrations to achieve high utility.

  10. Knowl 10 — PromptPATE Defense Effectiveness Against Membership Inference Attacks

    empirical result

    Evaluating membership inference attacks against public discrete student prompts generated by PromptPATE confirms that the framework prevents prompt data extraction.

    When adversaries conduct MIAs targeting the private demonstrations used in the teacher ensemble, the resulting ROC curves for the public student prompts align with the random-guess diagonal (y=xy = x), yielding near-chance AUC-ROC scores across all benchmark datasets:

    • AG News: AUC = 0.480.48
    • SST-2: AUC = 0.490.49
    • DBpedia: AUC = 0.520.52
    • TREC: AUC = 0.510.51

    These results empirically demonstrate that the combination of noisy teacher voting (Confident-GNMAX) and public-data-only student prompt selection mitigates membership inference risk on prompt data.

  11. Knowl 11 — Limitations of Differentially Private Prompt Learning

    limitation

    The private prompt learning methods have several documented limitations:

    • Trusted LLM API Provider Assumption: PromptPATE submits private teacher prompt demonstrations to the black-box LLM API provider. Differential privacy guarantees protect prompt data against end users querying the public student prompt, but trust in the API provider is assumed to protect private demonstrations sent during teacher inference.
    • Orthogonality to Pre-training Privacy: Private prompt learning exclusively prevents privacy leakage from downstream prompt demonstrations; it does not protect against privacy leakage or memorization stemming from the LLM's original pre-training corpora.
    • Fixed Prompt Templates: Discrete prompt instructions and templates rely on fixed hand-crafted formats without automated instruction optimization, which could potentially improve utility further.
    • Experimental Scale Constraints: Due to commercial API costs and rate limits, evaluations on larger models (such as GPT-4) and teacher ensembles beyond 400 models were constrained.

Coverage note — No substantial contributed material was omitted. All core algorithms (PromptDPSGD, PromptPATE), membership inference attacks, theoretical guarantees, and empirical results across models and benchmarks are represented.

References

  1. 1.M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security, pages 308–318, 2016.
  2. 2.R. Anil, B. Ghazi, V. Gupta, R. Kumar, and P. Manurangsi. Large-scale differentially private bert. arXiv preprint arXiv:2108.01624, 2021.
  3. 3.Antropic. Introducing claude. Antropic Website. 2023-03-14, https://www.anthropic.com/index/introducing-claude.
  4. 4.R. Bassily, A. Smith, and A. Thakurta. Private empirical risk minimization: Efficient algorithms and tight error bounds. In 2014 IEEE 55th annual symposium on foundations of computer science, pages 464–473. IEEE, 2014.
  5. 5.E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623, New York, NY, USA, 2021. Association for Computing Machinery.
  6. 6.T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  7. 7.N. Carlini, S. Chien, M. Nasr, S. Song, A. Terzis, and F. Tramer. Membership inference attacks from first principles. In 2022 IEEE Symposium on Security and Privacy (SP), pages 1897–1914. IEEE, 2022.
  8. 8.N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646, 2022.
  9. 9.N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650, 2021.
  10. 10.H.-T. Cheng and R. Thoppilan. Lamda: Towards safe, grounded, and high-quality dialog models for everything. Google Blog Post. 2023-05-09, https://ai.googleblog.com/2022/01/lamda-towards-safe-grounded-and-high.html.
  11. 11.O. chimaobi Samuel. news-data. Huggingface, 2022.
  12. 12.C. A. Choquette-Choo, N. Dullerud, A. Dziedzic, Y. Zhang, S. Jha, N. Papernot, and X. Wang. Capc learning: Confidential and private collaborative learning. In International Conference on Learning Representations, 2021.
  13. 13.J. Davison, J. Feldman, and A. M. Rush. Commonsense knowledge mining from pretrained models. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pages 1173–1178, 2019.
  14. 14.J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  15. 15.C. Dwork. Differential privacy. In Automata, Languages and Programming: 33rd International Colloquium, ICALP 2006, Venice, Italy, July 10-14, 2006, Proceedings, Part II 33, pages 1–12. Springer, 2006.
  16. 16.T. Gao, A. Fisch, and D. Chen. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723, 2020.
  17. 17.Google. Lamda: Towards safe, grounded, and high-quality dialog models for everything. Website, 2023. 2023, https://bard.google.com/.
  18. 18.H. Guo, B. Tan, Z. Liu, E. Xing, and Z. Hu. Efficient (soft) q-learning for text generation with limited good data. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 6969–6991, 2022.
  19. 19.S. Hoory, A. Feder, A. Tendler, S. Erell, A. Peled-Cohen, I. Laish, H. Nakhost, U. Stemmer, A. Benjamini, A. Hassidim, et al. Learning and evaluating a differentially private pre-trained language model. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 1178–1189, 2021.
  20. 20.D. Ippolito, F. Tramèr, M. Nasr, C. Zhang, M. Jagielski, K. Lee, C. A. Choquette-Choo, and N. Carlini. Preventing verbatim memorization in language models gives a false sense of privacy. arXiv preprint arXiv:2210.17546, 2022.
  21. 21.B. Jayaraman, L. Wang, K. Knipmeyer, Q. Gu, and D. Evans. Revisiting membership inference under realistic assumptions. Proceedings on Privacy Enhancing Technologies, 2021(2), 2021.
  22. 22.Z. Jiang, F. F. Xu, J. Araki, and G. Neubig. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438, 2020.
  23. 23.E. Kharitonov, M. Baroni, and D. Hupkes. How bpe affects memorization in transformers. arXiv preprint arXiv:2110.02782, 2021.
  24. 24.B. Lester, R. Al-Rfou, and N. Constant. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Nov. 2021.
  25. 25.X. Li, F. Tramer, P. Liang, and T. Hashimoto. Large language models can be strong differentially private learners. In International Conference on Learning Representations, 2022.
  26. 26.X. L. Li and P. Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, 2021.
  27. 27.J. Liu, D. Shen, Y. Zhang, B. Dolan, L. Carin, and W. Chen. What makes good in-context examples for gpt-3? arXiv preprint arXiv:2101.06804, 2021.
  28. 28.X. Liu, K. Ji, Y. Fu, W. Tam, Z. Du, Z. Yang, and J. Tang. P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 61–68, Dublin, Ireland, May 2022. Association for Computational Linguistics.
  29. 29.X. Liu, K. Ji, Y. Fu, W. L. Tam, Z. Du, Z. Yang, and J. Tang. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602, 2021.
  30. 30.Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov. Ro{bert}a: A robustly optimized {bert} pretraining approach, 2020.
  31. 31.A. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, pages 142–150, 2011.
  32. 32.R. T. McCoy, P. Smolensky, T. Linzen, J. Gao, and A. Celikyilmaz. How much do language models copy from their training data? evaluating linguistic novelty in text generation using raven. arXiv preprint arXiv:2111.09509, 2021.
  33. 33.F. Mireshghallah, A. Uniyal, T. Wang, D. Evans, and T. Berg-Kirkpatrick. Memorization in nlp fine-tuning methods. arXiv preprint arXiv:2205.12506, 2022.
  34. 34.R. Mitchell. Samsung fab data leak: How chatgpt exposed sensitive information. electropages.
  35. 35.OpenAI. Gpt-4 technical report, 2023.
  36. 36.N. Papernot, M. Abadi, Ú. Erlingsson, I. Goodfellow, and K. Talwar. Semi-supervised knowledge transfer for deep learning from private training data. In International Conference on Learning Representations, 2017.
  37. 37.N. Papernot, S. Song, I. Mironov, A. Raghunathan, K. Talwar, and U. Erlingsson. Scalable private learning with pate. In International Conference on Learning Representations, 2022.
  38. 38.F. Petroni, T. Rocktäschel, P. Lewis, A. Bakhtin, Y. Wu, A. H. Miller, and S. Riedel. Language models as knowledge bases? arXiv preprint arXiv:1909.01066, 2019.
  39. 39.A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al. Improving language understanding by generative pre-training. 2018.
  40. 40.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  41. 41.C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  42. 42.T. L. Scao and A. M. Rush. How many data points is a prompt worth? arXiv preprint arXiv:2103.08493, 2021.
  43. 43.V. Shejwalkar, H. A. Inan, A. Houmansadr, and R. Sim. Membership inference attacks against NLP classification models. In NeurIPS 2021 Workshop Privacy in Machine Learning, 2021.
  44. 44.T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020.
  45. 45.R. Shokri, M. Stronati, C. Song, and V. Shmatikov. Membership inference attacks against machine learning models. In 2017 IEEE symposium on security and privacy (SP), pages 3–18. IEEE, 2017.
  46. 46.R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642, 2013.
  47. 47.C. Song and V. Shmatikov. Auditing data provenance in text-generation models. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 196–206, 2019.
  48. 48.K. Tirumala, A. H. Markosyan, L. Zettlemoyer, and A. Aghajanyan. Memorization without overfitting: Analyzing the training dynamics of large language models. arXiv preprint arXiv:2205.10770, 2022.
  49. 49.F. Tramer and D. Boneh. Differentially private learning needs better features (or much more data). In International Conference on Learning Representations, 2021.
  50. 50.E. M. Voorhees and D. M. Tice. Building a question answering test collection. In Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval, pages 200–207, 2000.
  51. 51.A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations, 2019.
  52. 52.Z. Wang, W. Hamza, and R. Florian. Bilateral multi-perspective matching for natural language sentences. In Proceedings of the 26th International Joint Conference on Artificial Intelligence, pages 4144–4150, 2017.
  53. 53.S. Yeom, I. Giacomelli, M. Fredrikson, and S. Jha. Privacy risk in machine learning: Analyzing the connection to overfitting. In 2018 IEEE 31st computer security foundations symposium (CSF), pages 268–282. IEEE, 2018.
  54. 54.D. Yu, S. Naik, A. Backurs, S. Gopi, H. A. Inan, G. Kamath, J. Kulkarni, Y. T. Lee, A. Manoel, L. Wutschitz, S. Yekhanin, and H. Zhang. Differentially private fine-tuning of language models. In International Conference on Learning Representations, 2022.
  55. 55.C. Zhang, D. Ippolito, K. Lee, M. Jagielski, F. Tramèr, and N. Carlini. Counterfactual memorization in neural language models. arXiv preprint arXiv:2112.12938, 2021.
  56. 56.S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  57. 57.X. Zhang, J. Zhao, and Y. LeCun. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28, 2015.
  58. 58.Z. Zhao, E. Wallace, S. Feng, D. Klein, and S. Singh. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR, 2021.

Citation

MLA
Duan, H., et al. “Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 76852–71, https://proceedings.neurips.cc/paper_files/paper/2023/file/f26119b4ffe38c24d97e4c49d334b99e-Paper-Conference.pdf.
APA
Duan, H., Dziedzic, A., Papernot, N., & Boenisch, F. (2023). Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models. Advances in Neural Information Processing Systems, 36, 76852–76871. https://proceedings.neurips.cc/paper_files/paper/2023/file/f26119b4ffe38c24d97e4c49d334b99e-Paper-Conference.pdf
Chicago
Duan, H., A. Dziedzic, N. Papernot, and F. Boenisch. 2023. “Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models”. Advances in Neural Information Processing Systems 36: 76852–71. https://proceedings.neurips.cc/paper_files/paper/2023/file/f26119b4ffe38c24d97e4c49d334b99e-Paper-Conference.pdf.
Harvard
Duan, H. et al. (2023) “Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 76852–76871. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/f26119b4ffe38c24d97e4c49d334b99e-Paper-Conference.pdf.
Vancouver
1. Duan H, Dziedzic A, Papernot N, Boenisch F (2023) Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 76852–76871

BibTeX

@inproceedings{duan2023flocks,
  title = {Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models},
  author = {Duan, Haonan and Dziedzic, Adam and Papernot, Nicolas and Boenisch, Franziska},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {76852-76871},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/f26119b4ffe38c24d97e4c49d334b99e-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors