AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs

Madhava Gaikwad

article2025arXiv0 citations

Proposes AlignDP, a hybrid differential privacy framework that proactively prevents large language model extraction and unauthorized fine-tuning at the data interface by completely shielding rare inputs while maintaining accurate statistical utility for frequent patterns.

Listen

Large language models deployed across consumer and enterprise applications face critical security risks, including intellectual property theft, unauthorized model distillation, and sensitive training data extraction. Existing defense strategies primarily rely on post-incident mechanisms, such as output watermarking or activity monitoring, which attempt to detect abuse only after private information has already leaked. The article introduces and evaluates AlignDP, an architectural privacy framework designed to block unauthorized knowledge transfer directly at the data release interface before leakage can occur.

The framework implements a two-tier mechanism that differentiates sensitive, rare data from frequent, non-rare information based on an occurrence threshold. Highly identifying rare events are hidden using statistical indistinguishability guarantees that provide effective zero-leakage protection at the user level. Frequent events are privatized using a randomized response protocol (RAPPOR), which adds controlled mathematical noise to user telemetry while allowing a central aggregator to accurately compute aggregate frequency statistics. A global management layer regulates the total privacy budget across repeated user interactions to prevent adversaries from accumulating sensitive intelligence over time.

Empirical tests and theoretical evaluations confirm the effectiveness of this design. Simulated extraction attacks demonstrate that rare event detection remains locked at noise levels (mean absolute error of 0.001) even after 100 adversarial queries. Frequent categories are recovered with strong utility, achieving an 80% top-5 category identification accuracy, a rank correlation of 0.798, and an estimation error decay that follows theoretical bounds as sample sizes increase. Furthermore, the correlation for frequent patterns deliberately saturates at approximately 0.99, enforcing a strict upper bound on extraction capabilities regardless of repeated query attempts.

These findings indicate that organizations can prevent competitors or malicious actors from distilling proprietary models and conducting unauthorized fine-tuning, as the privatized outputs lack the precise label signals required for machine learning convergence. Rather than treating privacy solely as a compliance burden, this architectural approach provides proactive intellectual property protection at the interface level without destroying the utility of aggregate user telemetry.

Decision-makers should consider integrating rarity-aware privacy architectures into AI telemetry pipelines while planning for subsequent operational development. Future work must adapt fixed rarity thresholds to dynamic data distributions, develop sequence-level protections for correlated natural language outputs, and optimize communication overhead to scale from small categorical domains to full large language model vocabularies exceeding 50,000 tokens. While current results provide high confidence in the underlying mathematical proofs and small-scale feasibility, cautious validation on full-scale language models is recommended before enterprise-wide deployment.

arXiv: 2512.17251
  • Paper: Extracting alignment data in open models, Federico Barbero et al. (2025). This work investigates the extraction of post-training alignment data from open LLMs, directly illustrating the practical alignment leakage risks that AlignDP is designed to mitigate.
  • Paper: Large-scale online deanonymization with LLMs, Simon Lermen et al. (2026). This research evaluates large-scale deanonymization via LLMs on unstructured text, representing downstream re-identification risks that robust interface-level differential privacy defenses must counter.
Cover for AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs

Abstract

Large language models are exposed to risks of extraction, distillation, and unauthorized fine-tuning. Existing defenses use watermarking or monitoring, but these act after leakage. We design AlignDP, a hybrid privacy lock that blocks knowledge transfer at the data interface. The key idea is to separate rare and non-rare fields. Rare fields are shielded by PAC indistinguishability, giving effective zero-epsilon local DP. Non-rare fields are privatized with RAPPOR, giving unbiased frequency estimates under local DP. A global aggregator enforces composition and budget. This two-tier design hides rare events and adds controlled noise to frequent events. We prove limits of PAC extension to global aggregation, give bounds for RAPPOR estimates, and analyze utility trade-off. A toy simulation confirms feasibility: rare categories remain hidden, frequent categories are recovered with small error.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Model
  • 3.1 Rarity threshold
  • 3.2 Mechanism
  • 3.3 Adversary model
  • 3.4 Pipeline Illustration
  • 4 Theoretical Guarantees
  • 4.1 Rare events
  • 4.2 Non-rare events
  • 4.3 Global composition
  • 4.4 Summary
  • 5 Operational Illustration
  • 5.1 Basic Mechanism Validation
  • 5.2 Extraction Resistance
  • 5.3 Privacy Guarantees
  • 5.4 Utility Preservation
  • 6 Discussion and Future Work
  • 6.1 Key Findings and Implications
  • 6.2 Limitations and Extensions
  • 6.3 Future Directions
  • A Definitions
  • B Proof Sketches
  • C Extended Notes
  • D Relation to Lock-LLM Goals
  • E Extraction Resistance Numbers
  • F Code Availability and Reproduction
  • References

Knowls

  1. Knowl 1 — AlignDP Two-Tier Hybrid Privacy Mechanism

    model/method

    AlignDP is a hybrid privacy mechanism designed for telemetry and log data release from Large Language Model (LLM) interactions. It partitions categorical event domains based on a rarity threshold α>0\alpha > 0.

    For each user record X=(X1,…,Xd)X = (X_1, \dots, X_d) where field XiX_i takes values from domain Di\mathcal{D}_i with distribution μi\mu_i:

    1. Domain Partitioning: The domain Di\mathcal{D}_i is split into a rare event set Ri={x∈Di:μi(x)<α}R_i = \{x \in \mathcal{D}_i : \mu_i(x) < \alpha\} and a non-rare event set Ni=Di∖RiN_i = \mathcal{D}_i \setminus R_i.

    2. PAC Shielding for Rare Events: If x∈Rix \in R_i, the mechanism emits the symbol xx under Probably Approximately Correct (PAC) indistinguishability guarantees, bounding the distinguishability error to δ(n,α)\delta(n, \alpha) and providing effective zero-ϵ\epsilon local differential privacy (LDP).

    3. RAPPOR for Non-Rare Events: If x∈Nix \in N_i, xx is encoded as a bit vector v∈{0,1}mv \in \{0, 1\}^m, and each bit flips independently with probability pp. The privatized vector yy is transmitted under ϵ\epsilon-LDP.

    4. Global Aggregation: A central aggregator collects all privatized reports, computes debiased frequency estimates for non-rare categories, reports only PAC bounds for rare counts, and manages a global privacy budget across queries to prevent privacy loss accumulation.

  2. Knowl 2 — PAC Indistinguishability Guarantee for Rare Events

    theoretical result

    In AlignDP, rare events whose marginal distribution is below a rarity threshold α>0\alpha > 0 are protected via PAC indistinguishability.

    For any rare event xx satisfying μ(x)<α\mu(x) < \alpha, given nn independent and identically distributed (i.i.d.) user samples, the probability of an adversary distinguishing xx from other rare events is bounded by:

    δ(n,α)=exp⁡(−2n(α−μ(x))2)\delta(n, \alpha) = \exp\left(-2n(\alpha - \mu(x))^2\right)

    Because the distinguishing probability decays exponentially with sample size nn and margin (α−μ(x))(\alpha - \mu(x)), rare events satisfy effective zero-ϵ\epsilon local differential privacy.

  3. Knowl 3 — Unbiased Frequency Estimation and Local Differential Privacy for Non-Rare Events

    theoretical result

    For non-rare events x∈Ni={x∈Di:μi(x)≥α}x \in N_i = \{x \in \mathcal{D}_i : \mu_i(x) \ge \alpha\} over a discrete domain of size k=∣Di∣k = |\mathcal{D}_i|, AlignDP applies RAPPOR with bit flip probability p∈(0,0.5)p \in (0, 0.5), providing ϵ\epsilon-local differential privacy (ϵ\epsilon-LDP) with:

    ϵ=log⁡(1−pp)\epsilon = \log\left(\frac{1 - p}{p}\right)

    Let yjy_j be the observed fraction of user reports equal to category jj, and let q=1−p+pkq = 1 - p + \frac{p}{k}. The global aggregator computes the debiased frequency estimator μ^i(j)\hat{\mu}_i(j) as:

    μ^i(j)=yj−1k(1−q)q−1k(1−q)\hat{\mu}_i(j) = \frac{y_j - \frac{1}{k}(1 - q)}{q - \frac{1}{k}(1 - q)}

    This estimator is unbiased and its variance is bounded by:

    E[μ^i(j)]=μi(j),Var[μ^i(j)]≤p(1−p)n\mathbb{E}[\hat{\mu}_i(j)] = \mu_i(j), \quad \text{Var}[\hat{\mu}_i(j)] \le \frac{p(1 - p)}{n}

    where nn is the total number of user reports.

  4. Knowl 4 — Global Privacy Composition Across Queries

    theoretical result

    PAC indistinguishability protects rare mass (μ(x)<α \mu(x) < \alpha), but does not extend to non-rare events where adversaries succeed as sample size nn grows. Therefore, repeated query access to AlignDP must be bounded using differential privacy composition.

    For kk queries each satisfying ϵ\epsilon-local differential privacy, the cumulative global privacy loss ϵtot\epsilon_{\text{tot}} is bounded by:

    • Basic Composition: ϵtot≤kϵ\epsilon_{\text{tot}} \le k\epsilon

    • Advanced Composition (for failure parameter δ>0\delta > 0): ϵtot≤2klog⁡(1/δ) ϵ+kϵ(eϵ−1)\epsilon_{\text{tot}} \le \sqrt{2k \log(1/\delta)}\,\epsilon + k\epsilon(e^\epsilon - 1)

    The global aggregator enforces these bounds across all queries to restrict extraction.

  5. Knowl 5 — Extraction Resistance Under Repeated Query Budgets

    data/table

    AlignDP enforces a fixed ceiling on knowledge extraction under repeated adversary querying. In a categorical setup with n=1000n = 1000 users, domain size k=20k = 20 across 10 fields, rarity threshold α=0.01\alpha = 0.01 (with 4 rare categories having true mass 0.00250.0025 each), and flip probability p=0.25p = 0.25, the extraction accuracy varies with query count as follows:

    Queries 1 10 50 100
    Rare MAE 0.003 0.001 0.001 0.001
    Non-rare ρ\rho -0.08 0.85 0.99 0.99

    Rare Mean Absolute Error (MAE) stays at the noise floor (≤0.003\le 0.003), demonstrating effective zero-ϵ\epsilon shielding, while non-rare Spearman rank correlation ρ\rho saturates at 0.990.99 due to RAPPOR noise, preventing higher-fidelity reconstruction regardless of increased query counts.

  6. Knowl 6 — Empirical Utility and Error Decay Properties of AlignDP

    empirical result

    Empirical validation of AlignDP on categorical distributions (k=20k = 20, α=0.01\alpha = 0.01, p=0.25p = 0.25) confirms utility retention for frequent events alongside privacy enforcement:

    • MSE Scaling: The Mean Squared Error (MSE) of non-rare category frequency estimates decays strictly as O(1/n)O(1/n) across sample sizes n∈{200,400,600,800,1000,1500,2000}n \in \{200, 400, 600, 800, 1000, 1500, 2000\}, matching theoretical variance bounds.
    • Distributional Match: On 10,00010{,}000 evaluation samples, the Kullback-Leibler (KL) divergence between true and estimated non-rare distributions is 0.00130.0013.
    • Ranking Preservation: Identification of the top-5 frequent categories achieves 80%80\% accuracy, and overall Spearman rank correlation is ρ=0.798\rho = 0.798.
    • Empirical PAC Fit: The empirical indistinguishability of rare categories follows an exponential curve δ(n)=exp⁡(−n(α−μrare))\delta(n) = \exp\left(-n(\alpha - \mu_{\text{rare}})\right), fitting the theoretical Hoeffding-style bound.
  7. Knowl 7 — Mechanism Realization of Lock-LLM Objectives

    model/method

    AlignDP operationalizes the five Lock-LLM protection objectives at the data release mechanism interface:

    1. Un-distillable: PAC shielding conceals rare events while RAPPOR adds noise to non-rare events, causing clean knowledge distillation to fail on the privatized output.
    2. Un-finetunable: Supervised fine-tuning cannot access accurate labels due to non-rare noise and absent rare signals, preventing effective gradient optimization.
    3. Un-compressible: Outputs are randomized encodings; secondary compression removes signal further without improving upon debiased aggregate estimators.
    4. Un-editable: Every data release is mathematically auditable via PAC bounds for rare categories and calibrated noise for non-rare categories, enabling detection of unauthorized data injection.
    5. Un-usable: The central aggregator tracks cumulative privacy loss under DP composition theorems, cutting off repeated extraction queries when budget limits are reached.
  8. Knowl 8 — Limitations in Scaling AlignDP to LLM Token Sequences

    limitation

    Applying AlignDP to full-scale language model telemetry has three main limitations:

    1. Fixed Rarity Thresholding: A static rarity threshold α\alpha cannot adapt to dynamic data distributions or separate sensitive rare tokens from Zipfian statistical rarity typical of natural language vocabularies.
    2. Communication Overhead: RAPPOR encoding scales linearly with domain size kk. While feasible for small domains (k=20k = 20), the O(k)O(k) vector communication overhead becomes heavy for modern LLM vocabulary sizes exceeding 50,00050{,}000 tokens.
    3. Token Independence Assumption: AlignDP evaluates categorical tokens independently, which does not account for sequential token correlations or combinations of non-rare tokens that form rare, identifying phrases.

Coverage note — Proof sketches in Appendix B were omitted in accordance with the rule excluding derivations and intermediate proof steps. Implementation script logistics from Appendix F were omitted as they describe software environment setup rather than research findings.

References

  1. 1.Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security, pages 308–318, 2016.
  2. 2.Nicolas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting training data from diffusion models. In 32nd USENIX security symposium (USENIX Security 23), pages 5253–5270, 2023.
  3. 3.Cynthia Dwork, Aaron Roth, et al. The algorithmic foundations of differential privacy. Foundations and trends® in theoretical computer science, 9(3–4):211–407, 2014.
  4. 4.Úlfar Erlingsson, Vasyl Pihur, and Aleksandra Korolova. Rappor: Randomized aggregatable privacy-preserving ordinal response. In Proceedings of the 2014 ACM SIGSAC conference on computer and communications security, pages 1054–1067, 2014.
  5. 5.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
  6. 6.Matthew Jagielski, Nicholas Carlini, David Berthelot, Alex Kurakin, and Nicolas Papernot. High accuracy and high fidelity extraction of neural networks. In 29th USENIX security symposium (USENIX Security 20), pages 1345–1362, 2020.
  7. 7.John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. A watermark for large language models. In International Conference on Machine Learning, pages 17061–17084. PMLR, 2023.
  8. 8.Chris Liu, Yaxuan Wang, Jeffrey Flanigan, and Yang Liu. Large language model unlearning via embedding-corrupted prompts. Advances in Neural Information Processing Systems, 37:118198–118266, 2024.
  9. 9.Florian Tramèr, Fan Zhang, Ari Juels, Michael K Reiter, and Thomas Ristenpart. Stealing machine learning models via prediction {APIs}. In 25th USENIX security symposium (USENIX Security 16), pages 601–618, 2016.

Citation

MLA
Gaikwad, M. “AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs”. arXiv, 2025, http://arxiv.org/abs/2512.17251v1.
APA
Gaikwad, M. (2025). AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs. arXiv. http://arxiv.org/abs/2512.17251v1
Chicago
Gaikwad, M. 2025. “AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs”. arXiv. http://arxiv.org/abs/2512.17251v1.
Harvard
Gaikwad, M. (2025) “AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2512.17251v1.
Vancouver
1. Gaikwad M (2025) AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs. arXiv

BibTeX

@article{gaikwad2025aligndp,
  title = {AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs},
  author = {Gaikwad, Madhava},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2512.17251v1},
  eprint = {2512.17251}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/