LLM Dataset Inference: Did you train on my dataset?

Pratyush MainiHengrui JiaNicolas PapernotAdam Dziedzic

article2024NeurIPS114 citations

Demonstrates that standard membership inference attacks on large language models fail under matched data distributions and introduces a statistically grounded dataset inference framework that reliably detects copyright infringement across collections of texts.

Listen

The rapid deployment of large language models trained on vast web scrapes has sparked widespread legal and copyright disputes regarding the unauthorized use of proprietary content. Existing technical solutions predominantly rely on membership inference attacks, which attempt to determine whether specific, individual sentences or text sequences were used during training. However, these methods are fundamentally unreliable for modern models trained on trillions of tokens, leaving creators, developers, and arbiters without a dependable mechanism for data accountability.

The article demonstrates the systemic failures of existing sentence-level membership inference attacks and introduces dataset inference, a statistically grounded framework designed to reliably determine whether a curated collection of text—such as an entire book or body of work—was included in a model's training data.

To evaluate existing approaches, the researchers conducted controlled experiments using the Pythia model family and the Pile dataset, which contains over 20 distinct domain subsets with clearly separated training and validation splits. By testing across models ranging from 410 million to 12 billion parameters, the researchers showed how prior claims of membership inference success were confounded by temporal shifts in data. The proposed dataset inference method aggregates 52 distinct membership inference metrics, trains a lightweight linear model on a subset of suspect and validation data to learn domain-specific feature weights, and executes a statistical hypothesis test on held-out text sequences.

The findings demonstrate that sentence-level membership inference attacks perform no better than random guessing (achieving area-under-the-curve scores near 0.5) when evaluated on properly matched data from the same distribution. Prior reported successes merely detected broader conceptual or temporal shifts rather than actual training inclusion. Furthermore, no single attack metric functions consistently across different text domains. In contrast, dataset inference achieved strong statistical detection across all evaluated domains, consistently yielding p-values well below 0.1, while producing zero false positives when comparing separate subsets of unseen validation data. The framework proved highly sample-efficient, requiring as few as 100 to 1,000 text sequences to establish training data inclusion, with detection confidence strengthening as model parameter size increased.

These results establish that attempting to prove copyright infringement or training inclusion at the individual sentence level is technically unviable. By shifting the objective to aggregate dataset detection, organizations and legal arbiters gain an effective, mathematically rigorous tool to resolve intellectual property disputes without risking high rates of false accusations.

Organizations developing or deploying large language models should establish gray-box auditing procedures that allow authorized arbiters to query model output loss. Researchers and compliance teams evaluating training provenance should discontinue isolated sentence-level membership tests in favor of aggregated statistical dataset tests across strictly independent and identically distributed data splits.

The effectiveness of dataset inference relies on key operational conditions: the claimant must possess a private, unseen validation dataset drawn from the exact same distribution as the suspect data, and the arbiter must have gray-box access to inspect model loss. Where these conditions are fulfilled, confidence in the framework's ability to attribute training data ownership remains exceptionally high.

Cover for LLM Dataset Inference: Did you train on my dataset?

Abstract

The proliferation of large language models (LLMs) in the real world has come with a rise in copyright cases against companies for training their models on unlicensed data from the internet. Recent works have presented methods to identify if individual text sequences were members of the model’s training data, known as membership inference attacks (MIAs). We demonstrate that the apparent success of these MIAs is confounded by selecting non-members (text sequences not used for training) belonging to a different distribution from the members (e.g., temporally shifted recent Wikipedia articles compared with ones used to train the model). This distribution shift makes membership inference appear successful. However, most MIA methods perform no better than random guessing when discriminating between members and non-members from the same distribution (e.g., in this case, the same period of time). Even when MIAs work, we find that different MIAs succeed at inferring membership of samples from different distributions. Instead, we propose a new dataset inference method to accurately identify the datasets used to train large language models. This paradigm sits realistically in the modern-day copyright landscape, where authors claim that an LLM is trained over multiple documents (such as a book) written by them, rather than one particular paragraph. While dataset inference shares many of the challenges of membership inference, we solve it by selectively combining the MIAs that provide positive signal for a given distribution, and aggregating them to perform a statistical test on a given dataset. Our approach successfully distinguishes the train and test sets of different subsets of the Pile with statistically significant p-values < 0.1, without any false positives.

Table of Contents

  • 1 Introduction
  • 2 Background and Baselines
  • 2.1 Metrics for LLM Membership Inference
  • 3 Problem Setup
  • 4 Failure of Membership Inference
  • 5 LLMDataset Inference
  • 5.1 Procedure for the LLM Dataset Inference
  • 5.2 Assumptions for Dataset Inference
  • 5.3 Experimental Details
  • 5.4 Analysis and Results with Dataset Inference
  • 6 Discussions
  • 7 Acknowledgements
  • References
  • A Broader Impact
  • B Compute
  • C Additional Experiments

Knowls

  1. Knowl 1 — LLM Dataset Inference Procedure

    algorithm

    LLM Dataset Inference is a multi-stage statistical testing protocol designed to determine whether a collection of text documents (a suspect dataset) was included in the training corpus of a large language model fθf_\theta. The method leverages multiple membership inference attack (MIA) metrics, learns distribution-specific importance weights via linear regression on a partitioned subset of data, and conducts a hypothesis test on held-out evaluation samples.

    Input: Suspect dataset DsusD_{sus}, private validation dataset DvalD_{val}, suspect language model fθf_\theta, number of seeds nn
    Output: Combined statistical pp-value pcombinedp_{\text{combined}} testing whether fθf_\theta was trained on DsusD_{sus}
    for seed i=1i = 1 to nn do
        Partition DsusD_{sus} randomly into disjoint splits AsusA_{sus} and BsusB_{sus}
        Partition DvalD_{val} randomly into disjoint splits AvalA_{val} and BvalB_{val}
        
        Compute MM distinct MIA feature scores for each sample in AsusA_{sus} and AvalA_{val} using fθf_\theta
        Normalize each feature across Asus∪AvalA_{sus} \cup A_{val} to comparable scales
        Replace the top 2.5% and bottom 2.5% outlier values of each feature with that feature's mean
        
        Train a linear regression model g(x)=wTx+bg(x) = w^T x + b for 1000 updates where:
            Target y=0y = 0 for suspect samples x∈Asusx \in A_{sus}
            Target y=1y = 1 for validation samples x∈Avalx \in A_{val}
        
        Compute MM normalized MIA feature scores for samples in BsusB_{sus} and BvalB_{val} using fθf_\theta
        Compute membership prediction scores s=g(x)s = g(x) for each x∈Bsus∪Bvalx \in B_{sus} \cup B_{val}
        Discard the top 2.5% and bottom 2.5% outlier score values from BsusB_{sus} and BvalB_{val}
        
        Perform a one-sided two-sample tt-test between scores Ssus={g(x)∣x∈Bsus}S_{sus} = \{g(x) \mid x \in B_{sus}\} and Sval={g(x)∣x∈Bval}S_{val} = \{g(x) \mid x \in B_{val}\} to compute pip_i under H0:μ(Ssus)≥μ(Sval)H_0: \mu(S_{sus}) \ge \mu(S_{val}) vs H1:μ(Ssus)<μ(Sval)H_1: \mu(S_{sus}) < \mu(S_{val})
    end for
    Compute pcombined=1−exp⁡(∑i=1nlog⁡(1−pi))p_{\text{combined}} = 1 - \exp\left(\sum_{i=1}^n \log(1 - p_i)\right)
    return pcombinedp_{\text{combined}}
  2. Knowl 2 — Šidák Aggregation for Dependent Hypothesis Tests in Dataset Inference

    equation

    When dataset inference is repeated across nn different random splits of the dataset into training splits AA and evaluation splits BB, the resulting tt-test pp-values p1,p2,…,pnp_1, p_2, \dots, p_n are mutually dependent because the underlying example subsets overlap. To combine these dependent one-sided pp-values while conservatively controlling the Type-1 error rate, the aggregated pp-value pcombinedp_{\text{combined}} is computed via the Šidák-based aggregation formula:

    pcombined=1−exp⁡(∑i=1nlog⁡(1−pi))p_{\text{combined}} = 1 - \exp\left( \sum_{i=1}^{n} \log(1 - p_i) \right)

    where pi∈[0,1]p_i \in [0, 1] is the pp-value obtained from the ii-th random split, and nn is the total number of evaluated splits (e.g., n=10n = 10).

  3. Knowl 3 — Assumptions for LLM Dataset Inference

    assumption

    The LLM dataset inference framework operates under three necessary conditions:

    1. IID Data Distribution: The suspect dataset (hypothesized to be training data) and the validation dataset (unseen by the model) must be identically and independently distributed from the same distribution (e.g., drafts versus published text from the same author or era) to ensure tests are not confounded by temporal or domain shift.
    2. Private Validation Data: The validation dataset must be strictly private and inaccessible to the model trainer, ensuring zero leakage into the training set of the suspect LLM.
    3. Gray-Box Model Access: The arbiter evaluating the dispute must have gray-box access to the suspect LLM fθf_\theta, defined as the ability to compute output token probabilities, cross-entropy losses, or perplexities for arbitrary text sequences, without requiring model weights or gradients.
  4. Knowl 4 — Temporal Shift Confounder in LLM Membership Inference

    empirical result

    Prior assertions that single-sequence membership inference attacks (MIAs), such as MIN-K% PROB\text{MIN-}K\%\text{ PROB}, succeed on large language models are confounded by temporal distribution shifts in benchmark datasets. When evaluated on WikiMIA (where training data consists of pre-2023 Wikipedia articles and non-member data consists of post-2023 articles), MIN-K% PROB\text{MIN-}K\%\text{ PROB} achieves an apparent Area Under the ROC Curve (AUC) near 0.70.7.

    However, when evaluated in a confounder-free setting using identically distributed (IID) train and validation splits from the Pile dataset (both from the pre-2023 distribution), the AUC of MIN-K% PROB\text{MIN-}K\%\text{ PROB} drops to approximately 0.50.5 (equivalent to random guessing), with individual random subsets producing AUC values varying between 0.40.4 and 0.70.7.

    Furthermore, in a reversed setup where the pre-2023 Pile validation split (unseen by the Pythia model during training) is tested as the suspect set against post-2023 Wikipedia as non-members, MIN-K% PROB\text{MIN-}K\%\text{ PROB} achieves an AUC of 0.70.7 in falsely classifying the unseen pre-2023 validation text as training members. This demonstrates that such MIAs distinguish temporal concept shifts rather than individual sample membership.

  5. Knowl 5 — Cross-Distribution Brittleness of Single LLM Membership Inference Attacks

    empirical result

    Across 20 domain-specific subsets of the Pile dataset (including Arxiv, Books3, PhilPapers, PubMed Abstracts, and GitHub), no single membership inference attack (MIA) consistently distinguishes training members from non-members across all distributions. An individual attack metric that achieves high AUC on one domain often drops below 0.50.5 on another:

    • Perturbation-based MIA via synonym substitution achieves an AUC of approximately 0.700.70 on PhilPapers, but drops to 0.320.32 on PubMed Abstracts.
    • Loss and perplexity thresholding drop to an AUC near 0.400.40 on Arxiv and OpenWebText2, generating high false positive rates by assigning higher membership likelihood to predictable validation samples than to training members.

    Because different text domains exhibit divergent feature sensitivities, combining multiple diverse MIA metrics adaptively via domain-specific regression weighting is necessary for reliable attribution.

  6. Knowl 6 — Pile Dataset Attribution Efficacy and False Positive Control

    empirical result

    LLM Dataset Inference was evaluated using Pythia models (410M, 1.4B, 6.9B, and 12B parameters) across all 20 domain subsets of the Pile dataset using 1000 sequences per split:

    • True Positive Detection: When comparing training splits against unseen validation splits, dataset inference achieved statistically significant combined pp-values satisfying p<0.1p < 0.1 across all 20 subsets for Pythia-12B, with 19 of the 20 subsets achieving p<10−3p < 10^{-3} (and frequently below 10−3010^{-30}).
    • False Positive Control: When evaluating null conditions by comparing two disjoint 500-sample partitions of the unseen validation split against each other, the resulting pp-values were all strictly greater than 0.500.50 (ranging from 0.580.58 on arXiv to 1.001.00 on 17 subsets), demonstrating zero false positive attributions at the α=0.1\alpha = 0.1 significance threshold.
  7. Knowl 7 — Impact of Model Size, Sample Size, and Deduplication on Dataset Inference

    empirical result

    Ablation studies on dataset inference across the Pythia model family establish the following scaling behaviors:

    1. Sample Efficiency: For more than half of the Pile subsets, fewer than 100 query sequences are required to reject the null hypothesis at p<0.1p < 0.1. 1000 query sequences are sufficient to achieve p<0.1p < 0.1 across all 20 domain subsets.
    2. Model Parameter Scaling: Attribution confidence improves monotonically with parameter count (from 410M to 1.4B, 6.9B, and 12B parameters), with pp-value distributions concentrating further below the 0.10.1 threshold for larger models due to increased training memorization.
    3. Effect of Training Deduplication: Non-deduplicated models exhibit stronger dataset membership signals (lower and more concentrated pp-values at 500 query points) than deduplicated models, as repeated exposures during pretraining amplify model memorization signatures across intermediate MIA metrics.
  8. Knowl 8 — MIA Feature Preprocessing Pipeline for Dataset Inference

    model/method

    To prevent extreme loss values from distorting linear regression weights and tt-test statistics, dataset inference applies a two-stage outlier adjustment and normalization scheme to 52 candidate MIA metrics:

    1. Feature Normalization: Each raw MIA metric vector is z-score normalized across the combined stage-1 dataset Asus∪AvalA_{sus} \cup A_{val} to place heterogeneous metrics (e.g., cross-entropy loss, perplexity ratios, zlib compression ratios) on comparable scales.
    2. Classifier Input Outlier Adjustment: For each normalized feature column, values falling into the top 2.5% and bottom 2.5% percentiles of the distribution are replaced with the feature's empirical mean before training the linear regressor.
    3. Hypothesis Test Outlier Removal: After mapping evaluation samples BsusB_{sus} and BvalB_{val} to scalar confidence scores through the trained regressor g(x)g(x), the top 2.5% and bottom 2.5% of score values are discarded entirely before executing the two-sample tt-test.

    Ablations demonstrate that alternative strategies—such as mean correction or clipping without normalization—induce false positive attributions (p<0.1p < 0.1) when comparing two un-trained validation splits.

  9. Knowl 9 — Methodological Guidelines for LLM Membership Inference

    model/method

    To prevent experimental confounders in future evaluations of membership inference attacks (MIAs) on language models, researchers must follow four experimental standards:

    1. IID Member/Non-Member Splits: Training members and non-members must be drawn from identical underlying distributions and timeframes to eliminate temporal shifts, stylistic drift, and topic discrepancies.
    2. Multi-Split Replication: Experiments must be replicated across multiple distinct random data partitions and initializations to measure and report variance in classification metrics.
    3. Cross-Domain Evaluation: MIA techniques must be evaluated over diverse domain distributions rather than single narrow benchmarks to avoid overestimating generalization.
    4. Symmetric Error and False Positive Evaluation: Evaluations must include negative controls, such as testing disjoint subsets of non-member data or swapping member/non-member labels on unseen data, to confirm that high AUC reflects sample-level membership rather than concept familiarity.
  10. Knowl 10 — Limitations and Practical Constraints of LLM Dataset Inference

    limitation

    The dataset inference framework possesses specific operational and institutional constraints:

    1. Availability of Unseen IID Validation Sets: The method strictly requires that the claimant possess a private dataset drawn from the exact same distribution as the suspect dataset (e.g., unpublished drafts or rejected iterations of a published text). In settings where no unreleased, IID counter-dataset exists, the method cannot be executed.
    2. Gray-Box Loss Access: The protocol requires numerical sequence losses or token log-probabilities; it cannot operate under pure black-box text-generation access if the LLM API suppresses token likelihoods.
    3. Institutional and Legal Enforcement: Querying the suspect LLM requires legal arbitration or regulatory frameworks to compel proprietary model providers to permit evaluation queries.

Coverage note — None was omitted; all primary contributions, mathematical definitions, algorithms, empirical evaluations across the Pile and Pythia models, and experimental guidelines have been fully represented.

References

  1. 1.Gemini, https://gemini.google.com/. URL https://gemini.google.com/.
  2. 2.Introducing meta llama 3: The most capable openly available llm to date, https://ai.meta.com/blog/meta-llama-3/. URL https://ai.meta.com/blog/meta-llama-3/.
  3. 3.Openai, https://openai.com. URL https://openai.com/.
  4. 4.Introducing dbrx: A new state-of-the-art open llm, https://www.databricks.com/blog/introducing-dbrx-new-state-art-open-llm. URL https://www.databricks.com/blog/introducing-dbrx-new-state-art-open-llm.
  5. 5.Getty images vs. stability ai: A landmark case in copyright and ai, 2023. URL https://www.bakerlaw.com/getty-images-v-stability-ai/.
  6. 6.The times sues openai and microsoft over a.i. use of copyrighted work https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html. 2023. URL https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html.
  7. 7.Sarah silverman and authors sue openai and meta over copyright infringement. 2023. URL https://www.nytimes.com/2023/07/10/arts/sarah-silverman-lawsuit-openai-meta.html.
  8. 8.Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang. Fast-detectGPT: Efficient zero-shot detection of machine-generated text via conditional probability curvature. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Bpcgcr8E8Z.
  9. 9.Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. Pythia: a suite for analyzing large language models across training and scaling. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023.
  10. 10.Richard E Blakesley, Sati Mazumdar, Mary Amanda Dew, Patricia R Houck, Gong Tang, Charles F Reynolds III, and Meryl A Butters. Comparisons of methods for multiple hypothesis testing in neuropsychological research. Neuropsychology, 23(2):255–264, 2009. doi: 10.1037/a0012850.
  11. 11.Morton B Brown. 400: A method for combining non-independent, one-sided tests of significance. Biometrics, pages 987–992, 1975.
  12. 12.Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650. USENIX Association, August 2021. ISBN 978-1-939133-24-3. URL https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting.
  13. 13.Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramèr. Membership inference attacks from first principles. In 2022 IEEE Symposium on Security and Privacy (SP), pages 1897–1914, 2022. doi: 10.1109/SP46214.2022.9833649.
  14. 14.Kaustubh D. Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Srivastava, Samson Tan, Tongshuang Wu, Jascha Sohl-Dickstein, Jinho D. Choi, Eduard Hovy, Ondrej Dusek, Sebastian Ruder, Sajant Anand, Nagender Aneja, Rabin Banjade, Lisa Barthe, Hanna Behnke, Ian Berlot-Attwell, Connor Boyle, Caroline Brun, Marco Antonio Sobrevilla Cabezudo, Samuel Cahyawijaya, Emile Chapuis, Wanxiang Che, Mukund Choudhary, Christian Clauss, Pierre Colombo, Filip Cornell, Gautier Dagan, Mayukh Das, Tanay Dixit, Thomas Dopierre, Paul-Alexis Dray, Suchitra Dubey, Tatiana Ekeinhor, Marco Di Giovanni, Rishabh Gupta, Rishabh Gupta, Louanes Hamla, Sang Han, Fabrice Harel-Canada, Antoine Honore, Ishan Jindal, Przemyslaw K. Joniak, Denis Kleyko, Venelin Kovatchev, Kalpesh Krishna, Ashutosh Kumar, Stefan Langer, Seungjae Ryan Lee, Corey James Levinson, Hualou Liang, Kaizhao Liang, Zhexiong Liu, Andrey Lukyanenko, Vukosi Marivate, Gerard de Melo, Simon Meoni, Maxime Meyer, Afnan Mir, Nafise Sadat Moosavi, Niklas Muennighoff, Timothy Sum Hon Mun, Kenton Murray, Marcin Namysl, Maria Obedkova, Priti Oli, Nivranshu Pasricha, Jan Pfister, Richard Plant, Vinay Prabhu, Vasile Pais, Libo Qin, Shahab Raji, Pawan Kumar Rajpoot, Vikas Raunak, Roy Rinberg, Nicolas Roberts, Juan Diego Rodriguez, Claude Roux, Vasconcellos P. H. S., Ananya B. Sai, Robin M. Schmidt, Thomas Scialom, Tshephisho Sefara, Saqib N. Shamsi, Xudong Shen, Haoyue Shi, Yiwen Shi, Anna Shvets, Nick Siegel, Damien Sileo, Jamie Simon, Chandan Singh, Roman Sitelew, Priyank Soni, Taylor Sorensen, William Soto, Aman Srivastava, KV Aditya Srivatsa, Tony Sun, Mukund Varma T, A Tabassum, Fiona Anting Tan, Ryan Teehan, Mo Tiwari, Marie Tolkiehn, Athena Wang, Zijian Wang, Gloria Wang, Zijie J. Wang, Fuxuan Wei, Bryan Wilie, Genta Indra Winata, Xinyi Wu, Witold Wydmanski, Tianbao Xie, Usama Yaseen, M. Yee, Jing Zhang, and Yue Zhang. Nl-augmenter: A framework for tasksensitive natural language augmentation, 2023.
  15. 15.Haonan Duan, Adam Dziedzic, Nicolas Papernot, and Franziska Boenisch. Flocks of stochastic parrots: Differentially private prompt learning for large language models. In Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS), 2023a.
  16. 16.Haonan Duan, Adam Dziedzic, Mohammad Yaghini, Nicolas Papernot, and Franziska Boenisch. On the privacy risk of in-context learning. In The 61st Annual Meeting Of The Association For Computational Linguistics, 2023b.
  17. 17.Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. Do membership inference attacks work on large language models? 2024.
  18. 18.Adam Dziedzic, Haonan Duan, Muhammad Ahmad Kaleem, Nikita Dhawan, Jonas Guan, Yannis Cattan, Franziska Boenisch, and Nicolas Papernot. Dataset inference for self-supervised models. In NeurIPS (Neural Information Processing Systems), 2022.
  19. 19.Ronen Eldan and Yuanzhi Li. Tinystories: How small can language models be and still speak coherent english? arXiv preprint arXiv:2305.07759, 2023.
  20. 20.Jean-loup Gailly and Mark Adler. zlib compression library. 2004. URL http://www.dspace.cam.ac.uk/handle/1810/3486.
  21. 21.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  22. 22.James T Kost and Michael P McDermott. Combining dependent p-values. Statistics & Probability Letters, 60(2):183–190, 2002.
  23. 23.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.577. URL https://aclanthology.org/2022.acl-long.577.
  24. 24.Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee. Textbooks are all you need ii: phi-1.5 technical report. arXiv preprint arXiv:2309.05463, 2023.
  25. 25.Pratyush Maini, Mohammad Yaghini, and Nicolas Papernot. Dataset inference: Ownership resolution in machine learning. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=hvdKKV2yt7T.
  26. 26.Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schoelkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick. Membership inference attacks against language models via neighbourhood comparison. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023, pages 11330–11343, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.719. URL https://aclanthology.org/2023.findings-acl.719.
  27. 27.Xiao-Li Meng. Posterior predictive p-values. The annals of statistics, 22(3):1142–1160, 1994.
  28. 28.Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A Smith, and Luke Zettlemoyer. Silo language models: Isolating legal risk in a nonparametric datastore. arXiv preprint arXiv:2308.04430, 2023.
  29. 29.Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn. Detectgpt: zero-shot machine-generated text detection using probability curvature. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023.
  30. 30.Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, and Tatsunori Hashimoto. Proving test set contamination for black-box language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=KS8mIvetg2.
  31. 31.Noorjahan Rahman and Eduardo Santacana. Beyond fair use: Legal risk evaluation for training llms on copyrighted text. 2023. URL https://genlaw.org/CameraReady/57.pdf.
  32. 32.Ludger Rüschendorf. Random variables with maximum sums. Advances in Applied Probability, 14(3):623–632, 1982.
  33. 33.Avital Shafran, Shmuel Peleg, and Yedid Hoshen. Membership inference attacks are easier on difficult problems. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 14820–14829, October 2021.
  34. 34.Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. Detecting pretraining data from large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=zWqr3MQuNs.
  35. 35.R. Shokri, M. Stronati, C. Song, and V. Shmatikov. Membership inference attacks against machine learning models. In 2017 IEEE Symposium on Security and Privacy (SP), pages 3–18, Los Alamitos, CA, USA, may 2017. IEEE Computer Society. doi: 10.1109/SP.2017.41. URL https://doi.ieeecomputersociety.org/10.1109/SP.2017.41.
  36. 36.Thomas Steinke, Milad Nasr, and Matthew Jagielski. Privacy auditing with one (1) training run. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=f38EY21lBw.
  37. 37.Vladimir Vovk and Ruodu Wang. Combining p-values via averaging. Biometrika, 107(4):791–808, 2020.
  38. 38.Xiaodong Wu, Ran Duan, and Jianbing Ni. Unveiling security, privacy, and ethical concerns of chatgpt. Journal of Information and Intelligence, 2023.
  39. 39.S. Yeom, I. Giacomelli, M. Fredrikson, and S. Jha. Privacy risk in machine learning: Analyzing the connection to overfitting. In 2018 IEEE 31st Computer Security Foundations Symposium (CSF), pages 268–282, Los Alamitos, CA, USA, jul 2018. IEEE Computer Society. doi: 10.1109/CSF.2018.00027. URL https://doi.ieeecomputersociety.org/10.1109/CSF.2018.00027.
  40. 40.Zbyněk Šidák. Rectangular confidence regions for the means of multivariate normal distributions. Journal of the American Statistical Association, 62(318):626–633, 1967. ISSN 01621459, 1537274X. URL http://www.jstor.org/stable/2283989.

Citation

MLA
Maini, P., et al. “LLM Dataset Inference: Did You Train on My Dataset?”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 124069–92, https://proceedings.neurips.cc/paper_files/paper/2024/file/e01519b47118e2f51aa643151350c905-Paper-Conference.pdf.
APA
Maini, P., Jia, H., Papernot, N., & Dziedzic, A. (2024). LLM Dataset Inference: Did you train on my dataset?. Advances in Neural Information Processing Systems, 37, 124069–124092. https://proceedings.neurips.cc/paper_files/paper/2024/file/e01519b47118e2f51aa643151350c905-Paper-Conference.pdf
Chicago
Maini, P., H. Jia, N. Papernot, and A. Dziedzic. 2024. “LLM Dataset Inference: Did You Train on My Dataset?”. Advances in Neural Information Processing Systems 37: 124069–92. https://proceedings.neurips.cc/paper_files/paper/2024/file/e01519b47118e2f51aa643151350c905-Paper-Conference.pdf.
Harvard
Maini, P. et al. (2024) “LLM Dataset Inference: Did you train on my dataset?”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 124069–124092. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/e01519b47118e2f51aa643151350c905-Paper-Conference.pdf.
Vancouver
1. Maini P, Jia H, Papernot N, Dziedzic A (2024) LLM Dataset Inference: Did you train on my dataset?. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 124069–124092

BibTeX

@inproceedings{maini2024llm,
  title = {LLM Dataset Inference: Did you train on my dataset?},
  author = {Maini, Pratyush and Jia, Hengrui and Papernot, Nicolas and Dziedzic, Adam},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {124069-124092},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/e01519b47118e2f51aa643151350c905-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors