Foundation Models and Fair Use

Peter HendersonXuechen LiDan JurafskyTatsunori HashimotoMark A. LemleyPercy Liang

article2023JMLR173 citations

Examines the legal risks of training foundation models on copyrighted data under United States fair use law and proposes technical mitigation strategies alongside policy safe harbors to prevent copyright infringement in generative outputs.

Listen

Rapid advancements in foundation models—large machine learning models trained on vast web-scraped datasets—have sparked urgent legal and ethical concerns regarding copyright infringement. Developers commonly rely on the United States fair use doctrine to justify ingesting copyrighted material without direct licenses. However, fair use protection is not guaranteed when generative models produce outputs that closely resemble protected works or compete in the same commercial markets as the original data creators.

The article evaluates how U.S. copyright and fair use case law applies to modern foundation models generating text, source code, and visual art. It demonstrates the technical vulnerabilities of current models to verbatim and non-verbatim copying, while assessing methods to mitigate intellectual property risks.

The authors conducted a legal analysis of relevant U.S. fair use precedents, intermediate copying decisions, and the Digital Millennium Copyright Act. To ground this analysis empirically, they ran experiments using the Holistic Evaluation of Language Models benchmark and commercial application programming interfaces across various large models, including versions of Generative Pre-trained Transformer models and OpenAI Codex. They evaluated data regurgitation using string matching, plagiarism-detection software across Linux kernel source code licensed under the General Public License, and named entity recognition across 10 million user image prompts.

The analysis reveals several critical findings. First, existing foundation models can regurgitate verbatim training material, such as entire children's books or hundreds of lines of source code; in code generation tests, between 0.3% and 0.91% of samples exhibited more than a 20% match to reference source files, averaging 45% to 60% overlap among flagged outputs. Second, copyright infringement under U.S. law does not require exact verbatim reproduction; non-literal copying, such as abridgments, translations, or unauthorized adaptations that extract the qualitative core of a work, also fails fair use scrutiny. Third, larger models with longer context windows show an increased tendency to extract and reconstruct substantial spans of training data. Fourth, common technical guardrails, such as basic keyword or n-gram matching, are easily bypassed using simple instruction reframing or character substitutions. Finally, statutory protections such as Digital Millennium Copyright Act safe harbors are not guaranteed for automatically generated outputs, creating direct and secondary liability risks for deployers and creators.

These findings indicate that deploying generative artificial intelligence without robust safeguards poses substantial legal, financial, and compliance risks. Overly permissive interpretations could severely harm data creators and creative labor, whereas overly restrictive rulings might ban web-scale training entirely and concentrate market power among a few incumbent tech platforms. Simple contractual disclaimers or standard terms of service do not shield deployers from strict-liability copyright infringement.

To manage these risks, practitioners should implement multi-layered mitigation architectures. At training time, organizations should deduplicate data and incorporate feedback-based learning techniques that reward transformativeness while penalizing exact extraction. At inference time, deployers must move beyond simple string matching toward semantic output filters that detect non-literal transformations, protected characters, and uncredited source code. Policymakers and courts should co-evolve legal standards with technical realities by considering meaningful safe harbors for deployers who implement verified mitigation tools, while exploring broader policy remedies to protect creative labor.

The authors acknowledge key limitations in technical mitigations: formal methods like differential privacy introduce sharp trade-offs with utility, and instance attribution remains computationally expensive at scale. Furthermore, algorithmic filters cannot fully replicate nuanced, subjective judicial determinations of fair use. While technical mitigations significantly reduce operational exposure, they cannot entirely eliminate liability or resolve the broader socioeconomic disruptions facing creative industries.

arXiv: 2303.15715
  • Paper: On the Opportunities and Risks of Foundation Models, Rishi Bommasani et al. (2021). This foundational survey frames the capabilities and societal risks of foundation models, providing the conceptual basis for the source’s analysis of copyright risks in their development and deployment.
Cover for Foundation Models and Fair Use

Abstract

Existing foundation models are trained on copyrighted material. Deploying these models can pose both legal and ethical risks when data creators fail to receive appropriate attribution or compensation. In the United States and several other countries, copyrighted content may be used to build foundation models without incurring liability due to the fair use doctrine. However, there is a caveat: If the model produces output that is similar to copyrighted data, particularly in scenarios that affect the market of that data, fair use may no longer apply to the output of the model. In this work, we emphasize that fair use is not guaranteed, and additional work may be necessary to keep model development and deployment squarely in the realm of fair use. First, we survey the potential risks of developing and deploying foundation models based on copyrighted content. We review relevant U.S. case law, drawing parallels to existing and potential applications for generating text, source code, and visual art. Experiments confirm that popular foundation models can generate content considerably similar to copyrighted material. Second, we discuss technical mitigations that can help foundation models stay in line with fair use. We argue that more research is needed to align mitigation strategies with the current state of the law. Third, we suggest that the law and technical mitigations should co-evolve. For example, coupled with other policy mechanisms, the law could more explicitly consider safe harbors when strong technical tools are used to mitigate infringement harms. This co-evolution may help strike a balance between intellectual property and innovation, which speaks to the original goal of fair use. But we emphasize that the strategies we describe here are not a panacea and more work is needed to develop policies that address the potential harms of foundation models.

Table of Contents

  • 1. Introduction
  • 2. Foundation Models and Fair Use
  • 2.1 Foundation Models
  • 2.2 Definitions and Roles
  • 2.3 Fair Use
  • 2.4 Natural Language Text
  • 2.5 Code
  • 2.6 Generated Images
  • 3. Additional Considerations
  • 4. Technical Mitigation
  • 4.1 Data and Output Filtering
  • 4.2 Instance Attribution
  • 4.3 Differentially Private Training
  • 4.4 Learning from Human Feedback
  • 5. Forward-looking Agenda
  • 6. Related Work
  • 7. Conclusion
  • Acknowledgements
  • References
  • Appendix A. Experimental Setup
  • A.1 Book Extraction Experiments
  • A.2 Code Extraction Experiments
  • Appendix B. Examples of Reproduced Code
  • Appendix C. Additional Breakdowns of Prompt Entities
  • Appendix D. Additional Qualitative Examples

Knowls

  1. Knowl 1 — Fair use does not automatically protect foundation-model outputs

    definition

    Under U.S. copyright law, fair use is a case-specific defense assessed through four factors: (1) the purpose and character of the use, including commerciality and transformation; (2) the nature of the copyrighted work; (3) the amount and substantiality used; and (4) the effect on the work’s actual or potential market. The factors interact, and none supplies a categorical rule for foundation models. Training may have a different purpose from the source material, especially for non-generative tasks, but that does not establish that a generated output is fair use. An output that reproduces protected expression or competes in the source work’s market may face greater risk. The paper focuses primarily on output-related risk; the status of model training and model parameters remains unsettled.

  2. Knowl 2 — Transformation must be assessed at the level of protected expression

    model/method

    The paper’s case-law analysis treats transformation as more than low word- or pixel-level overlap: a use can retain the protected “heart” of a work even after substantial surface changes. Text translations, abridgements, and stories reusing protected characters can therefore remain risky; parody may support fair use when it comments on the source, while satire generally needs a stronger justification for borrowing. Facts, ideas, common plot structures, and an artist’s general style are not themselves protected expression, although distinctive expressive elements may be. Code has a narrower scope of copyright protection because functional elements are not protected, but substantial copying of expressive structure can still matter. For generated images, copying distinctive elements is riskier than producing sufficiently different work in a general artistic style. Across modalities, market overlap and the purpose of the output also affect the assessment.

  3. Knowl 3 — Text models can reproduce substantial copyrighted passages under targeted prompting

    empirical result

    The authors tested popular language models using random 125-token excerpts from popular books in Books3, random excerpts from the broader books corpus, and prompts based on the title and author of Dr. Seuss’s Oh, the Places You’ll Go!. Sampling used temperature T=0.2T=0.2; the benchmarked extraction analysis included outputs up to 1,024 tokens. Random excerpts generally produced little verbatim text, while popular works were more extractable. OPT-175B reproduced Oh, the Places You’ll Go! verbatim. In manual prompting, ChatGPT produced the entire story in two interactions using only its title and author; it also reproduced the first three pages of Harry Potter and the Sorcerer’s Stone before deviating. GPT-4 reproduced the Dr. Seuss story and, after a letter-substitution instruction, reproduced about three and a half chapters of Harry Potter and the Sorcerer’s Stone before deviating. The authors report approximate context windows of 4,000 tokens for ChatGPT and 8,000 for GPT-4, and associate longer extraction in these tests with the larger window and model capability. These results are specific to the tested models and prompts, not a general extraction rate for all deployments.

  4. Knowl 4 — Codex models sometimes generate substantial overlaps with GPL-licensed code

    empirical result

    The authors prompted three OpenAI Codex models—code-cushman-001, code-davinci-001, and code-davinci-002—with randomly selected Linux-kernel function signatures whose reference implementations exceeded 20 lines. For each signature, they sampled 10 completions at temperature 0.20.2, generated up to 1,800 tokens, and used MossPlus to select and measure the completion with the greatest overlap against the GPL-2.0 reference implementation. Defining a large match as a MossPlus overlap above 20%, the reported frequencies were 0.30% for code-cushman-001, 0.70% for code-davinci-001, and 0.91% for code-davinci-002. Among those large matches, average overlap was 57.83%, 59.43%, and 48.94%, respectively. MossPlus detects fuzzy as well as exact overlap and can produce false positives; manual inspection identified some false positives, including cases with extensive variable assignments. Thus, large overlaps were uncommon in this setup but occurred for all three models.

  5. Knowl 5 — Image-generation prompts often name artists and commercial works

    empirical result

    The authors analyzed 10 million prompts posted by community members to the Stable Diffusion Discord channel, using named-entity recognition as a proxy for users’ intended subjects and styles. Person names were the most common entity type; artist Greg Rutkowski appeared approximately 1.2 million times. The plotted entity breakdown also shows artist names among the frequently cited people and commercial franchises such as Star Wars, Twin Peaks, and Game of Thrones among frequently identified works. This pattern suggests that many prompts request particular artist styles, while some target specific existing works; the latter may pose greater copyright risk if the generated image reproduces protected expression. Named-entity counts indicate prompt references, not whether outputs actually copied a work or infringed copyright.

  6. Knowl 6 — Filtering should address both training data and generated outputs

    model/method

    The paper presents filtering as a two-stage risk-reduction strategy. At training time, developers can exclude copyrighted or restrictively licensed material, respect opt-out mechanisms such as robots.txt where applicable, and deduplicate repeated examples because repeated exposure can increase memorization. These measures do not guarantee fair use: underlying items may have different licenses from the dataset collection, attribution requirements can remain impractical, and deduplication may miss semantically similar examples or repeated protected elements. At deployment, output filters can block verbatim reproduction, but exact-match filters are inadequate when risky similarity survives paraphrase, translation, style transfer, or other surface changes; they can also be bypassed and add inference cost. The paper therefore calls for filters that assess higher-level expression and transformation, distinguish facts and parody where possible, and account for task and output structure. Such filtering should reduce risk, not purport to decide fair use conclusively.

  7. Knowl 7 — Instance attribution could identify sources and support output controls

    model/method

    Instance-attribution methods estimate how individual training examples, or groups of examples, contributed to a model prediction; existing approaches include leave-one-out retraining and influence functions, while retrieval-augmented models can provide attribution as part of their design. For generated content, attribution could identify whether an output depends heavily on a particular copyrighted work, enabling a filter against outputs dominated by one source. It could also support an output-level credit page listing influential works and their licensing information, or help identify training instances for deletion after a takedown request. The paper cautions that post-hoc methods can be computationally expensive, inaccurate on realistic models, and difficult to use at runtime. Attribution may also expose which training data drove an output, creating tensions with litigation-risk management.

  8. Knowl 8 — Differential privacy and near access-freeness depend on meaningful example units

    model/method

    Differentially private training aims to make a model difficult to distinguish from one trained without any particular example, which can limit memorization and make claims that an output was copied from one specific example harder to establish. Its usefulness for copyright depends on defining an example at a legally meaningful level: if a work is duplicated or represented across many examples, example-level privacy may not prevent its reproduction. The privacy–utility trade-off, training cost, and choice of privacy parameters are additional obstacles. The paper also discusses near access-freeness (NAF), a proposed guarantee under which a model trained on copyrighted material generates similarly to a model trained without that material. The described NAF algorithms require each copyrighted work to appear in at most one, or a constant number of, training examples; surface-level deduplication may not satisfy this condition. Both approaches therefore require careful, potentially semantic deduplication and do not by themselves settle whether a particular use is fair use.

  9. Knowl 9 — Human-feedback training can incorporate copyright-aware judgments

    model/method

    The paper proposes adapting learning from human feedback so that annotators evaluate not only helpfulness or instruction-following but also whether a generation is sufficiently transformative from its closest copyrighted source. For example, annotators could be shown the closest source material and asked to flag generations that are not sufficiently transformed; this signal could then inform the reward model and subsequent training. The proposal is intended to discourage outputs such as verbatim responses to requests to reproduce a book while retaining useful capabilities. It is not a certified guarantee: feedback can be inconsistent, the reward objective can be misspecified, and instruction-following rewards could otherwise favor reproducing requested copyrighted material.

  10. Knowl 10 — Copyright policy and technical mitigations should evolve together

    limitation

    The paper argues that technical safeguards and law should co-evolve. Strong, objectively assessable mitigation efforts could be considered in fair-use or indirect-liability analyses, and policymakers could clarify DMCA protections or create a model-specific safe harbor conditioned on sufficiently strong safeguards. The proposal is meant to avoid both extremes: treating all generative-model use as acceptable regardless of harm, or broadly excluding unlicensed data in ways that could concentrate model development among organizations with extensive licensed datasets. Safeguards should not be overbroad: filtering should preserve uses such as factual responses, parody, and limited quotation where appropriate. Neither technical measures nor fair use resolve all effects on creators, labor, or other stakeholders, so complementary policy responses may still be needed.

Coverage note — Omitted the paper’s detailed actor-by-actor liability scenarios, extended DMCA and copyright-management-information analysis, sovereign-immunity discussion, and non-U.S. legal comparisons; these are substantial but highly context-dependent legal surveys rather than standalone technical or empirical contributions in this selection.

References

  1. 1.Rohan Anil, Badih Ghazi, Vineet Gupta, Ravi Kumar, and Pasin Manurangsi. Large-scale differentially private bert. arXiv preprint arXiv:2108.01624, 2021.
  2. 2.Clark D Asay. Transformative use in software. Stan. L. Rev. Online, 70:9, 2017.
  3. 3.Clark D Asay, Arielle Sloan, and Dean Sobczak. Is transformative use eating the world. BCL Rev., 61:905, 2020.
  4. 4.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022a.
  5. 5.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022b.
  6. 6.Andy Baio. Ai data laundering: How academic and nonprofit researchers shield tech companies from accountability. https://waxy.org/2022/09/ai-data-laundering-how-academic-and-nonprofit-researchers-shield-tech-companies-from-accountability/, 2022.
  7. 7.Stephanie Plamondon Bair. Rational faith: The utility of fairness in copyright. BUL Rev., 97:1487, 2017.
  8. 8.John Bandy and Nicholas Vincent. Addressing "documentation debt" in machine learning: A retrospective datasheet for bookcorpus. In J. Vanschoren and S. Yeung, editors, Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021. URL https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/file/54229abfcfa5649e7003b83dd4755294-Paper-round1.pdf.
  9. 9.Taylor B Bartholomew. The death of fair use in cyberspace: Youtube and the problem with content id. Duke L. & Tech. Rev., 13:66, 2014.
  10. 10.Samyadeep Basu, Philip Pope, and Soheil Feizi. Influence functions in deep learning are fragile. arXiv preprint arXiv:2006.14651, 2020.
  11. 11.Barton Beebe. An empirical study of us copyright fair use opinions updated, 1978-2019. NYU J. Intell. Prop. & Ent. L., 10:1, 2020.
  12. 12.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623, 2021.
  13. 13.Stella Biderman, Kieran Bicheno, and Leo Gao. Datasheet for the pile. arXiv preprint arXiv:2201.07311, 2022.
  14. 14.Andrea Bittau, Úlfar Erlingsson, Petros Maniatis, Ilya Mironov, Ananth Raghunathan, David Lie, Mitch Rudominer, Ushasree Kode, Julien Tinnes, and Bernhard Seefeld. Prochlo: Strong privacy for analytics in the crowd. In Proceedings of the 26th symposium on operating systems principles, pages 441–459, 2017.
  15. 15.Joshua Bloch and Pamela Samuelson. Some misconceptions about software in the copyright literature. In CSLAW’22: Proceedings of the 2nd ACM Symposium on Computer Science and Law, 2022.
  16. 16.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. Language (technology) is power: A critical survey of “bias" in nlp. arXiv preprint arXiv:2005.14050, 2020.
  17. 17.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  18. 18.Benjamin Boroughf. The next great youtube: improving content id to foster creativity, cooperation, and fair compensation. Alb. LJ Sci. & Tech., 25:95, 2015.
  19. 19.Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. Machine unlearning. In 2021 IEEE Symposium on Security and Privacy (SP), pages 141–159. IEEE, 2021.
  20. 20.Peter Brown and Bob Mercer. Twenty years of bitext. https://www.cs.jhu.edu/~post/bitext/, 2013.
  21. 21.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  22. 22.Michelle Brownlee. Safeguarding style: What protection is afforded to visual artists by the copyright and trademark laws. Colum. L. Rev., 93:1157, 1993.
  23. 23.Dan L Burk. Algorithmic fair use. U. Chi. L. Rev., 86:283, 2019.
  24. 24.Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. The secret sharer: Evaluating and testing unintended memorization in neural networks. In 28th USENIX Security Symposium (USENIX Security 19), pages 267–284, 2019.
  25. 25.Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650, 2021.
  26. 26.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646, 2022.
  27. 27.Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting training data from diffusion models, 2023. URL https://arxiv.org/abs/2301.13188.
  28. 28.Michael W Carroll. Copyright and the progress of science: Why text and data mining is lawful. UC Davis L. Rev., 53:893, 2019.
  29. 29.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  30. 30.Albert Cheu, Adam Smith, Jonathan Ullman, David Zeber, and Maxim Zhilyaev. Distributed differential privacy via shuffling. In Annual International Conference on the Theory and Applications of Cryptographic Techniques, pages 375–403. Springer, 2019.
  31. 31.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  32. 32.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  33. 33.Dane S Ciolino. Rethinking the compatibility of moral rights and fair use. Wash. & Lee L. Rev., 54:33, 1997.
  34. 34.Samuel J Coe. The story of a character: Establishing the limits of independent copyright protection for literary characters. Chi.-Kent L. Rev., 86:1305, 2011.
  35. 35.Soham De, Leonard Berrada, Jamie Hayes, Samuel L Smith, and Borja Balle. Unlocking high-accuracy differentially private image classification through scale. arXiv preprint arXiv:2204.13650, 2022.
  36. 36.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  37. 37.Bolin Ding, Janardhan Kulkarni, and Sergey Yekhanin. Collecting telemetry data privately. Advances in Neural Information Processing Systems, 30, 2017.
  38. 38.F Jay Dougherty. All the world’s not a stooge: The transformativeness test for analyzing a first amendment defense to a right of publicity claim against distribution of a work of art. Colum. JL & Arts, 27:1, 2003.
  39. 39.Dr. Seuss Enters., L.P. v. ComicMix LLC. 983 F.3d 443, 9th Cir. 2020. URL https://www.copyright.gov/fair-use/summaries/drseuss-comicmix-9thcir2020.pdf.
  40. 40.Cynthia Dwork, Aaron Roth, et al. The algorithmic foundations of differential privacy. Foundations and Trends® in Theoretical Computer Science, 9(3–4):211–407, 2014.
  41. 41.Niva Elkin-Koren. Fair use by design. UCLA L. Rev., 64:1082, 2017.
  42. 42.Niva Elkin-Koren and Neil Weinstock Netanel. Transplanting fair use across the globe: A case study testing the credibility of us opposition. Hastings LJ, 72:1121, 2020.
  43. 43.Alexander v. Take-Two Interactive Software, Inc. 489 F. Supp. 3d 812, S.D. Ill. 2020.
  44. 44.Andersen et al. v. Stability AI et al. 3:23-cv-00201, N.D. Cal. 2023.
  45. 45.Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith. 598 U.S. __, 2023.
  46. 46.Authors Guild, Inc. v. Google, Inc. 804 F.3d 202, 2d Cir. 2015.
  47. 47.Chabon v. OpenAI, Inc.,. 3:23-cv-04625, (N.D. Cal.), 2023.
  48. 48.Concord Music Group, Inc. v. Anthropic PBC. 3:23-cv-01092, 2023.
  49. 49.DOE 1 v. GitHub, Inc. 4:22-cv-06823, N.D. Cal. 2022.
  50. 50.Dr. Seuss Enters., LP v. Penguin Books USA, Inc. 109 F.3d 1394, 9th Cir. 1997.
  51. 51.Fox News Network, LLC v. TVEyes, Inc. Nos. 15-3885, 15-3886, 2d Cir. Feb. 27, 2018.
  52. 52.Getty Images (US), Inc. v. Stability AI, Inc. 1:99-mc-09999, D. Del. 2023.
  53. 53.Google LLC v. Oracle America Inc. 141 S. Ct. 1183, 593 U.S., 209 L. Ed. 2d 311, 2021.
  54. 54.Hall v. Swift. No. 18-55426, 9th Cir., Oct. 28, 2019.
  55. 55.Kadrey v. Meta Platforms, Inc. 3:23-cv-03417, 2023.
  56. 56.Nihon Keizai Shimbun, Inc. v. Comline Bus. Data Inc. 166 F.3d 65, 69, 2d Cir. 1999.
  57. 57.Paramount Pictures Corp. v. Axanar Prods., Inc. No. 2:15-cv-09938-RGK-E, C.D. Cal. Jan. 3, 2017.
  58. 58.Penguin Grp. (USA), Inc. v. Am. Buddha. No. 4:13-cv-02075-JGZ, D. Ariz. May 11, 2015.
  59. 59.Penguin Random House LLC, et al. v. Frederik Colting and Melissa Medina, d/b/a Moppet Books. No. 17-cv-386, S.D.N.Y. Sept. 8, 2017.
  60. 60.The New York Times Company v. Microsoft Corporation. 1:23-cv-11195, 2023.
  61. 61.Tremblay v. OpenAI, Inc.,. 23-cv-03416-AMO, (N.D. Cal.), 2023.
  62. 62.Warner Bros. Entertainment Inc. v. RDR Books. 575 F. Supp. 2d 513, S.D.N.Y. 2008.
  63. 63.Úlfar Erlingsson, Vasyl Pihur, and Aleksandra Korolova. Rappor: Randomized aggregatable privacy-preserving ordinal response. In Proceedings of the 2014 ACM SIGSAC conference on computer and communications security, pages 1054–1067, 2014.
  64. 64.European Parliament and Council of the European Union. Proposal for a regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts. 21-01-2024 at 17h11, 2021. DRAFT [Final draft as updated on 21/01].
  65. 65.European Union. Directive (EU) 2019/790 of the European Parliament and of the Council of 17 April 2019 on copyright and related rights in the Digital Single Market and amending Directives 96/9/EC and 2001/29/EC, 2019.
  66. 66.Vitaly Feldman and Chiyuan Zhang. What neural networks memorize and why: Discovering the long tail via influence estimation. Advances in Neural Information Processing Systems, 33:2881–2891, 2020.
  67. 67.Carlos Muñoz Ferrandis, Danish Contractor, Huu Nguyen, and David Lansky. The BigScience RAIL License. https://bigscience.huggingface.co/blog/the-bigscience-rail-license, 2022.
  68. 68.Giorgio Franceschelli and Mirco Musolesi. Copyright in generative deep learning. Data & Policy, 4, 2022.
  69. 69.Iason Gabriel. Artificial intelligence, values, and alignment. Minds and machines, 30(3):411–437, 2020.
  70. 70.Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al. Predictability and surprise in large generative models. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 1747–1764, 2022.
  71. 71.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  72. 72.Amirata Ghorbani and James Zou. Data shapley: Equitable valuation of data for machine learning. In International Conference on Machine Learning, pages 2242–2251. PMLR, 2019.
  73. 73.Amirata Ghorbani, Abubakar Abid, and James Zou. Interpretation of neural networks is fragile. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 3681–3688, 2019.
  74. 74.Jessica L Gillotte. Copyright infringement in AI-generated artworks. UC Davis L. Rev., 53:2655, 2019.
  75. 75.Aaron Gokaslan and Vanya Cohen. Openwebtext corpus, 2019.
  76. 76.James Grimmelmann. Copyright for literate robots. Iowa L. Rev., 101:657, 2015.
  77. 77.Andres Guadamuz. Do androids dream of electric copyright? comparative analysis of originality in artificial intelligence generated works. Intellectual property quarterly, 2017.
  78. 78.Andres Guadamuz. A scanner darkly: Copyright infringement in artificial intelligence inputs and outputs. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4371204, 2023.
  79. 79.Chuan Guo, Brian Karrer, Kamalika Chaudhuri, and Laurens van der Maaten. Bounding training data reconstruction in private (deep) learning. arXiv preprint arXiv:2201.12383, 2022.
  80. 80.Kelvin Guu, Tatsunori B Hashimoto, Yonatan Oren, and Percy Liang. Generating sentences by editing prototypes. Transactions of the Association for Computational Linguistics, 6:437–450, 2018.
  81. 81.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR, 2020.
  82. 82.Peter Henderson, Koustuv Sinha, Nicolas Angelard-Gontier, Nan Rosemary Ke, Genevieve Fried, Ryan Lowe, and Joelle Pineau. Ethical challenges in data-driven dialogue systems. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, pages 123–129, 2018.
  83. 83.Peter Henderson, Mark S Krass, Lucia Zheng, Neel Guha, Christopher D Manning, Dan Jurafsky, and Daniel E Ho. Pile of law: Learning responsible data filtering from the law and a 256gb open-source legal dataset. arXiv preprint arXiv:2207.00220, 2022.
  84. 84.Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt. Unsolved problems in ml safety. arXiv preprint arXiv:2109.13916, 2021.
  85. 85.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. An empirical analysis of compute-optimal large language model training. In Advances in Neural Information Processing Systems.
  86. 86.Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry. Data-models: Predicting predictions from training data. arXiv preprint arXiv:2202.00622, 2022.
  87. 87.Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A. Choquette-Choo, and Nicholas Carlini. Preventing verbatim memorization in language models gives a false sense of privacy, 2022. URL https://arxiv.org/abs/2210.17546.
  88. 88.Russell W Jacobs. Gutters and hyperlinks: The dmca and proper position of copyright management information. Nw. J. Tech. & Intell. Prop., 11:xxi, 2012.
  89. 89.Steven D Jamar and Christen Glenn. When the author owns the world: Copyright issues arising from monetizing fan fiction. Tex. A&M L. Rev., 1:959, 2013.
  90. 90.Japan. 2018 amendment to the japanese copyright act, 2018.
  91. 91.Yacine Jernite, Huu Nguyen, Stella Biderman, Anna Rogers, Maraim Masoud, Valentin Danchev, Samson Tan, Alexandra Sasha Luccioni, Nishant Subramani, Isaac Johnson, et al. Data governance in the age of large-scale data-driven language technology. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 2206–2222, 2022.
  92. 92.Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, and Costas J Spanos. Towards efficient data valuation based on the shapley value. In The 22nd International Conference on Artificial Intelligence and Statistics, pages 1167–1176. PMLR, 2019.
  93. 93.Nikhil Kandpal, Eric Wallace, and Colin Raffel. Deduplicating training data mitigates privacy risks in language models. arXiv preprint arXiv:2202.06539, 2022.
  94. 94.Sayash Kapoor and Arvind Narayanan. Artists can now opt out of generative ai. it’s not enough. https://aisnakeoil.substack.com/p/artists-can-now-opt-out-of-generative, 2023.
  95. 95.Daniel Kifer and Ashwin Machanavajjhala. No free lunch in data privacy. In Proceedings of the 2011 ACM SIGMOD International Conference on Management of data, pages 193–204, 2011.
  96. 96.Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. The stack: 3 tb of permissively licensed source code. Preprint, 2022.
  97. 97.Pang Wei Koh and Percy Liang. Understanding black-box predictions via influence functions. In International conference on machine learning, pages 1885–1894. PMLR, 2017.
  98. 98.Jaewoo Lee and Chris Clifton. How much is enough? choosing ε for differential privacy. In International Conference on Information Security, pages 325–340. Springer, 2011.
  99. 99.Jooyoung Lee, Thai Le, Jinghui Chen, and Dongwon Lee. Do language models plagiarize? arXiv preprint arXiv:2203.07618, 2022.
  100. 100.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499, 2021.
  101. 101.Katherine Lee, A Feder Cooper, and James Grimmelmann. Talkin”bout ai generation: Copyright and the generative-ai supply chain. arXiv preprint arXiv:2309.08133, 2023.
  102. 102.Kalev Leetaru. Common crawl and unlocking web archives for research. https://www.forbes.com/sites/kalevleetaru/2017/09/28/common-crawl-and-unlocking-web-archives-for-research/?sh=7b9bbae63b83, 2017.
  103. 103.Mark A Lemley and Bryan Casey. Remedies for robots. The University of Chicago Law Review, 86 (5):1311–1396, 2019.
  104. 104.Mark A Lemley and Bryan Casey. Fair learning. Tex. L. Rev., 99:743, 2020.
  105. 105.Amanda Levendowski. How copyright law can fix artificial intelligence’s implicit bias problem. Wash. L. Rev., 93:579, 2018.
  106. 106.Amanda Levendowski. Resisting face surveillance with copyright law. NCL Rev., 100:1015, 2021.
  107. 107.Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. Large language models can be strong differentially private learners. arXiv preprint arXiv:2110.05679, 2021.
  108. 108.Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. Competition-level code generation with alphacode. arXiv preprint arXiv:2203.07814, 2022.
  109. 109.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
  110. 110.Daryl Lim. Ai, equity, and the ip gap. SMU Law Review, 75(4):815, 2022.
  111. 111.Susuk Lim. A survey of the dcma’s copyright management information protections: The dmca’s cmi landscape after all headline news and mcclatchey. Wash. JL Tech. & Arts, 6:297, 2010.
  112. 112.Jacqueline D Lipton. Copyright and the commercialization of fanfiction. Hous. L. Rev., 52:425, 2014.
  113. 113.Keoni Mahelona, Gianna Leoni, Suzanne Duncan, and Miles Thompson. Openai’s whisper is another case study in colonisation. Papa Reo, 2023. URL https://blog.papareo.nz/whisper-is-another-case-study-in-colonisation/.
  114. 114.Mango v. BuzzFeed, Inc. 970 F.3d 167, 2020.
  115. 115.Tony Mason, Ada Gavrilovska, and David A Joyner. Collaboration versus cheating: Reducing code plagiarism in an online ms computer science program. In Proceedings of the 50th ACM Technical Symposium on Computer Science Education, pages 1004–1010, 2019.
  116. 116.Mavrix Photographs, LLC v. Livejournal Inc. 873 F.3d 1045, 9th Cir. 2017.
  117. 117.Sancho McCann. Copyright throughout a creative ai pipeline. Canadian JL & Tech, 2021.
  118. 118.Stephen McJohn and Ian McJohn. Fair use and machine learning. NEULR, 12:99, 2020.
  119. 119.Eric Miraglia. Privacy that works for everyone. 2019. URL https://blog.google/technology/safety-security/privacy-everyone-io/.
  120. 120.Pamela Mishkin, Lama Ahmad, Miles Brundage, Gretchen Krueger, and Girish Sastry. Dall·e 2 preview - risks and limitations. 2022.
  121. 121.Gary Myers. Muddy waters: Fair use implications of google llc v. oracle america, inc. Nw. J. Tech. & Intell. Prop., 19:155, 2021.
  122. 122.Alex Nichol. Dalle 2 pre-training mitigations. https://openai.com/blog/dall-e-2-pre-training-mitigations/, 2022.
  123. 123.Tyler T Ochoa. Dr. seuss, the juice and fair use revisited: Two decades of parody and satire in copyright law. IDEA, 59:233, 2018.
  124. 124.OpenAI. Gpt-4 technical report, 2023. URL https://arxiv.org/abs/2303.08774.
  125. 125.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  126. 126.Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna. Data and its (dis) contents: A survey of dataset development and use in machine learning research. Patterns, 2(11):100336, 2021.
  127. 127.Shira Perlmutter. Copyright and state sovereign immunity: A report of the register of copyrights. 2021.
  128. 128.Pouya Pezeshkpour, Sarthak Jain, Byron C Wallace, and Sameer Singh. An empirical comparison of instance attribution methods for nlp. arXiv preprint arXiv:2104.04128, 2021.
  129. 129.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022.
  130. 130.Shawn Presser. Books3. https://twitter.com/theshawwn/status/1320282149329784833, 2020.
  131. 131.Public Resource. The General Index. https://archive.org/details/GeneralIndex, 2021.
  132. 132.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  133. 133.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints, 2019.
  134. 134.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021.
  135. 135.Lisa P Ramsey. Brandjacking on social networks: Trademark infringement by impersonation of markholders. Buff. L. Rev., 58:851, 2010.
  136. 136.In re Aimster Copyright Litig. 334 F.3d 643, 645-646, 7th Cir. 2003.
  137. 137.Trevor G Reed. Fair use as cultural appropriation. Cal. L. Rev., 109:1373, 2021.
  138. 138.Jerome H Reichman and Paul F Uhlir. Database protection at the crossroads: recent development and their impact on science and technology. Berkeley Tech. LJ, 14:793, 1999.
  139. 139.Religious Technology Center v. Netcom On-line Communication Services, Inc. 907 F. Supp. 1361 (N.D. Cal. 1995), 1995.
  140. 140.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2021.
  141. 141.Betsy Rosenblatt. Considering the role of fairness in copyright fair use. Houston Law Review, 61(2):261–293, 2023.
  142. 142.Matthew Sag. Copyright and copy-reliant technology. Nw. UL Rev., 103:1607, 2009.
  143. 143.Matthew Sag. Predicting fair use. Ohio St. LJ, 73:47, 2012.
  144. 144.Matthew Sag. The new legal landscape for text mining and machine learning. J. Copyright Soc’y USA, 66:291, 2018.
  145. 145.Matthew Sag. Copyright safety for generative ai. Houston Law Review, 2024.
  146. 146.Pamela Samuelson. Text and data mining of in-copyright works: is it legal? Communications of the ACM, 64(11):20–22, 2021.
  147. 147.Pamela Samuelson and Clark D Asay. Saving software’s fair use future. Harv. JL & Tech., 31:535, 2017.
  148. 148.Tom Sander, Pierre Stock, and Alexandre Sablayrolles. Tan without a burn: Scaling laws of dp-sgd. arXiv preprint arXiv:2210.03403, 2022.
  149. 149.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207, 2021.
  150. 150.Sarah Scheffler, Eran Tromer, and Mayank Varia. Formalizing human ingenuity: A quantitative framework for coyright law’s substantial similarity. arXiv preprint arXiv:2206.01230, 2022.
  151. 151.Saul Schleimer, Daniel S Wilkerson, and Alex Aiken. Winnowing: local algorithms for document fingerprinting. In Proceedings of the 2003 ACM SIGMOD international conference on Management of data, pages 76–85, 2003.
  152. 152.Christoph Schmon, Filip Lukáš, and Corynne McSherry. The eu’s copyright directive is still about filters, but eu’s top court limits its use. https://www.eff.org/deeplinks/2022/05/eus-copyright-directive-still-about-filters-eus-top-court-limits-its-use, 2022.
  153. 153.Christoph Schuhmann. Laion-400-million open dataset, 2021.
  154. 154.John Schulman, Barret Zoph, Christina Kim, Jacob Hilton, Jacob Menick, Jiayi Weng, Juan Felipe Ceron Uribe, Liam Fedus, Luke Metz, Michael Pokorny, Rapha Gontijo Lopes, Shengjia Zhao, Arun Vijayvergiya, Eric Sigler, Adam Perelman, Chelsea Voss, Mike Heaton, Joel Parish, Dave Cummings, Rajeev Nayak, Valerie Balcom, David Schnurr, Tomer Kaftan, Chris Hallacy, Nicholas Turley, Noah Deutsch, and Vik Goel. Chatgpt: Optimizing language models for dialogue. https://openai.com/blog/chatgpt/.
  155. 155.Benjamin LW Sobel. Artificial intelligence’s fair use crisis. Colum. JL & Arts, 41:45, 2017.
  156. 156.Anders Søgaard et al. Revisiting methods for finding influential examples. arXiv preprint arXiv:2111.04683, 2021.
  157. 157.Solid Oak Sketches, LLC v. 2K Games, Inc. 449 F. Supp. 3d 333, S.D.N.Y. 2020.
  158. 158.Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Diffusion art or digital forgery? investigating data replication in diffusion models. arXiv preprint arXiv:2212.03860, 2022.
  159. 159.Daxton R Stewart. Rise of the copyleft trolls: When photographers sue after creative commons licenses go awry. Ohio St. Tech. LJ, 18:333, 2021.
  160. 160.Mukund Sundararajan and Amir Najmi. The many shapley values for model explanation. In International conference on machine learning, pages 9269–9278. PMLR, 2020.
  161. 161.Xiyin Tang. Copyright’s techno-pessimist creep. Fordham L. Rev., 90:1151, 2021.
  162. 162.Xiyin Tang. The class action as licensing and reform device. Columbia Law Review, 122(6):1627–1690, 2022.
  163. 163.Allen v. Cooper, 2020.
  164. 164.Arista Records LLC v. Lime Group LLC. 784 F. Supp. 2d 398, S.D.N.Y. 2011.
  165. 165.Associated Press v. Meltwater U.S. Holdings, Inc. No. 1:12-cv-01087, 156, S.D.N.Y. Mar. 21, 2013.
  166. 166.Authors Guild v. Google Inc. 804 F.3d 202, 220, 221, 2d. Cir. 2015.
  167. 167.Cambridge University Press v. Mark P. Becker. No. 1:08-cv-01425-ODE, N.D. Ga. Mar. 31, 2016.
  168. 168.Campbell v. Acuff-Rose Music, Inc. 510 U.S. 569, 1994.
  169. 169.Computer Associates Intern., Inc. v. Altai, Inc. 982 F.2d 693, 2d Cir. 1992.
  170. 170.Davis v. Elec. Arts Inc. 775 F.3d 1172, 9th Cir. 2015.
  171. 171.DC Comics v. Towle. 802 F.3d 1012, 9th Cir. 2015.
  172. 172.Field v. Google, Inc. 412 F.Supp. 2d 1106, D. Nev. 2006.
  173. 173.Golan v. Holder. 565 U.S. 302, 2012.
  174. 174.Harper & Row, Publishers, Inc. v. Nation Enters. 471 U.S. 539, 562, 1985.
  175. 175.Harper & Row v. Nation Enterprises. 471 U.S. 539, 1985.
  176. 176.Hart v. Elec. Arts, Inc. 717 F.3d 141, 3d Cir. 2013.
  177. 177.Kelly v. Arriba Soft Corp. 77 F.Supp.2d 1116, 1122, aff’d and rev’d in part on other grounds, 336 F.3d 811 (9th Cir. 2003), C.D. Cal. 1999.
  178. 178.Kirk Kara Corp. v. W. Stone & Metal Corp. No.CV 20-1931-DMG (EX), 2020 WL 5991503, C.D. Cal. Aug. 14, 2020.
  179. 179.Metro-Goldwyn-Mayer Studios, Inc. v. Grokster, Ltd. 454 F. Supp. 2d 966, 974, C.D. Cal. 2006.
  180. 180.Righthaven LLC v. Choudhry. No. 2:10-CV-2155 JCM PAL, 2011 WL 2976800, (D. Nev. July 21, 2011).
  181. 181.Salinger v. Colting. 607 F.3d 68, 2d Cir. 2010.
  182. 182.Sega Enterprises Ltd. v. Accolade, Inc. 977 F. 2d 1510, 9th Cir. 1992.
  183. 183.Sony Computer Entertainment v. Connectix Corp. 203 F. 3d 496, 9th Cir. 2000.
  184. 184.Tiffany v. eBay. 600 F.3d 93, 2d Cir. 2010.
  185. 185.Mikael Thalen. Artists fed up with ai-image generators use mickey mouse to goad copyright lawsuits. DailyDot, 2022. URL https://www.dailydot.com/debug/ai-art-protest-disney-characters-mickey-mouse/.
  186. 186.The Guardian. The top 100 bestselling books of all time. https://www.theguardian.com/news/datablog/2012/aug/09/best-selling-books-all-time-fifty-shades-grey-compare, 2012.
  187. 187.U.K. Intellectual Property Office. Copyright exemptions. https://www.gov.uk/guidance/exceptions-to-copyright#text-and-data-mining-for-non-commercial-research, 2021.
  188. 188.U.S. Copyright Office. Faq. https://www.copyright.gov/help/faq/faq-general.html, 2022.
  189. 189.James Vincent. Getty images bans ai-generated content over fears of legal challenges. https://www.theverge.com/2022/9/21/23364696/getty-images-ai-ban-generated-artwork-illustration-copyright, 2022.
  190. 190.James Vincent. Getty images is suing the creators of ai art tool stable diffusion for scraping its content. https://www.theverge.com/2023/1/17/23558516/ai-art-copyright-stable-diffusion-getty-images-lawsuit, 2023.
  191. 191.Eugene Volokh. Freedom of speech and the right of publicity. Hous. L. Rev., 40:903, 2003.
  192. 192.Nikhil Vyas, Sham Kakade, and Boaz Barak. Provable copyright protection for generative models, 2023. URL https://arxiv.org/abs/2302.10870.
  193. 193.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  194. 194.Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al. Taxonomy of risks posed by language models. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 214–229, 2022.
  195. 195.Reid Kress Weisbord. A copyright right of publicity. Fordham L. Rev., 84:2803, 2015.
  196. 196.Williams-Sonoma, Inc. v. Amazon.com, Inc. 3:18-cv-07548, N.D. Cal. 2021. URL https://www.courtlistener.com/docket/8418854/125/williams-sonoma-inc-v-amazoncom-inc/.
  197. 197.Da Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi, Huseyin A Inan, Gautam Kamath, Janardhan Kulkarni, Yin Tat Lee, Andre Manoel, Lukas Wutschitz, et al. Differentially private fine-tuning of language models. arXiv preprint arXiv:2110.06500, 2021.
  198. 198.Weichen Yu, Tianyu Pang, Qian Liu, Chao Du, Bingyi Kang, Yan Huang, Min Lin, and Shuicheng Yan. Bag of tricks for training data extraction from language models. arXiv preprint arXiv:2302.04460, 2023.
  199. 199.Eliezer Yudkowsky. The ai alignment problem: why it is hard, and where to start. Symbolic Systems Distinguished Speaker, 2016.
  200. 200.Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini. Counterfactual memorization in neural language models. arXiv preprint arXiv:2112.12938, 2021.
  201. 201.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  202. 202.Albert Ziegler. Github copilot research recitation. https://github.blog/2021-06-30-github-copilot-research-recitation/, 2021.
  203. 203.Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Henderson, P., et al. “Foundation Models and Fair Use”. Journal of Machine Learning Research, vol. 24, no. 400, 2023, pp. 1–9, https://www.jmlr.org/papers/v24/23-0569.html.
APA
Henderson, P., Li, X., Jurafsky, D., Hashimoto, T., Lemley, M. A., & Liang, P. (2023). Foundation Models and Fair Use. Journal of Machine Learning Research, 24(400), 1–79. https://www.jmlr.org/papers/v24/23-0569.html
Chicago
Henderson, P., X. Li, D. Jurafsky, T. Hashimoto, M. A. Lemley, and P. Liang. 2023. “Foundation Models and Fair Use”. Journal of Machine Learning Research 24 (400): 1–79. https://www.jmlr.org/papers/v24/23-0569.html.
Harvard
Henderson, P. et al. (2023) “Foundation Models and Fair Use”, Journal of Machine Learning Research, 24(400), pp. 1–79. Available at: https://www.jmlr.org/papers/v24/23-0569.html.
Vancouver
1. Henderson P, Li X, Jurafsky D, Hashimoto T, Lemley MA, Liang P (2023) Foundation Models and Fair Use. Journal of Machine Learning Research 24:1–79

BibTeX

@article{JMLR:v24:23-0569,
  author  = {Peter Henderson and Xuechen Li and Dan Jurafsky and Tatsunori Hashimoto and Mark A. Lemley and Percy Liang},
  title   = {Foundation Models and Fair Use},
  journal = {Journal of Machine Learning Research},
  year    = {2023},
  volume  = {24},
  number  = {400},
  pages   = {1--79},
  url     = {http://jmlr.org/papers/v24/23-0569.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/