Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline

Shivani KapaniaStephanie BallardAlex KesslerJennifer Wortman Vaughan

article2025Conference on Fairness, Accountability and Transparency22 citations

Reveals how industry practitioners use synthetic data throughout AI development pipelines, identifying critical challenges in demographic representation, output control, and validation to provide practical guidance for responsible use.

Listen

As generative artificial intelligence development expands, engineering teams face significant bottlenecks in acquiring large-scale, high-quality human datasets due to high costs, strict privacy regulations, and overall data scarcity. To overcome these constraints, organizations are increasingly turning to synthetic data—machine-generated content and automated evaluation scores produced by large "auxiliary" generative models. However, organizational practices, standards, and policy frameworks have struggled to keep pace with this rapid shift, leaving major uncertainties regarding the reliability, safety, and governance of these systems.

The article aims to empirically examine why and how practitioners integrate synthetic data across modern AI development pipelines, identify the core challenges they face during generation and validation, and evaluate the resulting ethical and sociotechnical risks.

To conduct this evaluation, the authors performed a qualitative study involving 29 participants across 14 United States-based organizations between May and August 2024. The research unfolded in two phases: the first consisted of semi-structured interviews with 19 active AI practitioners focusing on workflows and technical challenges, while the second involved 10 responsible AI experts who analyzed the broader societal, governance, and safety implications using real-world use-case vignettes.

The findings reveal that auxiliary models are now deeply embedded across every phase of AI development, serving not only to generate training and test sets but also to simulate interactive user sessions and score model outputs (acting as an automated judge). Practitioners report that synthetic workflows provide massive efficiency gains; for example, manual red-teaming often requires skilled experts spending up to 45 minutes to craft a single test sample, whereas automated auxiliary models generate test cases nearly instantly at minimal inference cost. However, generating controlled data remains problematic because auxiliary models are brittle, highly sensitive to minor prompt changes, and prone to creating inaccurate caricatures or stereotypes rather than authentic representations of underrepresented populations. Crucially, validation practices remain a major bottleneck: despite acknowledging data quality risks, most teams rely on superficial "spot-checking" or manual "eyeballing" because market pressures prioritize rapid scaling over rigorous verification. Furthermore, practitioners frequently "chain" models—using the same auxiliary architecture to both generate and evaluate outputs—which creates systemic feedback loops and increases the risk of long-term model degradation.

These findings indicate that while synthetic data drastically lowers development costs and accelerates release timelines, it introduces hidden systemic risks. Over-reliance on synthetic evaluations can provide false confidence regarding system safety, while removing real human subjects from datasets eliminates avenues for individuals to contest data usage or exercise agency. Additionally, using foundation models to generate data concentrates market power among a few large cloud and model providers, reinforcing cultural dominance and narrowing output diversity.

To mitigate these risks, the article recommends establishing domain-specific standards for synthetic data use, differentiating between high-risk applications (such as medical tools) that require strict data operationalization and low-risk environments where noise is tolerable. Organizations must shift away from ad-hoc spot-checking toward standardized, multi-method validation protocols and implement formal documentation detailing prompts, parameters, model choices, and rejected iterations. Regulators and industry leaders should also enforce transparency disclosures for synthetic datasets and incorporate expert or community consultation frameworks when modeling affected demographic groups.

Readers should interpret these insights within the context of the study's scope: the participant sample was restricted to United States-based professionals working primarily on text-based systems. While the evidence offers a strong, confident qualitative picture of current industry workflows, quantitative validation metrics and concrete best-practice frameworks across multimodal and global settings remain emerging areas that warrant ongoing experimentation.

arXiv: 2501.18493
Cover for Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline

Abstract

Alongside the growth of generative AI, we are witnessing a surge in the use of synthetic data across all stages of the AI development pipeline. It is now common practice for researchers and practitioners to use one large generative model (which we refer to as an auxiliary model) to generate synthetic data that is used to train or evaluate another, reconfiguring AI workflows and reshaping the very nature of data. While scholars have raised concerns over the risks of synthetic data, policy guidance and best practices for its responsible use have not kept up with these rapidly evolving industry trends, in part because we lack a clear picture of current practices and challenges. Our work aims to address this gap. Through 29 interviews with AI practitioners and responsible AI experts, we examine the expanding role of synthetic data in AI development. Our findings reveal how auxiliary models are now widely used across the AI development pipeline. Practitioners describe synthetic data as crucial for addressing data scarcity and providing a competitive edge, noting that evaluation of generative AI systems at scale would be infeasible without auxiliary models. However, they face challenges controlling the outputs of auxiliary models, generating data that accurately depict underrepresented groups, and scaling data validation practices that are based primarily on manual inspection. We detail general limitations of and ethical considerations for synthetic data and conclude with a proposal of concrete steps towards the development of best practices for its responsible use.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Synthetic Data Usage in Modern AI Development
  • 2.2 Critical Studies of Synthetic Data
  • 2.3 AI Development Practices
  • 3 Methods
  • 4 Findings
  • 4.1 The Expanding Role of Synthetic Data
  • 4.1.1 Use Cases Across the AI Pipeline.
  • 4.1.2 Perceived Promises of Synthetic Data.
  • 4.2 Generating Synthetic Data
  • 4.2.1 Desiderata for Synthetic Data.
  • 4.2.2 Approaches to Data Generation.
  • 4.2.3 Challenges with Synthetic Data Generation.
  • 4.3 Validating Synthetic Data
  • 4.3.1 Approaches to Data Validation.
  • 4.3.2 Challenges in Data Validation.
  • 4.4 General Limitations and Ethical Considerations of Synthetic Data
  • 4.4.1 ‘Chaining’ Auxiliary Models.
  • 4.4.2 Removed Avenues for Exercising Agency.
  • 4.4.3 Stereotyping and Cultural Hegemony.
  • 4.4.4 Organizational Pressure to Scale.
  • 5 Discussion
  • 5.1 Interrogating Impacts to the AI Supply Chain
  • 5.2 Toward Considerations for Responsible Use of Synthetic Data
  • 6 Conclusion
  • 7 End Matter
  • References
  • A Additional Study Details
  • A.1 Interview Protocol for Phase 1
  • A.1.1 Introduction
  • A.1.2 ML pipeline
  • A.1.3 Model training
  • A.1.4 Model evaluation
  • A.1.5 Prompt-generated data
  • A.1.6 Understanding synthetic data practices
  • A.1.7 Responsible AI considerations
  • A.2 Interview Protocol and Vignettes for Phase 2
  • A.2.1 Introduction
  • A.2.2 Reflecting on vignettes
  • A.2.3 Responsible AI considerations

Knowls

  1. Knowl 1 — Auxiliary models are used across training and evaluation

    empirical result

    In interviews with AI practitioners, auxiliary generative models—models used to create data for another, primary model—appeared at nearly every stage of AI development. Practitioners used their outputs to create or augment pretraining corpora, make specialized fine-tuning datasets, and distill knowledge into smaller models. For evaluation, auxiliary models generated fixed test prompts, simulated multi-turn user interactions and adversarial inputs, and assigned labels or scores to primary-model outputs. Some practitioners initially did not regard model-generated scores as synthetic data, but recognized them as synthetic evaluation labels on reflection. These practices increasingly substitute model-generated data and judgments for data and labels traditionally produced by people.

  2. Knowl 2 — Interview study of synthetic-data practices

    experimental setup

    The study comprised 29 semi-structured interviews with participants from 14 organizations in the United States. The first phase interviewed 19 practitioners involved in generative-AI development or evaluation; the second interviewed 10 responsible-AI experts with experience in dataset creation or model evaluation. The second phase used hypothetical scenarios about generating fine-tuning data, evaluating harms through simulated interactions and model scoring, and simulating organizational user data. Interviews were conducted in English by video call from May to August 2024, lasted 40–60 minutes, and were analyzed using an inductive, iterative reflexive thematic-analysis process. The first author developed codes and the research team refined themes; the analysis surfaced 762 first-level codes. Recruitment continued until saturation, using social-network advertisements, direct contacts, and snowball sampling. The sample was limited: all practitioners were US-based, and most worked with language technologies and text data.

  3. Knowl 3 — Practitioners value synthetic data for scale and scarce resources

    empirical result

    Practitioners described synthetic data as a way to address shortages of data, time, money, and human annotation capacity. They expected auxiliary models to produce specialized or expert-like examples without recruiting scarce domain experts, and to make large-scale model evaluation feasible when human review could not keep pace with frequent testing across teams. Automated red-teaming was described as substantially cheaper than hiring cybersecurity experts to produce and test individual examples. Participants also saw synthetic data as a potential route to improved performance, competitive differentiation, greater control over dataset composition, and reduced data-rights or privacy-related compliance burdens. These were practitioners’ perceived benefits, not benefits established by the interview study; participants also reported that the promised control and quality were not consistently achieved.

  4. Knowl 4 — Generation practices trade control for flexibility

    model/method

    Practitioners most often wanted generated data to be diverse—covering varied scenarios, input types, and demographic characteristics—and to resemble natural or human-produced data without reproducing the auxiliary model’s training data. To constrain generation, teams commonly asked an auxiliary model to vary manually curated seed examples, such as by paraphrasing them, or used structured templates to steer topics, formats, or demographic attributes. More open-ended workflows asked the model to adopt a persona or scenario, accepting less predictability in exchange for potentially broader and more natural-seeming outputs; practitioners iteratively revised prompts when results missed their expectations. For model-generated evaluation scores, a common approach was to give the auxiliary model annotation instructions and a rubric similar to those used for human raters. Auxiliary-model selection was often ad hoc, based on availability, perceived state-of-the-art quality, cost, and task complexity; some practitioners preferred models they believed were less aligned when generating harmful content for safety testing.

  5. Knowl 5 — Generation is brittle and difficult to control

    limitation

    Practitioners reported that synthetic data reflects the auxiliary model’s capabilities and limitations, and that small prompt changes—such as reordering words, changing punctuation, or rephrasing—could substantially alter output quality. Teams therefore repeatedly experimented with prompts to correct inconsistent or unexpected behavior. Even targeted generation did not guarantee semantic control: outputs could follow unintended associations or be shaped by an injected example in the generation pipeline. Practitioners also doubted that synthetic data distributions would exactly match real-world distributions. Attempts to generate data about smaller or underrepresented populations could fail to represent those groups accurately and, as participants noted, might produce identifying outputs that raise privacy concerns. These reports qualify practitioners’ perception that synthetic data generation is inherently controllable.

  6. Knowl 6 — Validation remains a largely manual bottleneck

    empirical result

    Practitioners most commonly validated generated data through manual inspection, described as spot-checking or eyeballing samples for coherence and compliance with instructions. This review was often periodic and subjective, and participants said different reviewers could apply it inconsistently. Less common alternatives included consultation with domain experts, small downstream experiments testing the effect of synthetic training data on primary-model performance, and use of an auxiliary model’s own confidence scores—a potentially circular way to judge its outputs. Teams struggled to define data quality when real examples or ground truth were unavailable. Some treated diversity as the number of unique outputs, although uniqueness alone does not establish meaningful coverage. Validation was a bottleneck because data was generated at scale, difficult cases could require specialized expertise, and organizational urgency often prioritized generation over review.

  7. Knowl 7 — Chaining auxiliary models can conceal errors and shift data distributions

    limitation

    Practitioners and responsible-AI experts raised concerns about using the same or similar auxiliary models at multiple pipeline stages—for example, generating training or evaluation data and then scoring model outputs. If a model scores outputs produced by itself or a similar model, its judgments may favor those outputs; repeated reliance on related models can also create feedback loops that hide problems that a more independent evaluation might reveal. Participants additionally worried that repeated training on synthetic data could shift models away from real-world data distributions and contribute to compounding errors or model collapse. Views on the severity of this risk differed, including whether combining synthetic data with real data could prevent serious degradation. The interviews document these concerns and disagreement; they do not establish that collapse occurs in every such workflow.

  8. Knowl 8 — Identity prompting can flatten groups into stereotypes

    limitation

    Participants and responsible-AI experts warned that prompting an auxiliary model to represent a demographic identity can produce stereotyped behavior or generic responses rather than nuanced representation. One concern was that models learn about groups through material produced about them, rather than through the groups’ own perspectives, and may therefore generate caricatures. Experts further argued that outputs can default to normative assumptions embedded in the auxiliary model, reinforcing dominant cultural viewpoints and marginalizing alternatives. These risks challenge the use of generated identity-related data as a reliable substitute for data created by members of the groups it is intended to represent.

  9. Knowl 9 — Synthetic generation can weaken data subjects’ agency

    limitation

    Responsible-AI experts argued that transforming people’s activities, decisions, or creative work into synthetic data can distance those people from the resulting dataset and downstream AI systems. That abstraction may make it harder for affected people to understand, contest, or seek redress for how their data is used, removing a potential route for participation. Experts also questioned claims that synthetic data is privacy-preserving when an auxiliary model may have been trained on people’s or creators’ work: producing useful data about a group can depend on information derived from that group. Thus, synthetic data may reduce direct links to individuals while leaving unresolved questions about the source data, consent, and accountability.

  10. Knowl 10 — Responsible use requires context-specific choices and stronger processes

    model/method

    The paper proposes considerations for developing responsible synthetic-data practices, rather than settled best practices. Teams should assess whether synthetic data suits the particular purpose and domain, and determine what properties—such as diversity—must be operationalized and validated, especially in high-risk applications. They should choose auxiliary models in light of their capabilities and limitations for the task instead of assuming that the newest or most powerful model is suitable. Validation should receive adequate resources and combine methods where appropriate, such as manual review and calibration experiments; criteria should account for underrepresented and edge cases, not merely random samples. Documentation should capture the iterative process, including prompts, revisions, rejected examples, generation parameters, auxiliary-model choice, and the reasons for design decisions. Meaningful engagement with affected communities and data workers should also be considered, including who participates and at what stage. The authors note that these practices cost time and resources, creating tension with the efficiency that motivates synthetic-data use.

  11. Knowl 11 — Synthetic data may redistribute power and labor across the AI supply chain

    theoretical result

    The paper’s discussion identifies competing possible effects of synthetic data beyond the teams generating it. In principle, access to generated data could lower the cost and data-access barriers facing smaller developers. However, if those developers depend on proprietary auxiliary-model services controlled by large technology companies, dependence on APIs, usage restrictions, and compute access may instead reinforce the power of existing incumbents. Replacing human annotation may change data workers’ roles toward quality review, although the consequences for employment and labor conditions remain underexplored; automation could also reduce workers’ exposure to harmful content. For people affected by downstream models, synthetic-data distribution shifts or inadequate synthetic evaluation could leave performance and fairness failures undetected. These are implications and research questions raised by the paper, not outcomes measured by its interview study.

Coverage note — No substantial contributed material was deliberately omitted; the interview protocols and illustrative quotations were not extracted as separate knowls because they support, rather than add to, the study’s findings.

References

  1. 1.Nazmiye Ceren Abay, Yan Zhou, Murat Kantarcioglu, Bhavani Thuraisingham, and Latanya Sweeney. 2019. Privacy preserving synthetic data release using deep learning. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2018, Dublin, Ireland, September 10–14, 2018, Proceedings, Part I 18. Springer, 510–526.
  2. 2.Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. 2024. Phi-4 technical report. arXiv preprint arXiv:2412.08905 (2024).
  3. 3.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023).
  4. 4.William Agnew, A Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R McKee. 2024. The illusion of artificial inclusion. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–12.
  5. 5.Samuel A Assefa, Danial Dervovic, Mahmoud Mahfouz, Robert E Tillman, Prashant Reddy, and Manuela Veloso. 2020. Generating synthetic data in finance: opportunities, challenges and pitfalls. In Proceedings of the First ACM International Conference on AI in Finance. 1–8.
  6. 6.Pierre Azoulay, Joshua L Krieger, and Abhishek Nagaraj. 2024. Old moats for new models: Openness, control, and competition in generative ai. Technical Report. National Bureau of Economic Research.
  7. 7.Gwangbin Bae, Martin de La Gorce, Tadas Baltrušaitis, Charlie Hewitt, Dong Chen, Julien Valentin, Roberto Cipolla, and Jingjing Shen. 2023. DigiFace-1M: 1 Million Digital Face Images for Face Recognition. In 2023 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE.
  8. 8.Chelsea Barabas, Colin Doyle, JB Rubinovitz, and Karthik Dinakar. 2020. Studying up: reorienting the study of algorithmic fairness around issues of power. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. 167–176.
  9. 9.Emily M Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics 6 (2018), 587–604.
  10. 10.Virginia Braun and Victoria Clarke. 2012. Thematic analysis. American Psychological Association.
  11. 11.Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis. Qualitative research in sport, exercise and health 11, 4 (2019), 589–597.
  12. 12.Sarah Burkhardt and Bernhard Rieder. 2024. Foundation models are platform models: Prompting and the political economy of AI. Big Data & Society 11, 2 (2024), 20539517241247839.
  13. 13.Cheng-Han Chiang and Hung-yi Lee. 2023. Can large language models be an alternative to human evaluations? arXiv preprint arXiv:2305.01937 (2023).
  14. 14.Kasia Chmielinski, Sarah Newman, Chris N. Kranzinger, Michael Hind, Jennifer Wortman Vaughan, Margaret Mitchell, Julia Stoyanovich, Angelina McMillan-Major, Emily McReynolds, Kathleen Esfahany, Mary L. Gray, Audrey Chang, , and Maui Hudson. 2024. The CLeAR Documentation Framework for AI Transparency: Recommendations for Practitioners & Context for Policymakers. Harvard Kennedy School Shorenstein Center discussion paper.
  15. 15.Anamaria Crisan, Brittany Fiore-Gartland, and Melanie Tory. 2020. Passing the data baton: A retrospective analysis on data science work and workers. IEEE Transactions on Visualization and Computer Graphics 27, 2 (2020), 1860–1870.
  16. 16.Jessamyn Dahmen and Diane Cook. 2019. SynSys: A synthetic data generation system for healthcare applications. Sensors 19, 5 (2019), 1181.
  17. 17.Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2023. The participatory turn in ai design: Theoretical foundations and the current state of practice. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. 1–23.
  18. 18.Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh, Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu. 2022. Exploring how machine learning practitioners (try to) use fairness toolkits. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 473–484.
  19. 19.Wesley Hanwen Deng, Nur Yildirim, Monica Chang, Motahhare Eslami, Kenneth Holstein, and Michael Madaio. 2023. Investigating Practices and Opportunities for Cross-functional Collaboration around AI Fairness in Industry Practice. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 705–716.
  20. 20.Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. 2024. Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems 36 (2024).
  21. 21.Madeleine Clare Elish and Danah Boyd. 2018. Situating methods in the magic of Big Data and AI. Communication monographs 85, 1 (2018), 57–80.
  22. 22.Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics 9 (2021), 391–409.
  23. 23.Shangbin Feng, Vidhisha Balachandran, Yuyang Bai, and Yulia Tsvetkov. 2023. Factkb: Generalizable factuality evaluation using language models enhanced with factual knowledge. arXiv preprint arXiv:2305.08281 (2023).
  24. 24.Andrew Fitzgerald. 2024. Why Synthetic Data Can Never Be Ethical: A Lesson from Media Ethics. Surveillance & Society 22, 4 (Dec. 2024), 477–482. https://doi.org/10.24908/ss.v22i4.18324
  25. 25.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM 64, 12 (December 2021), 86–92.
  26. 26.Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences 120, 30 (2023), e2305016120.
  27. 27.Lisa Gitelman. 2013. “Raw Data” Is an Oxymoron.
  28. 28.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020. Generative adversarial networks. Commun. ACM 63, 11 (2020), 139–144.
  29. 29.Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dacheng Tao. 2021. Knowledge Distillation: A Survey. International Journal of Computer Vision 129 (2021), 1789–1819.
  30. 30.Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024).
  31. 31.Igor Grossmann, Matthew Feinberg, Dawn C Parker, Nicholas A Christakis, Philip E Tetlock, and William A Cunningham. 2023. AI and the transformation of social science research. Science 380, 6650 (2023), 1108–1109.
  32. 32.Tom Gunter, Zirui Wang, Chong Wang, Ruoming Pang, Andy Narayanan, Aonan Zhang, Bowen Zhang, Chen Chen, Chung-Cheng Chiu, David Qiu, et al. 2024. Apple intelligence foundation language models. arXiv preprint arXiv:2407.21075 (2024).
  33. 33.Alon Halevy, Peter Norvig, and Fernando Pereira. 2009. The unreasonable effectiveness of data. IEEE intelligent systems 24, 2 (2009), 8–12.
  34. 34.Lei Han, Tianwa Chen, Gianluca Demartini, Marta Indulska, and Shazia Sadiq. 2023. A data-driven analysis of behaviors in data curation processes. ACM Transactions on Information Systems 41, 3 (2023), 1–35.
  35. 35.Alex Hanna and Tina M Park. 2020. Against scale: Provocations and resistances to scale thinking. arXiv preprint arXiv:2010.08850 (2020).
  36. 36.Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. http://arxiv.org/abs/2203.09509 arXiv:2203.09509 [cs].
  37. 37.Amy K Heger, Liz B Marquis, Mihaela Vorvoreanu, Hanna Wallach, and Jennifer Wortman Vaughan. 2022. Understanding machine learning practitioners’ data documentation perceptions, needs, challenges, and desiderata. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2 (2022), 1–29.
  38. 38.Paula Helm, Benjamin Lipp, and Roser Pujadas. 2024. Generating reality and silencing debate: Synthetic data as discursive device. Big Data & Society 11, 2 (June 2024), 20539517241249447. https://doi.org/10.1177/20539517241249447 Publisher: SAGE Publications Ltd.
  39. 39.Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015. Distilling the Knowledge in a Neural Network. arXiv preprint arXiv:1503.02531.
  40. 40.Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. 2019. Improving fairness in machine learning systems: What do industry practitioners need?. In Proceedings of the 2019 CHI conference on human factors in computing systems. 1–16.
  41. 41.Qisheng Hu, Kaixin Li, Xu Zhao, Yuxi Xie, Tiedong Liu, Hui Chen, Qizhe Xie, and Junxian He. 2023. Instructcoder: Empowering language models for code editing. arXiv preprint arXiv:2310.20329 (2023).
  42. 42.Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al. 2022. Inner monologue: Embodied reasoning through planning with language models. arXiv preprint arXiv:2207.05608 (2022).
  43. 43.IBM. 2023. What is Synthetic Data? | IBM. https://www.ibm.com/think/topics/synthetic-data [Online; accessed 2025-01-21].
  44. 44.Apple Inc. 2024. Introducing Apple’s On-Device and Server Foundation Models - Apple Machine Learning Research. https://machinelearning.apple.com/research/introducing-apple-foundation-models [Online; accessed 2025-01-21].
  45. 45.Benjamin N Jacobsen. 2023. Machine learning and the politics of synthetic data. Big Data & Society 10, 1 (Jan. 2023), 20539517221145372. https://doi.org/10.1177/20539517221145372 Publisher: SAGE Publications Ltd.
  46. 46.Nikita Jaipuria, Xianling Zhang, Rohan Bhasin, Mayar Arafa, Punarjay Chakravarty, Shubham Shrivastava, Sagar Manglani, and Vidya N Murali. 2020. Deflating dataset bias using synthetic data augmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. 772–773.
  47. 47.Harry H Jiang, Lauren Brown, Jessica Cheng, Mehtab Khan, Abhishek Gupta, Deja Workman, Alex Hanna, Johnathan Flowers, and Timnit Gebru. 2023. AI Art and its Impact on Artists. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. 363–374.
  48. 48.James Jordon, Lukasz Szpruch, Florimond Houssiau, Mirko Bottarelli, Giovanni Cherubin, Carsten Maple, Samuel N Cohen, and Adrian Weller. 2022. Synthetic Data–what, why and how? arXiv preprint arXiv:2205.03257 (2022).
  49. 49.James Jordon, Jinsung Yoon, and Mihaela Van Der Schaar. 2018. PATE-GAN: Generating synthetic data with differential privacy guarantees. In International conference on learning representations.
  50. 50.Shivani Kapania, Alex S Taylor, and Ding Wang. 2023. A hunt for the Snark: Annotator Diversity in Data Practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Machinery, New York, NY, USA, 1–15. https://doi.org/10.1145/3544548.3580645
  51. 51.Diederik P Kingma. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013).
  52. 52.Laura Koesten, Elena Simperl, Tom Blount, Emilia Kacprzak, and Jeni Tennison. 2020. Everything you always wanted to know about a dataset: Studies in data summarisation. International Journal of Human-Computer Studies 135 (March 2020), 102367. https://doi.org/10.1016/j.ijhcs.2019.10.004
  53. 53.Adam Kortylewski, Bernhard Egger, Andreas Schneider, Thomas Gerig, Andreas Morel-Forster, and Thomas Vetter. 2019. Analyzing and Reducing the Damage of Dataset Bias to Face Recognition With Synthetic Data. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 2261–2268. https://doi.org/10.1109/CVPRW.2019.00279
  54. 54.Francis Lee, Saghi Hajisharif, and Ericka Johnson. 2025. The ontological politics of synthetic data: Normalities, outliers, and intersectional hallucinations. Big Data & Society 12, 2 (2025), 20539517251318289.
  55. 55.Peter Lee. 2024. Synthetic Data and the Future of AI. (2024).
  56. 56.Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. 2023. Code as policies: Language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 9493–9500.
  57. 57.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. Advances in neural information processing systems 36 (2024).
  58. 58.Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zhang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, and Andrew M. Dai. 2024. Best Practices and Lessons Learned on Synthetic Data for Language Models. https://doi.org/10.48550/arXiv.2404.07503 arXiv:2404.07503 [cs].
  59. 59.Dieuwertje Luitse and Wiebke Denkena. 2021. The great transformer: Examining the role of large language models in the political economy of AI. Big Data & Society 8, 2 (2021), 20539517211047734.
  60. 60.Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. arXiv preprint arXiv:2308.09583 (2023).
  61. 61.Michael Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan, and Hanna Wallach. 2022. Assessing the Fairness of AI Systems: AI Practitioners’ Processes, Challenges, and Needs for Support. Proceedings of the ACM on Human-Computer Interaction 6, CSCW1 (2022), 26 pages.
  62. 62.Michael A Madaio, Luke Stark, Jennifer Wortman Vaughan, and Hanna Wallach. 2020. Co-designing checklists to understand organizational challenges and opportunities around fairness in ai. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. 1–14.
  63. 63.Daniella Meeker, Crystal Kallem, Yan Heras, Stephanie Garcia, and Casey Thompson. 2022. Case report: evaluation of an open-source synthetic data platform for simulation studies. JAMIA open 5, 3 (2022), ooac067.
  64. 64.Cade Metz, Cecilia Kang, Sheera Frenkel, Stuart A. Thompson, and Nico Grant. 2024. How Tech Giants Cut Corners to Harvest Data for A.I. - The New York Times. https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html [Online; accessed 2025-01-21].
  65. 65.Milagros Miceli, Martin Schuessler, and Tianling Yang. 2020. Between subjectivity and imposition: Power dynamics in data annotation for computer vision. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2 (2020), 1–25.
  66. 66.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency. 220–229.
  67. 67.Michael Muller, Ingrid Lange, Dakuo Wang, David Piorkowski, Jason Tsay, Q Vera Liao, Casey Dugan, and Thomas Erickson. 2019. How data science workers work with data: Discovery, capture, curation, design, creation. In Proceedings of the 2019 CHI conference on human factors in computing systems. 1–15.
  68. 68.Laura Nader. 1972. Up the anthropologist: Perspectives gained from studying up. (1972).
  69. 69.Sergey I Nikolenko. 2021. Synthetic data for deep learning. Vol. 174. Springer.
  70. 70.Code of Practice Working Groups. 2024. Second Draft of the General-Purpose AI Code of Practice. https://digital-strategy.ec.europa.eu/en/library/second-draft-general-purpose-ai-code-practice-published-written-independent-experts [Online; accessed 2025-01-20].
  71. 71.Office of the Assistant Secretary for Planning and Evaluation (ASPE). 2022. A Synthetic Health Data Generation Engine to Accelerate Patient-Centered Outcomes Research | ASPE. https://aspe.hhs.gov/synthetic-health-data-generation-engine-accelerate-patient-centered-outcomes-research [Online; accessed 2025-01-22].
  72. 72.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 (2022), 27730–27744.
  73. 73.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311–318.
  74. 74.Samir Passi and Steven Jackson. 2017. Data Vision: Learning to See Through Algorithmic Abstraction. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing (CSCW ’17). Association for Computing Machinery, New York, NY, USA, 2436–2447. https://doi.org/10.1145/2998181.2998331
  75. 75.Samir Passi and Steven J Jackson. 2018. Trust in data science: Collaboration, translation, and accountability in corporate data science projects. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–28.
  76. 76.Personal Data Protection Commission (PDPC). 2024. Privacy Enhancing Technology (PET): Proposed Guide On Synthetic Data Generation. https://www.pdpc.gov.sg/-/media/files/pdpc/pdf-files/other-guides/proposed-guide-on-synthetic-data-generation.pdf [Online; accessed 2025-01-20].
  77. 77.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red Teaming Language Models with Language Models. arXiv:2202.03286 [cs] (Feb. 2022). http://arxiv.org/abs/2202.03286 arXiv: 2202.03286.
  78. 78.Crystal Qian, Michael Xieyang Liu, Emily Reif, Grady Simon, Nada Hussein, Nathan Clement, James Wexler, Carrie J Cai, Michael Terry, and Minsuk Kahng. 2024. The Evolution of LLM Adoption in Industry Data Curation Practices. arXiv preprint arXiv:2412.16089 (2024).
  79. 79.Bogdana Rakova, Jingying Yang, Henriette Cramer, and Rumman Chowdhury. 2021. Where responsible AI meets reality: Practitioner perspectives on enablers for shifting organizational practices. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1 (2021), 1–23.
  80. 80.Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot, Djamé Seddah, and Jacopo Staiano. 2020. Synthetic data augmentation for zero-shot cross-lingual question answering. arXiv preprint arXiv:2010.12643 (2020).
  81. 81.Kevin Roose. 2024. Data for A.I. Training Is Disappearing Fast, Study Shows - The New York Times. https://www.nytimes.com/2024/07/19/technology/ai-data-restrictions.html [Online; accessed 2025-01-21].
  82. 82.Roy A Ruddle, James Cheshire, and Sara Johansson Fernstad. 2023. Tasks and visualizations used for data profiling: A survey and interview study. IEEE Transactions on Visualization and Computer Graphics (2023).
  83. 83.Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. In proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–15.
  84. 84.Siamak Shakeri, Noah Constant, Mihir Sanjay Kale, and Linting Xue. 2020. Towards zero-shot multilingual synthetic question and answer generation for cross-lingual reading comprehension. arXiv preprint arXiv:2010.12008 (2020).
  85. 85.Hoo-Chang Shin, Neil A Tenenholtz, Jameson K Rogers, Christopher G Schwarz, Matthew L Senjem, Jeffrey L Gunter, Katherine P Andriole, and Mark Michalski. 2018. Medical image synthesis for data augmentation and anonymization using generative adversarial networks. In Simulation and Synthesis in Medical Imaging: Third International Workshop, SASHIMI 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 16, 2018, Proceedings 3. Springer, 1–11.
  86. 86.Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. 2024. AI models collapse when trained on recursively generated data. Nature 631 (2024), 755–759.
  87. 87.Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The psychological well-being of content moderators: the emotional labor of commercial moderation and avenues for improving support. In Proceedings of the 2021 CHI conference on human factors in computing systems. 1–14.
  88. 88.Daniel Susser, Daniel S Schiff, Sara Gerke, Laura Y Cabrera, I Glenn Cohen, Megan Doerr, Jordan Harrod, Kristin Kostick-Quenet, Jasmine McNealy, Michelle N Meyer, et al. 2024. Synthetic Health Data: Real Ethical Promise and Peril. Hastings Center Report 54, 5 (2024), 8–13.
  89. 89.Daniel Susser, Daniel S. Schiff, Sara Gerke, Laura Y. Cabrera, I. Glenn Cohen, Megan Doerr, Jordan Harrod, Kristin Kostick-Quenet, Jasmine McNealy, Michelle N. Meyer, W. Nicholson Price II, and Jennifer K. Wagner. 2024. Synthetic Health Data: Real Ethical Promise and Peril. Hastings Center Report 54, 5 (2024), 8–13. https://doi.org/10.1002/hast.4911 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/hast.4911.
  90. 90.Daniel Susser and Jeremy Seeman. [n. d.]. Dialogue Critical Provocations for Synthetic Data. ([n. d.]).
  91. 91.Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023).
  92. 92.Risto Uuk, Annemieke Brouwer, Tim Schreier, Noemi Dreksler, Valeria Pulignano, and Rishi Bommasani. 2024. Effective Mitigations for Systemic Risks from General-Purpose AI. arXiv preprint arXiv:2412.02145 (2024).
  93. 93.Janet Vertesi and Paul Dourish. 2011. The value of data: considering the context of production in data economies. In Proceedings of the ACM 2011 conference on Computer supported cooperative work. 533–542.
  94. 94.Angelina Wang, Jamie Morgenstern, and John P. Dickerson. 2024. Large language models should not replace human participants because they can misportray and flatten identity groups. arXiv preprint arXiv:2402.01908.
  95. 95.Ding Wang, Shantanu Prabhat, and Nithya Sambasivan. 2022. Whose AI Dream? In search of the aspiration in data annotation.. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–16.
  96. 96.Cedric Deslandes Whitney and Justin Norman. 2024. Real Risks of Fake Data: Synthetic Data, Diversity-Washing and Consent Circumvention. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM, Rio de Janeiro Brazil, 1733–1744. https://doi.org/10.1145/3630106.3659002
  97. 97.David Gray Widder and Dawn Nafus. 2023. Dislocated accountabilities in the “AI supply chain”: Modularity and developers’ notions of responsibility. Big Data & Society 10, 1 (2023), 20539517231177620.
  98. 98.David Gray Widder, Sarah West, and Meredith Whittaker. 2023. Open (for business): Big tech, concentrated power, and the political economy of open AI. Concentrated Power, and the Political Economy of Open AI (August 17, 2023) (2023).
  99. 99.David Gray Widder and Richmond Wong. 2023. Thinking upstream: Ethics and policy opportunities in AI supply chains. arXiv preprint arXiv:2303.07529 (2023).
  100. 100.David Gray Widder, Derrick Zhen, Laura Dabbish, and James Herbsleb. 2023. It’s about power: What ethical concerns do software engineers have, and what do they (feel they can) do about them?. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 467–479.
  101. 101.Tanja Wiehn. 2024. Synthetic Data: From Data Scarcity to Data Pollution. Surveillance & Society 22, 4 (Dec. 2024), 472–476. https://doi.org/10.24908/ss.v22i4.18327
  102. 102.Matteo Wong. 2024. The GPT Era Is Already Ending - The Atlantic. https://www.theatlantic.com/technology/archive/2024/12/openai-o1-reasoning-models/680906/?utm_source=chatgpt.com [Online; accessed 2025-01-22].
  103. 103.Kanit Wongsuphasawat, Yang Liu, and Jeffrey Heer. 2019. Goals, process, and challenges of exploratory data analysis: An interview study. arXiv preprint arXiv:1911.00568 (2019).
  104. 104.Sierra Wyllie, Ilia Shumailov, and Nicolas Papernot. 2024. Fairness feedback loops: training on synthetic data amplifies bias. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2113–2147.
  105. 105.Liyang Xie, Kaixiang Lin, Shu Wang, Fei Wang, and Jiayu Zhou. 2018. Differentially private generative adversarial network. arXiv preprint arXiv:1802.06739 (2018).
  106. 106.Zhenpei Yang, Yuning Chai, Dragomir Anguelov, Yin Zhou, Pei Sun, Dumitru Erhan, Sean Rafferty, and Henrik Kretzschmar. 2020. Surfelgan: Synthesizing realistic sensor data for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11118–11127.
  107. 107.Nur Yildirim, Mahima Pushkarna, Nitesh Goyal, Martin Wattenberg, and Fernanda Viégas. 2023. Investigating How Practitioners Use Human-AI Guidelines: A Case Study on the People+ AI Guidebook. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–13.
  108. 108.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284 (2023).
  109. 109.Amy X Zhang, Michael Muller, and Dakuo Wang. 2020. How do data science workers collaborate? Roles, workflows, and tools. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1 (2020), 1–23.
  110. 110.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2023), 46595–46623.

Citation

MLA
Kapania, S., et al. “Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline”. arXiv, 2025, http://arxiv.org/abs/2501.18493v2.
APA
Kapania, S., Ballard, S., Kessler, A., & Vaughan, J. W. (2025). Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline. arXiv. http://arxiv.org/abs/2501.18493v2
Chicago
Kapania, S., S. Ballard, A. Kessler, and J. W. Vaughan. 2025. “Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline”. arXiv. http://arxiv.org/abs/2501.18493v2.
Harvard
Kapania, S. et al. (2025) “Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.18493v2.
Vancouver
1. Kapania S, Ballard S, Kessler A, Vaughan JW (2025) Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline. arXiv

BibTeX

@article{kapania2025examining,
  title = {Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline},
  author = {Kapania, Shivani and Ballard, Stephanie and Kessler, Alex and Vaughan, Jennifer Wortman},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.18493v2},
  eprint = {2501.18493}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/