Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline
Shivani KapaniaStephanie BallardAlex KesslerJennifer Wortman Vaughan
Reveals how industry practitioners use synthetic data throughout AI development pipelines, identifying critical challenges in demographic representation, output control, and validation to provide practical guidance for responsible use.
As generative artificial intelligence development expands, engineering teams face significant bottlenecks in acquiring large-scale, high-quality human datasets due to high costs, strict privacy regulations, and overall data scarcity. To overcome these constraints, organizations are increasingly turning to synthetic data—machine-generated content and automated evaluation scores produced by large "auxiliary" generative models. However, organizational practices, standards, and policy frameworks have struggled to keep pace with this rapid shift, leaving major uncertainties regarding the reliability, safety, and governance of these systems.
The article aims to empirically examine why and how practitioners integrate synthetic data across modern AI development pipelines, identify the core challenges they face during generation and validation, and evaluate the resulting ethical and sociotechnical risks.
To conduct this evaluation, the authors performed a qualitative study involving 29 participants across 14 United States-based organizations between May and August 2024. The research unfolded in two phases: the first consisted of semi-structured interviews with 19 active AI practitioners focusing on workflows and technical challenges, while the second involved 10 responsible AI experts who analyzed the broader societal, governance, and safety implications using real-world use-case vignettes.
The findings reveal that auxiliary models are now deeply embedded across every phase of AI development, serving not only to generate training and test sets but also to simulate interactive user sessions and score model outputs (acting as an automated judge). Practitioners report that synthetic workflows provide massive efficiency gains; for example, manual red-teaming often requires skilled experts spending up to 45 minutes to craft a single test sample, whereas automated auxiliary models generate test cases nearly instantly at minimal inference cost. However, generating controlled data remains problematic because auxiliary models are brittle, highly sensitive to minor prompt changes, and prone to creating inaccurate caricatures or stereotypes rather than authentic representations of underrepresented populations. Crucially, validation practices remain a major bottleneck: despite acknowledging data quality risks, most teams rely on superficial "spot-checking" or manual "eyeballing" because market pressures prioritize rapid scaling over rigorous verification. Furthermore, practitioners frequently "chain" models—using the same auxiliary architecture to both generate and evaluate outputs—which creates systemic feedback loops and increases the risk of long-term model degradation.
These findings indicate that while synthetic data drastically lowers development costs and accelerates release timelines, it introduces hidden systemic risks. Over-reliance on synthetic evaluations can provide false confidence regarding system safety, while removing real human subjects from datasets eliminates avenues for individuals to contest data usage or exercise agency. Additionally, using foundation models to generate data concentrates market power among a few large cloud and model providers, reinforcing cultural dominance and narrowing output diversity.
To mitigate these risks, the article recommends establishing domain-specific standards for synthetic data use, differentiating between high-risk applications (such as medical tools) that require strict data operationalization and low-risk environments where noise is tolerable. Organizations must shift away from ad-hoc spot-checking toward standardized, multi-method validation protocols and implement formal documentation detailing prompts, parameters, model choices, and rejected iterations. Regulators and industry leaders should also enforce transparency disclosures for synthetic datasets and incorporate expert or community consultation frameworks when modeling affected demographic groups.
Readers should interpret these insights within the context of the study's scope: the participant sample was restricted to United States-based professionals working primarily on text-based systems. While the evidence offers a strong, confident qualitative picture of current industry workflows, quantitative validation metrics and concrete best-practice frameworks across multimodal and global settings remain emerging areas that warrant ongoing experimentation.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). Self-Instruct makes synthetic data generation for model training concrete, showing how model-generated instructions and examples can scale beyond costly human-created datasets.
- Paper: Data Feedback Loops: Model-driven Amplification of Dataset Biases, Rohan Taori et al. (2023). This analysis of model-driven data feedback loops supplies a key risk framework for understanding how synthetic data can propagate and amplify biases across AI development.
- Paper: Is GPT-3 a Good Data Annotator?, Bosheng Ding et al. (2023). Its comparison of LLM-generated labels and synthetic training examples with human-annotated data provides an early practical case of auxiliary models in the development pipeline.
- Paper: Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions, John Joon Young Chung et al. (2023). This study clarifies why synthetic text generation requires balancing diversity against correctness and human review, issues central to the source's account of validation challenges.
- Paper: Modeling Tabular data using Conditional GAN, Lei Xu et al. (2019). CTGAN establishes a concrete approach to generating synthetic datasets for model training, offering useful context for the source's broader survey of synthetic-data practices.
- Paper: SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations, Shuaiqi Wang et al. (2026). SynAE turns the source's concern about manually validating synthetic evaluation data into a quantitative framework for assessing validity, fidelity, diversity, and downstream agent performance.
