Data Curation Alone Can Stabilize In-context Learning
Ting-Yun ChangRobin Jia
Demonstrates that selecting high-value training subsets through individual example scoring significantly reduces performance variance and increases accuracy in in-context learning without requiring dynamic prompt retrieval or model calibration.
In-context learning allows large language models to tackle new classification tasks simply by receiving a few reference examples inside the prompt, without modifying model parameters. While this approach avoids costly retraining, standard in-context learning is notoriously fragile. Minor adjustments to the selected prompt examples, their order, or prompt formatting frequently cause substantial fluctuations in accuracy. Organizations deploying language models therefore face significant performance variability and operational risk.
The article demonstrates that carefully curating a small subset of high-quality training examples can resolve this instability without requiring complex inference-time retrieval systems, prompt calibration, or model fine-tuning. The primary objective is to evaluate whether automated data valuation techniques can extract small pools of consistently effective examples from larger training datasets to stabilize and improve few-shot performance.
To identify stable subsets, the authors evaluated two scoring methods on large language models including GPT-J (6 billion parameters) and OPT (13 billion parameters) across five classification benchmarks. The first method, conditional accuracy, scores each training example by its average validation accuracy when combined with other random examples. The second method, data modeling, fits linear regression models to predict language model output margins based on the presence and position of specific training examples. The highest-scoring examples per class are selected to form a curated subset of 20 examples. Prompts sampled from these curated pools were tested against baselines using uncurated datasets, calibration, one-shot scoring, and top-performing prompt unions.
The results show that data curation alone substantially boosts reliability and accuracy. Across all tasks and models, prompts randomly drawn from subsets curated via conditional accuracy and data modeling improved average accuracy by 7.7% and 6.3%, respectively, over sampling from the full, uncurated training set. These curated subsets raised worst-case performance floors and reduced variance across prompt combinations. When evaluated on out-of-distribution target datasets, curated prompts maintained superior performance, demonstrating that the selected examples teach task-level definitions rather than overfitting to specific distributions. Even in an unlabeled setup—where inputs were paired with all candidate labels—the curated selection method improved performance by 5.7% over the fully gold-labeled baseline. Surprisingly, analyses showed that stable examples are not defined by high text diversity or low perplexity; instead, effective examples cluster tightly in representation space.
These findings indicate that in-context learning sensitivity stems largely from low-quality data rather than an inherent flaw in few-shot prompting. For practitioners, establishing a fixed pool of curated examples removes the operational complexity, latency, and infrastructure costs associated with dynamic example-retrieval pipelines. Furthermore, the ability to curate high-performing prompts without gold labels lowers data annotation costs. Curated prompts also proved more compact, with 4 curated examples outperforming 16 to 24 uncurated examples in several tasks, conserving prompt context limits.
Organizations leveraging in-context learning should prioritize offline data curation to build vetted example pools for deployment prompts. If resources permit, teams should evaluate conditional accuracy scoring on a representative validation set before deployment. However, decision-makers should account for significant upfront computational costs: profiling example effectiveness required running tens of thousands of validation prompts, consuming hundreds of GPU hours in larger configurations. Additional exploration is recommended to test more efficient search methods during data collection and to determine if these findings transfer to generative tasks and models exceeding 100 billion parameters.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Its experiments show which parts of demonstrations matter—and that gold labels may matter little—providing key context for this paper’s data-quality and unlabeled-setting results.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). Its analysis of selecting useful in-context examples establishes the example-quality problem that this paper addresses with offline data valuation rather than query-time retrieval.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). Its contextual calibration method is an important response to few-shot instability and clarifies the calibration baseline against which this paper evaluates curation.
- Paper: LESS: Selecting Influential Data for Targeted Instruction Tuning, Mengzhou Xia et al. (2024). LESS carries the small-subset, data-influence perspective into targeted instruction tuning, extending the curation question beyond in-context examples to model training.
- Paper: In-Context Learning with Long-Context Models: An In-Depth Exploration, Amanda Bertsch et al. (2025). Its long-context experiments revisit example selection and prompt robustness at thousands-of-demonstrations scale, extending the source’s few-shot curation findings to a different context regime.
