A Survey on Data Selection for LLM Instruction Tuning
Bolin ZhangJiahao WangQianlong DuJiajun ZhangZhiying TuDianhui Chu
Presents a structured taxonomy and comparative analysis of data selection strategies for large language model instruction tuning, providing clear guidance on how to filter high-impact training samples to cut computational costs while boosting model performance.
Training modern artificial intelligence models to reliably follow user instructions traditionally requires fine-tuning them on massive datasets containing tens or hundreds of thousands of examples. However, relying purely on large data volumes introduces severe problems, including high computational and electricity costs, data redundancy, and inconsistent quality. Manual curation by humans produces high-quality data but is too expensive and slow to scale, and it introduces subjective human bias. Consequently, researchers have shifted focus toward automated data selection—identifying small, highly effective subsets of data that allow models to achieve strong performance at a fraction of the training cost.
The main objective of the article is to provide a comprehensive survey and taxonomy of automated data selection methods for instruction tuning. It evaluates how different selection strategies function, compares their effectiveness across standard industry benchmarks, and identifies open operational and research challenges.
The article conducts a systematic literature review and comparative analysis across four distinct categories of selection methods: indicator-based systems that score data using mathematical metrics (such as length and perplexity); trainable model approaches where dedicated language models learn to assess instruction difficulty; methods driven by powerful external models (such as GPT-4) using specialized prompts; and small-model approaches that use compact networks to cluster data and ensure diversity. The authors assess these methodologies across three standard evaluation schemes: win rates against reference models, internal comparisons against models trained on full datasets, and external comparisons against independent state-of-the-art benchmarks.
The key findings demonstrate that data quality decisively outweighs data quantity in instruction tuning. Fine-tuning models on carefully selected subsets of only 5% to 20% of an original dataset consistently matches or outperforms models trained on the entire unfiltered dataset, whereas random downsampling degrades performance. In particular, methods like Instruction Following Difficulty (IFD) and Model-Oriented Data Selection (MoDS) substantially surpass models trained on full datasets. Second, simply increasing the subset size does not guarantee better model capabilities, as redundant or low-quality data can dilute performance gains. Third, foundational model capability matters: more advanced base models extract significantly greater learning value from identical high-quality subsets than older architectures. Finally, no single selection methodology is universally superior across all operational needs; for instance, indicator methods offer high processing speed and transparency but low domain adaptability, whereas advanced model filters deliver the highest selection quality but suffer from severe cost and latency bottlenecks.
These findings have immediate practical implications for engineering timelines, budgets, and operational risk. Organizations fine-tuning language models can drastically lower computing expenses, energy consumption, and infrastructure costs by prioritizing rigorous data selection over massive data acquisition. Furthermore, training on smaller, curated subsets accelerates deployment cycles and reduces the risk of models learning errors or toxic patterns from low-quality data. However, over-reliance on proprietary commercial application programming interfaces (APIs) for data filtering introduces vendor dependence, recurring expenses, and latency risks.
To maximize efficiency and performance, technical leaders should adopt hybrid data selection workflows that combine fast, transparent heuristic filters for initial data pruning with targeted model-based scoring for final curation. Teams should also explore lightweight or distilled selector models to reduce dependence on expensive external APIs. Further development is necessary before standardized enterprise deployment, specifically establishing universal evaluation benchmarks and expanding selection frameworks beyond English to specialized domains such as law and medicine.
While the findings are strongly supported by cross-benchmark comparisons, readers should interpret current results with moderate caution. The field currently lacks standardized, uniform evaluation protocols, meaning that data selection methods are tested across varying model architectures and subjective judging criteria. Additionally, existing research is heavily concentrated on general English-language tasks, and confidence in cross-domain or multilingual performance remains limited until more standardized benchmarks are established.
- Paper: From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning, Ming Li et al. (2024). Its Instruction-Following Difficulty method is a central example of self-guided selection that the survey examines and contextualizes.
- Paper: LESS: Selecting Influential Data for Targeted Instruction Tuning, Mengzhou Xia et al. (2024). LESS provides a key gradient-influence approach to instruction-data selection, clarifying one of the survey’s major method families.
- Paper: Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning, Ming Li et al. (2024). Superfiltering grounds the survey’s discussion of lightweight selectors by showing how a small model can filter instruction data for larger models.
- Paper: Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning, Hao Zhao et al. (2024). This work supplies a simple length-based selection baseline that helps explain the survey’s indicator-driven methods and comparisons.
- Paper: LIMA: Less Is More for Alignment, Chunting Zhou et al. (2023). LIMA’s finding that a small, carefully curated instruction set can align a model establishes important context for the survey’s data-efficiency question.
- Paper: WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions, Can Xu et al. (2023). WizardLM’s Evol-Instruct pipeline and resulting dataset provide a concrete source of instruction data whose selection the survey’s methods address.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). Self-Instruct’s generation and heuristic filtering pipeline introduces the synthetic instruction-data setting that motivates later automated selection.
No sufficiently relevant recommendations were found.
