keyword
instruction data selection
Instruction data selection refers to the process of identifying, filtering, and curating a high-quality, diverse subset of prompt-response pairs from larger training datasets for the purpose of instruction tuning machine learning models, particularly large language models. Rather than training models on massive, uncurated collections of instructional examples, this technique evaluates candidate samples using criteria such as response quality, task complexity, semantic diversity, and model-specific learning difficulty. By prioritizing informative and challenging samples over sheer quantity, instruction data selection aims to improve the model's ability to follow human instructions accurately while substantially reducing computational costs, training time, and the risks associated with redundant or noisy data.
2 items

A Survey on Data Selection for LLM Instruction Tuning
Bolin Zhang, Jiahao Wang, Qianlong Du, Jiajun Zhang, Zhiying Tu, Dianhui Chu
Why you should read this
Presents a structured taxonomy and comparative analysis of data selection strategies for large language model instruction tuning, providing clear guidance on how to filter high-impact training samples to cut computational costs while boosting model performance.
Instruction tuning is a vital step of training large language models (LLMs), so how to enhance the effect of instruction tuning has received increased attention. Existing works indicate that the quality of the dataset is more crucial than the quantity during instruction tuning of LLMs. Therefore, recently a lot of studies focus on exploring the methods of selecting high-quality subset from instruction datasets, aiming to reduce training costs and enhance the instruction-following capabilities of LLMs. This paper presents a comprehensive survey on data selection for LLM instruction tuning. Firstly, we introduce the wildly used instruction datasets. Then, we propose a new taxonomy of the data selection methods and provide a detailed introduction of recent advances, and the evaluation strategies and results of data selection methods are also elaborated in detail. Finally, we emphasize the open challenges and present new frontiers of this task.
Added
2026-10-02

From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, Jing Xiao
Why you should read this
Proposes the Instruction-Following Difficulty metric to let large language models autonomously select high-impact training samples, outperforming full-dataset fine-tuning on Alpaca and WizardLM while using only ten percent of the data.
In the realm of Large Language Models (LLMs), the balance between instruction data quality and quantity is a focal point. Recognizing this, we introduce a self-guided methodology for LLMs to autonomously discern and select cherry samples from open-source datasets, effectively minimizing manual curation and potential cost for instruction tuning an LLM. Our key innovation, the Instruction-Following Difficulty (IFD) metric, emerges as a pivotal metric to identify discrepancies between a model's expected responses and its intrinsic generation capability. Through the application of IFD, cherry samples can be pinpointed, leading to a marked uptick in model training efficiency. Empirical validations on datasets like Alpaca and WizardLM underpin our findings; with a mere 10% of original data input, our strategy showcases improved results. This synthesis of self-guided cherry-picking and the IFD metric signifies a transformative leap in the instruction tuning of LLMs, promising both efficiency and resource-conscious advancements. Codes, data, and models are available.
Added
2026-09-28
