A Survey on Data Selection for LLM Instruction Tuning

Bolin ZhangJiahao WangQianlong DuJiajun ZhangZhiying TuDianhui Chu

article2025JAIR105 citations

Presents a structured taxonomy and comparative analysis of data selection strategies for large language model instruction tuning, providing clear guidance on how to filter high-impact training samples to cut computational costs while boosting model performance.

Listen

Training modern artificial intelligence models to reliably follow user instructions traditionally requires fine-tuning them on massive datasets containing tens or hundreds of thousands of examples. However, relying purely on large data volumes introduces severe problems, including high computational and electricity costs, data redundancy, and inconsistent quality. Manual curation by humans produces high-quality data but is too expensive and slow to scale, and it introduces subjective human bias. Consequently, researchers have shifted focus toward automated data selection—identifying small, highly effective subsets of data that allow models to achieve strong performance at a fraction of the training cost.

The main objective of the article is to provide a comprehensive survey and taxonomy of automated data selection methods for instruction tuning. It evaluates how different selection strategies function, compares their effectiveness across standard industry benchmarks, and identifies open operational and research challenges.

The article conducts a systematic literature review and comparative analysis across four distinct categories of selection methods: indicator-based systems that score data using mathematical metrics (such as length and perplexity); trainable model approaches where dedicated language models learn to assess instruction difficulty; methods driven by powerful external models (such as GPT-4) using specialized prompts; and small-model approaches that use compact networks to cluster data and ensure diversity. The authors assess these methodologies across three standard evaluation schemes: win rates against reference models, internal comparisons against models trained on full datasets, and external comparisons against independent state-of-the-art benchmarks.

The key findings demonstrate that data quality decisively outweighs data quantity in instruction tuning. Fine-tuning models on carefully selected subsets of only 5% to 20% of an original dataset consistently matches or outperforms models trained on the entire unfiltered dataset, whereas random downsampling degrades performance. In particular, methods like Instruction Following Difficulty (IFD) and Model-Oriented Data Selection (MoDS) substantially surpass models trained on full datasets. Second, simply increasing the subset size does not guarantee better model capabilities, as redundant or low-quality data can dilute performance gains. Third, foundational model capability matters: more advanced base models extract significantly greater learning value from identical high-quality subsets than older architectures. Finally, no single selection methodology is universally superior across all operational needs; for instance, indicator methods offer high processing speed and transparency but low domain adaptability, whereas advanced model filters deliver the highest selection quality but suffer from severe cost and latency bottlenecks.

These findings have immediate practical implications for engineering timelines, budgets, and operational risk. Organizations fine-tuning language models can drastically lower computing expenses, energy consumption, and infrastructure costs by prioritizing rigorous data selection over massive data acquisition. Furthermore, training on smaller, curated subsets accelerates deployment cycles and reduces the risk of models learning errors or toxic patterns from low-quality data. However, over-reliance on proprietary commercial application programming interfaces (APIs) for data filtering introduces vendor dependence, recurring expenses, and latency risks.

To maximize efficiency and performance, technical leaders should adopt hybrid data selection workflows that combine fast, transparent heuristic filters for initial data pruning with targeted model-based scoring for final curation. Teams should also explore lightweight or distilled selector models to reduce dependence on expensive external APIs. Further development is necessary before standardized enterprise deployment, specifically establishing universal evaluation benchmarks and expanding selection frameworks beyond English to specialized domains such as law and medicine.

While the findings are strongly supported by cross-benchmark comparisons, readers should interpret current results with moderate caution. The field currently lacks standardized, uniform evaluation protocols, meaning that data selection methods are tested across varying model architectures and subjective judging criteria. Additionally, existing research is heavily concentrated on general English-language tasks, and confidence in cross-domain or multilingual performance remains limited until more standardized benchmarks are established.

No sufficiently relevant recommendations were found.

Cover for A Survey on Data Selection for LLM Instruction Tuning

Abstract

Instruction tuning is a vital step of training large language models (LLMs), so how to enhance the effect of instruction tuning has received increased attention. Existing works indicate that the quality of the dataset is more crucial than the quantity during instruction tuning of LLMs. Therefore, recently a lot of studies focus on exploring the methods of selecting high-quality subset from instruction datasets, aiming to reduce training costs and enhance the instruction-following capabilities of LLMs. This paper presents a comprehensive survey on data selection for LLM instruction tuning. Firstly, we introduce the wildly used instruction datasets. Then, we propose a new taxonomy of the data selection methods and provide a detailed introduction of recent advances, and the evaluation strategies and results of data selection methods are also elaborated in detail. Finally, we emphasize the open challenges and present new frontiers of this task.

Table of Contents

  • 1 Introduction
  • 2 Instruction Datasets
  • 3 Data Selection Methods
  • 3.1 Methods Based on a System of Indicators
  • 3.2 Methods Based on Trainable LLMs
  • 3.3 Methods Based on Powerful LLMs like ChatGPT
  • 3.4 Methods Based on Small Models
  • 3.5 Comparative Analysis of Data Selection Methods
  • 4 Evaluation methods and Result Analysis
  • 4.1 Wining Rate
  • 4.2 Inner Comparison
  • 4.3 External Comparison
  • 4.4 Result Analysis
  • 5 Discussion and Open Challenges
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Four paradigms organize instruction-data selection and expose different trade-offs

    model/method

    Instruction-data selection methods are grouped by the scoring mechanism and model resources they use: (1) indicator-based methods combine predefined or learned measures of sample quality; (2) trainable-LLM methods adapt a language model or selector to score instructions; (3) powerful-LLM methods use an external model such as GPT-4 through prompts or API calls; and (4) small-model methods use compact pretrained models or modular pipelines. The groups have different practical strengths. Indicator methods are generally scalable and interpretable but have low domain adaptability and variable selection effectiveness. Trainable-LLM methods are characterized by high domain adaptability and effectiveness, but low scalability and interpretability. Powerful-LLM methods also tend toward high selection effectiveness, but depend on costly, opaque external APIs and have low scalability. Small-model approaches offer moderate scalability and interpretability, low-to-medium domain adaptability, and medium effectiveness. No category dominates across all dimensions, so the choice involves balancing efficiency, adaptability, interpretability, and selection quality.

  2. Knowl 2 — Selection quality is defined by downstream performance at a fixed subset size

    definition

    Let X={x1,…,xn}X=\{x_1,\ldots,x_n\} be an instruction dataset of nn examples, and let π\pi be a selection method that returns a subset of size mm, where 0≤m≤n0\leq m\leq n. Write this subset as Sπ(m)=π(X)S_\pi^{(m)}=\pi(X). A predefined evaluation metric QQ measures the quality of the selected subset, typically through downstream model performance after fine-tuning. The ideal selector at size mm is therefore defined as

    π∗=arg⁡max⁡πQ ⁣(Sπ(m)).\pi^*=\arg\max_{\pi} Q\!\left(S_\pi^{(m)}\right).

    This formulation makes selection effectiveness conditional on both the chosen evaluation metric and the subset size; it does not define data quality independently of downstream use.

  3. Knowl 3 — Trainable-LLM selectors learn instruction scores, with IFD measuring instruction dependence

    model/method

    Trainable-LLM approaches use a language model adapted to score instruction examples rather than relying only on hand-designed features. In the Instruction Following Difficulty (IFD) method, a language model is first fine-tuned on a small, clustered subset to acquire basic instruction-following ability. For instruction QQ and target answer A=(w1,…,wN)A=(w_1,\ldots,w_N), where wiw_i is the iith answer token, NN is the number of answer tokens, and θ\theta denotes the model parameters, the conditioned answer score is

    sθ(A∣Q)=1N∑i=1Nlog⁡Pθ(wi∣Q,w1,…,wi−1).s_\theta(A\mid Q)=\frac{1}{N}\sum_{i=1}^{N}\log P_\theta(w_i\mid Q,w_1,\ldots,w_{i-1}).

    The unconditioned score sθ(A)s_\theta(A) is calculated analogously without the instruction QQ. IFD defines the instruction-dependence score as

    rθ(Q,A)=sθ(A∣Q)sθ(A).r_\theta(Q,A)=\frac{s_\theta(A\mid Q)}{s_\theta(A)}.

    IFD interprets higher scores as cases where the instruction has greater influence on generating the answer, and applies a threshold to select examples for further fine-tuning. Other trainable-LLM methods use different signals: Instruction Backtranslation generates instructions from web documents and filters them with a self-scoring model; Nuggets compares zero-shot and one-shot performance on predefined tasks to form a “gold score”; DIVERSEEVOL repeatedly expands a subset using instruction embeddings and k-center-greedy selection; TEGIT trains a task generator and task discriminator from a text-grounded meta-dataset; and Active Instruction Tuning prioritizes tasks with high prompt uncertainty, estimated from output-probability changes under randomly deleted words.

  4. Knowl 4 — Powerful external LLMs score, label, or diversify candidate instructions

    model/method

    Methods based on powerful external LLMs use prompt-driven scoring or analysis to select instruction data. AlpaGasus prompts ChatGPT to score each instruction–input–response tuple and discards examples below a quality threshold. INSTAG uses ChatGPT to produce and normalize open-ended instruction labels, then constructs a subset by sorting instructions by label count and adding examples that contribute distinctive labels until the desired size is reached. DEITA scores complexity and response quality, multiplies those scores into a combined score, and then uses embedding distances to add sufficiently diverse examples until the target subset size is met. LIFT first expands the instruction distribution and selects by row variance, then uses ChatGPT ratings of accuracy, interpretability, clarity, difficulty, and length to curate the result. InstructionNode measures complexity by the number of nodes in GPT-4-generated semantic parse trees, and can increase complexity by adding nodes before converting the tree back into an instruction. WaveCoder uses a GPT-4-based discriminator to assess generated code instructions against criteria organized into subtopics. These mechanisms trade access to strong scoring or analysis capabilities for dependence on external models and their associated API costs and latency.

  5. Knowl 5 — Indicator-based selection combines explicit quality measures and diversity objectives

    model/method

    Indicator-based methods assign each instruction multiple scores and combine them into an overall selection score. For an instruction xx, let I1(x),…,IK(x)I_1(x),\ldots,I_K(x) be KK indicators, such as length, perplexity, reward score, or a learned feature, and let GG be their aggregation function. The overall score is G(I1(x),…,IK(x))G(I_1(x),\ldots,I_K(x)), and thresholding this score yields a selected subset. INSTRUCTMINING fits a linear scoring rule from natural-language indicators using fine-tuning experiments: candidate subsets are evaluated on a test set, their performance supplies quality labels, and least squares estimates the rule's parameters. InstructionGPT-4 combines visual and text metrics, including CLIP-based features and instruction length, in a trainable selector whose quality labels are derived from fine-tuning and evaluating clustered subsets. DQ instead emphasizes dataset diversity. For candidate example xx, its gain is

    P(x)=∑p∈S∥f(p)−f(x)∥22−∑p∈D∖S∥f(p)−f(x)∥22,P(x)=\sum_{p\in S}\|f(p)-f(x)\|_2^2-\sum_{p\in D\setminus S}\|f(p)-f(x)\|_2^2,

    where DD is the full dataset, SS is the current selected subset, pp ranges over examples in those sets, and ff maps an example to its feature representation. DQ iteratively partitions the dataset into non-overlapping subsets using this gain and uniformly selects a representative from each partition, aiming to preserve dataset-wide diversity.

  6. Knowl 6 — Small-model pipelines approximate selection with quality, coverage, and necessity signals

    model/method

    Small-model approaches seek useful selection without relying on a powerful external API or training a large selector. MoDS combines three criteria: quality, coverage, and necessity. It first uses a reward model to retain examples above a quality threshold, then uses k-center-greedy selection to choose diverse seed instructions. A pretrained LLM is fine-tuned on the seeds; the resulting model is applied to the quality-filtered pool, and a reward model identifies examples with low scores as potentially important for further learning. The seed examples and selected augmented instructions form the final subset. A coreset-based alternative embeds examples with a pretrained language model such as BERT, identifies a center using unsupervised clustering, and applies KCenterGreedy to select representative core samples. These methods aim to reduce training-data volume while maintaining or potentially improving model performance.

  7. Knowl 7 — Common instruction datasets vary in source, scale, and curation

    data/table

    The surveyed instruction datasets illustrate contrasting construction strategies and scales. Self-Instruct contains 52,000 training instructions and 252 test instructions; it uses seed tasks and an LLM to generate inputs and outputs, followed by post-processing for uniqueness and relevance. Alpaca contains 52,002 instruction–response pairs generated with the Self-Instruct framework and text-davinci-003. WizardLM contains 250,000 pairs generated with ChatGPT using depth- and breadth-evolution procedures to increase instruction complexity and scope. LIMA has 1,000 training, 300 test, and 50 development examples, including human-authored examples and carefully selected Q&A material. Dolly-V2 contains 15,000 human-authored pairs covering tasks such as brainstorming, classification, question answering, and summarization; its creators were restricted to Wikipedia as a source and were not to use generative AI to write responses. P3 combines 170 NLP datasets with 2,052 handcrafted prompts that convert conventional tasks into natural-language inputs and outputs. Together, these datasets range from model-generated collections to manually curated examples and templated NLP tasks, so their construction choices imply different quality-control and diversity considerations.

  8. Knowl 8 — Selection is evaluated by judged win rate and same-model or cross-model comparisons

    definition

    The survey distinguishes three evaluation paradigms. Win rate compares a model fine-tuned on a selected subset (LLM-sub) with a reference model, often one fine-tuned on the full dataset or on a regular-sampling subset. It is defined as

    WinRate=Num(win)−Num(lose)Num(all)+1,\mathrm{WinRate}=\frac{\mathrm{Num(win)}-\mathrm{Num(lose)}}{\mathrm{Num(all)}}+1,

    where the terms count wins by LLM-sub, losses by LLM-sub, and all benchmark cases, respectively. A score of 11 indicates parity, a value above 11 indicates an advantage, and a value below 11 indicates worse performance. Outputs are commonly rated by a judge such as GPT-4 on a 1–10 scale; presenting the two outputs in both orders can reduce positional bias. Inner comparison compares LLM-sub with the same model fine-tuned on the full dataset or a same-size regularly sampled subset. External comparison compares LLM-sub with a different LLM, testing transfer across models or architectures. The three paradigms answer different questions and are not interchangeable.

  9. Knowl 9 — Reported results show subset selection can beat random sampling but does not guarantee gains

    empirical result

    Reported aggregate win scores show advantages for several selected subsets over their reference models, while also showing that results depend on the selector, subset size, and base model. With IFD selection from Alpaca, Llama-7B trained on 5%, 10%, and 15% of the data had aggregate win scores of 1.04, 1.097, and 1.064, respectively, versus the full-data Llama-7B reference; the 10% subset therefore scored higher than either the 5% or 15% subset. For Llama2-7B trained on 5% of Alpaca with IFD, the reported score was 1.4311. MoDS with 2,000 Alpaca examples and Llama2-7B reported 1.4786, while random sampling with 5% of Alpaca and Llama-7B reported 0.9. These aggregate scores use 1 as parity under the survey's win-score convention.

    Other comparisons also show mixed outcomes across tasks. On Platypus, LIFT's selected 15,000 examples scored 0.844/0.643/0.49/0.645 on HellaSwag/ARC/TruthfulQA/MMLU, compared with 0.82/0.607/0.438/0.625 for 15,000 randomly selected examples. In an external comparison on MT-bench and AlpacaEval, DEITA's mixture-trained Llama2-13B scored 6.79 and 81.09, respectively; TAGLM-13B-v1.0 scored 6.44 ± 0.04 and 72.8. The survey also reports that increasing subset size does not necessarily improve performance: IFD's Llama-7B Alpaca results were not monotonic across 5%, 10%, and 15%, whereas the reported AlpaGasus results improved with subset size in that setting. Overall, selection can reduce data requirements and outperform regular selection, but effectiveness varies by task and setup.

  10. Knowl 10 — Evaluation consistency, selection cost, and domain coverage remain open problems

    limitation

    The survey identifies three unresolved limitations in instruction-data selection. First, methods use different benchmarks and evaluation procedures, making results difficult to compare fairly; a standardized benchmark suite should cover diverse tasks, model architectures, and languages, and should assess task-conditioned or real-world usefulness rather than relying only on coarse overall scores. Second, scoring hundreds of thousands of examples with large LLMs or external APIs can be slow and expensive; proposed directions include compact selectors distilled from powerful models and retrieval-augmented pruning that applies expensive scoring only to highly ranked candidates. Third, most methods focus on English and general-purpose tasks, leaving multilingual and specialized domains such as biomedical and legal instruction underexplored; multilingual encoders, domain-specific priors, and modular components are suggested as ways to improve transfer. The survey also calls for reproducible, open selection pipelines, including code, selector models, scores, and evaluation benchmarks.

Coverage note — No substantial contribution is omitted; the survey's individual method descriptions are summarized by selection family, emphasizing their selection signals and distinguishing mechanisms rather than reproducing every implementation detail.

References

  1. 1.Y. Cao, Y. Kang, and L. Sun. 2023. Instruction mining: high-quality instruction data selection for large language models. CoRR, abs/2307.06290.
  2. 2.H. Chen, Y. Zhang, Q. Zhang, H. Yang, X. Hu, X. Ma, Y. Yanggong, and J. Zhao. 2023. Maybe only 0.5% data is needed: A preliminary exploration of low training data instruction tuning. CoRR, abs/2305.09246.
  3. 3.L. Chen et al. 2024. Alpagasus: training a better alpaca with fewer data. In The Twelfth International Conference on Learning Representations.
  4. 4.Y. Chen, H. Jiang, X. Huang, S. Shi, and G. Qi. 2023. Tegit: generating high-quality instruction-tuning data with text-grounded task design. CoRR, abs/2309.05447.
  5. 5.A. Chowdhery et al. 2023. Palm: scaling language modeling with pathways. Journal of Machine Learning Research, 24, 240, 1–113.
  6. 6.M. Conover et al. 2023. Free dolly: introducing the world’s first truly open instruction-tuned llm. (2023).
  7. 7.W. Dong, M. Charikar, and K. Li. 2011. Efficient k-nearest neighbor graph construction for generic similarity measures. In Proceedings of the 20th International Conference on World Wide Web, WWW 2011, Hyderabad, India, March 28 - April 1, 2011. S. Srinivasan, K. Ramamritham, A. Kumar, M. P. Ravindra, E. Bertino, and R. Kumar, (Eds.) ACM, 577–586.
  8. 8.Q. Du, C. Zong, and J. Zhang. 2023. Mods: model-oriented data selection for instruction tuning. CoRR, abs/2311.15653.
  9. 9.S. Ghosh, C. K. R. Evuru, S. Kumar, S. Ramaneswaran, D. Aneja, Z. Jin, R. Duraiswami, and D. Manocha. 2024. A closer look at the limitations of instruction tuning. In International Conference on Machine Learning. PMLR, 15559–15589.
  10. 10.P. Kung, F. Yin, D. Wu, K. Chang, and N. Peng. 2023. Active instruction tuning: improving cross-task generalization by training on prompt sensitive tasks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023. H. Bouamor, J. Pino, and K. Bali, (Eds.) Association for Computational Linguistics, 1813–1829.
  11. 11.M. Li, Y. Zhang, Z. Li, J. Chen, L. Chen, N. Cheng, J. Wang, T. Zhou, and J. Xiao. 2024. From quantity to quality: boosting llm performance with self-guided data selection for instruction tuning. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7595–7628.
  12. 12.X. Li, P. Yu, C. Zhou, T. Schick, O. Levy, L. Zettlemoyer, J. E. Weston, and M. Lewis. 2024. Self-alignment with instruction backtranslation. In The Twelfth International Conference on Learning Representations.
  13. 13.Y. Li et al. 2024. One-shot learning as instruction data prospector for large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4586–4601.
  14. 14.W. Liu, W. Zeng, K. He, Y. Jiang, and J. He. 2024. What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning. In The Twelfth International Conference on Learning Representations.
  15. 15.S. Longpre et al. 2023. The flan collection: designing data and methods for effective instruction tuning. In International Conference on Machine Learning. PMLR, 22631–22648.
  16. 16.K. Lu, H. Yuan, Z. Yuan, R. Lin, J. Lin, C. Tan, C. Zhou, and J. Zhou. 2024. # instag: instruction tagging for analyzing supervised fine-tuning of large language models. In The Twelfth International Conference on Learning Representations.
  17. 17.OpenAI. 2023. Gpt-4 technical report. CoRR, abs/2303.08774.
  18. 18.L. Ouyang et al. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  19. 19.L. Ouyang et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35, 27730–27744.
  20. 20.A. Radford et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine Learning Research). M. Meila and T. Zhang, (Eds.) Vol. 139. PMLR, 8748–8763.
  21. 21.N. Reimers and I. Gurevych. 2019. Sentence-bert: sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3982–3992.
  22. 22.V. Sanh et al. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  23. 23.O. Sener and S. Savarese. 2018. Active learning for convolutional neural networks: A core-set approach. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  24. 24.O. Sener and S. Savarese. 2018. Active learning for convolutional neural networks: A core-set approach. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  25. 25.R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto. 2023. Stanford alpaca: an instruction-following llm model. https://github.com/tatsu-lab/stanford_alpaca. (2023).
  26. 26.H. Touvron et al. 2023. Llama: open and efficient foundation language models. CoRR, abs/2302.13971.
  27. 27.Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi. 2023. Self-instruct: aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023. A. Rogers, J. L. Boyd-Graber, and N. Okazaki, (Eds.) Association for Computational Linguistics, 13484–13508.
  28. 28.J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le. 2021. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  29. 29.L. Wei, Z. Jiang, W. Huang, and L. Sun. 2023. Instructiongpt-4: A 200-instruction paradigm for fine-tuning minigpt-4. CoRR, abs/2308.12067.
  30. 30.S. Wu, K. Lu, B. Xu, J. Lin, Q. Su, and C. Zhou. 2023. Self-evolved diverse data sampling for efficient instruction tuning. CoRR, abs/2311.08182.
  31. 31.C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, and D. Jiang. 2023. Wizardlm: empowering large language models to follow complex instructions. CoRR, abs/2304.12244.
  32. 32.Y. Xu, Y. Yao, Y. Huang, M. Qi, M. Wang, B. Gu, and N. Sundaresan. 2023. Rethinking the instruction quality: LIFT is what you need. CoRR, abs/2312.11508.
  33. 33.Z. Yu, X. Zhang, N. Shang, Y. Huang, C. Xu, Y. Zhao, W. Hu, and Q. Yin. 2023. Wavecoder: widespread and versatile enhanced instruction tuning with refined data generation. CoRR, abs/2312.14187. arXiv: 2312.14187. doi: 10.48550/ARXIV.2312.14187.
  34. 34.S. Zhang et al. 2023. Instruction tuning for large language models: A survey. CoRR, abs/2308.10792.
  35. 35.Y. Zhao, B. Yu, B. Hui, H. Yu, F. Huang, Y. Li, and N. L. Zhang. 2023. A preliminary study of the intrinsic relationship between complexity and alignment. (2023). arXiv: 2308.05696 [cs.CL].
  36. 36.C. Zhou et al. 2023. LIMA: less is more for alignment. In Thirty-seventh Conference on Neural Information Processing Systems.
  37. 37.D. Zhou, K. Wang, J. Gu, X. Peng, D. Lian, Y. Zhang, Y. You, and J. Feng. 2023. Dataset quantization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 17205–17216.
  38. 38.D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny. 2024. Minigpt-4: enhancing vision-language understanding with advanced large language models. In The Twelfth International Conference on Learning Representations.

Citation

MLA
Zhang, B., et al. “A Survey on Data Selection for LLM Instruction Tuning”. Journal of Artificial Intelligence Research, vol. 83, 2025, https://doi.org/10.1613/jair.1.17625.
APA
Zhang, B., Wang, J., Du, Q., Zhang, J., Tu, Z., & Chu, D. (2025). A Survey on Data Selection for LLM Instruction Tuning. Journal of Artificial Intelligence Research, 83. https://doi.org/10.1613/jair.1.17625
Chicago
Zhang, B., J. Wang, Q. Du, J. Zhang, Z. Tu, and D. Chu. 2025. “A Survey on Data Selection for LLM Instruction Tuning”. Journal of Artificial Intelligence Research 83. https://doi.org/10.1613/jair.1.17625.
Harvard
Zhang, B. et al. (2025) “A Survey on Data Selection for LLM Instruction Tuning”, Journal of Artificial Intelligence Research, 83. Available at: https://doi.org/10.1613/jair.1.17625.
Vancouver
1. Zhang B, Wang J, Du Q, Zhang J, Tu Z, Chu D (2025) A Survey on Data Selection for LLM Instruction Tuning. Journal of Artificial Intelligence Research. https://doi.org/10.1613/jair.1.17625

BibTeX

@article{Zhang_2025, title={A Survey on Data Selection for LLM Instruction Tuning}, volume={83}, ISSN={1076-9757}, url={http://dx.doi.org/10.1613/jair.1.17625}, DOI={10.1613/jair.1.17625}, journal={Journal of Artificial Intelligence Research}, publisher={AI Access Foundation}, author={Zhang, Bolin and Wang, Jiahao and Du, Qianlong and Zhang, Jiajun and Tu, Zhiying and Chu, Dianhui}, year={2025}, month=Aug }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/