From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Ming LiYong ZhangZhitao LiJiuhai ChenLichang ChenNing ChengJianzong WangTianyi ZhouJing Xiao

article2024NAACL408 citations

Proposes the Instruction-Following Difficulty metric to let large language models autonomously select high-impact training samples, outperforming full-dataset fine-tuning on Alpaca and WizardLM while using only ten percent of the data.

Listen

Training large language models to follow human instructions typically relies on fine-tuning them on massive datasets. However, curating and training on these large volumes is computationally expensive and introduces low-quality or redundant data. While recent research shows that small sets of high-quality data can outperform large, indiscriminate datasets, identifying these high-value samples automatically has remained an unsolved operational challenge that often requires costly manual curation or dependence on external models.

The article introduces and evaluates an automated, self-guided data selection method that enables language models to identify and select the most effective training samples—termed cherry data—from large open-source datasets without relying on external teacher models.

The authors developed a three-stage framework based on a novel metric called Instruction-Following Difficulty. First, a base language model undergoes brief training on a diverse subset of approximately 300 to 1,000 samples to establish basic instruction-following behavior. Second, this model evaluates target dataset samples by comparing the loss when generating a response with the instruction versus generating it in isolation. The resulting difficulty score identifies samples where following the prompt requires substantial alignment effort, while filtering out misaligned pairs. Finally, the model is retrained exclusively on the top-scoring difficulty samples. The approach was tested across standard benchmarks and five human-curated test suites using standard open-source datasets including Alpaca (52,000 samples) and WizardLM (around 64,000 samples), evaluated by automated judges and human evaluators.

The analysis reveals several key findings. First, models trained on only 5% of Alpaca data and 10% of WizardLM data surpassed the performance of baseline models trained on 100% of the original data across standard benchmarks and head-to-head win rates. Second, selecting data using the difficulty metric consistently outperformed alternative pruning methods, including random sampling, diversity-only clustering, and standard loss-based filtering. Third, exposure to a minimal set of roughly 300 diverse samples during the initial phase was sufficient to maximize data-selection performance. Linguistic analysis also showed that the highest-value data concentrated heavily in creative, multi-step tasks requiring deep reasoning, whereas low-scoring samples consisted primarily of simple, rule-based text edits.

These findings demonstrate that training efficiency can be dramatically improved by prioritizing instruction complexity over dataset volume. For organizations developing specialized models, this approach offers substantial reductions in compute costs, training timelines, and human data curation expenses, while simultaneously improving model accuracy. The method also shifts data strategy from indiscriminate data gathering toward targeted selection of high-complexity instructions.

Organizations should consider adopting the Instruction-Following Difficulty metric to prune existing fine-tuning corpora down to the top 10% most difficult samples before initiating full-scale training. However, decision-makers must weigh specific performance trade-offs. The article observed that pruning reduced accuracy in specialized domains such as mathematics and computer programming, which require extensive domain-specific data volume for 7-billion parameter base models. For production deployments, teams can choose between training a brief initial model for maximum selection accuracy or calculating difficulty scores directly from modern base models to maximize operational throughput.

Confidence in these findings is high for general instruction following across diverse open-domain tasks, supported by consistent outcomes across multiple model architectures and evaluation protocols. Nevertheless, stakeholders should note that the approach exhibits limitations on highly technical sub-categories and that determining the exact optimal data percentage remains dependent on the underlying dataset distribution.

Cover for From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Abstract

In the realm of Large Language Models (LLMs), the balance between instruction data quality and quantity is a focal point. Recognizing this, we introduce a self-guided methodology for LLMs to autonomously discern and select cherry samples from open-source datasets, effectively minimizing manual curation and potential cost for instruction tuning an LLM. Our key innovation, the Instruction-Following Difficulty (IFD) metric, emerges as a pivotal metric to identify discrepancies between a model's expected responses and its intrinsic generation capability. Through the application of IFD, cherry samples can be pinpointed, leading to a marked uptick in model training efficiency. Empirical validations on datasets like Alpaca and WizardLM underpin our findings; with a mere 10% of original data input, our strategy showcases improved results. This synthesis of self-guided cherry-picking and the IFD metric signifies a transformative leap in the instruction tuning of LLMs, promising both efficiency and resource-conscious advancements. Codes, data, and models are available.

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 Learning from Brief Experience
  • 2.2 Evaluating Based on Experience
  • 3 Experimental Setup
  • 3.1 Datasets
  • 3.2 Implementation Details
  • 3.3 Evaluation Metrics
  • 3.3.1 Pair-wise Comparison
  • 3.3.2 Benchmarks
  • 3.3.3 Human Evaluation
  • 4 Experimental Results
  • 4.1 Main Results
  • 4.2 Ablation on Data Selection Mechanism
  • 4.2.1 Data Randomly Selected
  • 4.2.2 Data with Diversity
  • 4.2.3 Data with Low IFD Score
  • 4.2.4 Data with High CA Scores
  • 4.3 Ablation on Pre-Experienced Data
  • 4.3.1 Number of Pre-Experience Data
  • 4.3.2 Distribution of Pre-Experience Data
  • 4.4 Results on LLaMA2 Models
  • 5 Cherry Data Characteristics
  • 5.1 Distribution Characteristics
  • 5.2 Pattern Characteristics
  • 6 Related Work
  • 6.1 Data-driven Instruction Tuning
  • 6.2 Coreset Selection
  • 6.3 Instruction Data Selection
  • 6.4 Pointwise Mutual Information
  • 7 Conclusion
  • Limitation
  • Acknowledgement
  • References
  • A Implementation Details
  • B Ablation on Data Selection Mechanism
  • C Performance across Sub-Categories
  • D Results with Official WizardLM
  • E Cherry Data General Characteristics
  • F Cherry Example Analysis
  • G Additional Discussion
  • G.1 Fully-trained Model as Pre-Experienced Model?
  • G.2 How Many Cherry Samples are Required?
  • H Prompt for Evaluation
  • I Detailed Main Comparison
  • I.1 Comparison with the Official Alpaca
  • I.2 Comparison with the Reimplemented WizardLM
  • I.3 Comparison with the Official WizardLM
  • J Detailed Ablation Comparison
  • J.1 Data Randomly Selected
  • J.2 Data with Low IFD Score
  • J.3 Data with High CA Scores
  • J.4 Number of Pre-Experienced Data
  • J.5 Distribution of Pre-Experience Data
  • J.6 Fully-trained Model as Pre-Experienced Models
  • K More Examples

Knowls

  1. Knowl 1 — Self-Guided Data Selection Pipeline for Instruction Tuning

    model/method

    The self-guided instruction data selection methodology identifies high-impact training samples (termed "cherry data") from large open-source datasets without relying on external teacher models. The process operates in three distinct phases:

    1. Learning from Brief Experience: A pre-trained base model LLMθ0\text{LLM}_{\theta_0} is trained for a single epoch on a small, diverse subset of instruction triplets (e.g., 1,000 samples selected via clustering instruction representations). This creates a pre-experienced model LLMθ\text{LLM}_{\theta} equipped with a rudimentary instruction-following capability.

    2. Evaluating Based on Experience: The pre-experienced model scores every candidate sample (Q,A)(Q, A) in the target dataset using the Instruction-Following Difficulty (IFD) metric. Pairs with IFD>1.0\text{IFD} > 1.0 are filtered out as misaligned samples where the instruction does not assist response generation.

    3. Retraining from Self-Guided Experience: Samples with the highest valid IFD scores (typically the top 5% to 10%) are extracted as the cherry data. The original raw pre-trained base model LLMθ0\text{LLM}_{\theta_0} is then fine-tuned from scratch exclusively on this subset to create the final "cherry model".

  2. Knowl 2 — Instruction-Following Difficulty (IFD) Metric

    equation

    Given an instruction-response pair (Q,A)(Q, A), where QQ is the instruction prompt and A=(w1A,w2A,…,wNA)A = (w_1^A, w_2^A, \dots, w_N^A) is the ground-truth answer of length NN tokens, the Instruction-Following Difficulty (IFD) metric evaluates the difficulty of following instruction QQ by isolating the instructional component from the intrinsic complexity of generating AA.

    The Conditioned Answer Score sθ(A∣Q)s_\theta(A \mid Q) is the average cross-entropy loss of generating the answer given the instruction context:

    sθ(A∣Q)=−1N∑i=1Nlog⁡P(wiA∣Q,w1A,…,wi−1A;θ)s_\theta(A \mid Q) = -\frac{1}{N} \sum_{i=1}^N \log P(w_i^A \mid Q, w_1^A, \dots, w_{i-1}^A; \theta)

    The Direct Answer Score sθ(A)s_\theta(A) is the model's intrinsic average cross-entropy loss when predicting the answer tokens autoregressively without the instruction context:

    sθ(A)=−1N∑i=1Nlog⁡P(wiA∣w1A,…,wi−1A;θ)s_\theta(A) = -\frac{1}{N} \sum_{i=1}^N \log P(w_i^A \mid w_1^A, \dots, w_{i-1}^A; \theta)

    The Instruction-Following Difficulty (IFD) score is defined as the ratio:

    IFDθ(Q,A)=sθ(A∣Q)sθ(A)\text{IFD}_\theta(Q, A) = \frac{s_\theta(A \mid Q)}{s_\theta(A)}

    where θ\theta represents the parameters of the evaluating model (such as a pre-experienced model). A higher IFD score indicates that the instruction provides less relative help in predicting AA, indicating a more challenging instruction-following task. When IFDθ(Q,A)>1.0\text{IFD}_\theta(Q, A) > 1.0, the Conditioned Answer Score exceeds the Direct Answer Score, signaling semantic misalignment between instruction QQ and response AA.

  3. Knowl 3 — Instruction Embedding and Pre-Experience Clustering Algorithm

    algorithm

    To build the pre-experienced model, a diverse set of pre-experienced training samples is extracted from an initial dataset by clustering sentence embeddings derived from the base pre-trained language model.

    Input: Target dataset D0={(Instructionj,Inputj,Answerj)}j=1nD_0 = \{(Instruction_j, Input_j, Answer_j)\}_{j=1}^n, pre-trained base model LLMθ0\text{LLM}_{\theta_0}, cluster count K=100K = 100, samples per cluster M=10M = 10
    Output: Brief pre-experience dataset DbriefD_{\text{brief}} of size K×M=1000K \times M = 1000
    for each sample j∈{1,…,n}j \in \{1, \dots, n\} do
        Qj←map(Instructionj,Inputj)Q_j \leftarrow \text{map}(Instruction_j, Input_j) with token sequence (wj,1Q,…,wj,mQ)(w_{j,1}^Q, \dots, w_{j,m}^Q)
        [hj,1Q,…,hj,mQ]←LLMθ0(wj,1Q,…,wj,mQ)[h_{j,1}^Q, \dots, h_{j,m}^Q] \leftarrow \text{LLM}_{\theta_0}(w_{j,1}^Q, \dots, w_{j,m}^Q)
        hjQ←1m∑i=1mhj,iQh_j^Q \leftarrow \frac{1}{m} \sum_{i=1}^m h_{j,i}^Q
    end for
    Clusters {C1,…,CK}←K-Means({hjQ}j=1n,K)\{C_1, \dots, C_K\} \leftarrow \text{K-Means}(\{h_j^Q\}_{j=1}^n, K)
    Dbrief←∅D_{\text{brief}} \leftarrow \emptyset
    for each cluster Ck∈{C1,…,CK}C_k \in \{C_1, \dots, C_K\} do
        Sk←UniformlySample(Ck,min⁡(M,∣Ck∣))S_k \leftarrow \text{UniformlySample}(C_k, \min(M, |C_k|))
        Dbrief←Dbrief∪SkD_{\text{brief}} \leftarrow D_{\text{brief}} \cup S_k
    end for
    return DbriefD_{\text{brief}}

    Here hj,iQh_{j,i}^Q represents the last hidden state of token wj,iQw_{j,i}^Q from LLMθ0\text{LLM}_{\theta_0}, and hjQh_j^Q is the mean-pooled instruction embedding. The base model LLMθ0\text{LLM}_{\theta_0} is trained on DbriefD_{\text{brief}} for 1 epoch to produce the pre-experienced model.

  4. Knowl 4 — Data Efficiency of Cherry Models on Alpaca and WizardLM Datasets

    empirical result

    Selecting instruction data via high IFD scores enables language models trained on small data subsets to outperform models trained on full datasets:

    1. Alpaca Dataset (52,002 samples): A LLaMA-7B model trained on the top 5% (~2,600 samples) of IFD-selected cherry data outperforms the official Alpaca model trained on 100% of the dataset. Across five benchmark test sets (Vicuna, Koala, WizardLM, Self-instruct, and LIMA) judged by GPT-4 pairwise evaluation with position-bias correction, the 5% cherry model achieves a winning score greater than 1.0. In human evaluation across 100 sampled test instructions, the 5% cherry model achieves 49 wins, 25 ties, and 26 losses against the 100% Alpaca baseline.

    2. WizardLM Dataset (63,655 filtered samples): A LLaMA-7B model trained on the top 10% (~6,365 samples) of IFD-selected data outperforms the reimplemented WizardLM model trained on 100% data. In human evaluation across 100 test instructions, the 10% cherry model achieves 37 wins, 32 ties, and 31 losses against the 100% WizardLM model.

  5. Knowl 5 — Open LLM and AlpacaEval Benchmark Performance Across LLaMA and LLaMA2 Models

    data/table

    Across LLaMA-7B, LLaMA2-7B, and LLaMA2-13B models, cherry models trained on only 5% to 15% of instruction data match or surpass baselines trained on 100% of the data on the Huggingface Open LLM Leaderboard (ARC, HellaSwag, MMLU, TruthfulQA) and AlpacaEval:

    Model Avg ARC HellaSwag MMLU TruthfulQA AlpacaEval
    Official Alpaca 50.21 42.65 76.91 41.73 39.55 26.46
    Cherry Alpaca (5%) 52.06 53.92 79.49 36.51 38.33 34.74
    Reimplemented WizardLM 52.79 53.07 77.44 37.75 42.90 61.99
    Cherry WizardLM (10%) 51.59 52.90 78.95 33.08 41.41 61.44
    Alpaca LLaMA2-7B (100%) 55.25 54.35 78.65 47.02 40.98 27.75
    LLaMA2-7B Cherry (5%) 55.78 57.94 80.37 44.19 40.62 36.78
    LLaMA2-7B Cherry (10%) 56.31 58.02 80.42 46.64 40.18 -
    LLaMA2-7B Cherry (15%) 56.37 57.42 80.68 46.40 40.95 -
    Alpaca LLaMA2-13B (100%) 58.78 57.59 81.98 54.05 41.49 35.00
    LLaMA2-13B Cherry (5%) 61.21 62.37 84.00 55.65 42.82 46.82
    LLaMA2-13B Cherry (10%) 61.02 62.97 83.88 55.29 41.93 -
    LLaMA2-13B Cherry (15%) 61.23 62.37 83.48 55.56 43.42 -

    For LLaMA-7B, 5% cherry data raises Open LLM average performance from 50.21 to 52.06 and AlpacaEval win rate from 26.46% to 34.74%. For LLaMA2-13B, 5% cherry data raises Open LLM average from 58.78 to 61.21 and AlpacaEval from 35.00% to 46.82%.

  6. Knowl 6 — Ablation of Instruction Data Selection Strategies Against IFD

    empirical result

    Comparing different 5% data selection heuristics on the Alpaca dataset demonstrates the superiority of the IFD metric over random selection, diversity-only clustering, and perplexity-based selection:

    Selection Strategy (5%) Avg ARC HellaSwag MMLU TruthfulQA Human Win/Tie/Loss vs. Cherry 5% Winning Score
    Cherry 5% (High IFD) 52.06 53.92 79.49 36.51 38.33 - -
    Random 5% 50.61 53.52 79.33 32.90 36.67 19 / 23 / 58 1.39
    Diversity-only 5% 49.48 53.41 79.29 29.19 36.04 18 / 21 / 61 1.43
    Low IFD 5% 50.77 53.92 79.09 34.83 35.25 5 / 8 / 87 1.82
    High CA Score 5% 47.51 51.45 75.50 35.41 26.67 9 / 15 / 76 1.67

    Key observations:

    1. Diversity-only baseline: Selecting data via K-means clustering without difficulty scoring achieves 49.48 average, trailing random selection (50.61).
    2. High Conditioned Answer (CA) baseline: Selecting data based solely on high cross-entropy loss sθ(A∣Q)s_\theta(A \mid Q) produces the lowest score (47.51), because answer loss alone reflects intrinsic text complexity rather than instruction difficulty.
    3. Low IFD baseline: Selecting low-IFD data yields the worst pairwise human evaluation against Cherry 5% (87 losses to 5 wins), confirming that IFD correlates monotonically with instruction utility.
  7. Knowl 7 — Ablation of Sample Size, Distribution, and Architecture in the Brief Experience Phase

    empirical result

    Ablations on the brief experience phase demonstrate how sample volume, sample distribution, and model training stage affect downstream cherry model quality:

    1. Number of Pre-experienced Samples: Using 0 pre-experienced samples (scoring via the raw base model) selects effective data that outperforms full Alpaca at 10% data, but performs worse than pre-experienced scoring. Downstream performance improves at 100 samples and reaches a performance plateau at 300 samples; expanding pre-experience data to 500 or 1,000 samples yields no further gains, showing that 300 samples are sufficient to equip the model with the necessary instruction-following prior.

    2. Distribution of Pre-experienced Samples: Selecting 1,000 pre-experienced samples by difficulty (IFD on base model), diversity (KK-means), or uniform random sampling yields comparable downstream winning scores against full Alpaca (at 10% cherry data: 1.072 for Difficulty, 1.097 for Diversity, and 1.047 for Random), proving the pre-experience process is robust to subset sampling strategies.

    3. Fully-Trained Evaluator Failure: Using a fully-trained Alpaca model as the evaluator fails to identify effective cherry data (downstream winning scores against official Alpaca: 0.968 at 5%, 0.999 at 10%, 1.005 at 15%). This breakdown occurs because the distribution gap between a fully-trained model and the initial raw base model prevents the evaluator from estimating difficulty relevant to the raw model.

  8. Knowl 8 — Distributional, Syntactic, and Qualitative Characteristics of High-IFD Cherry Data

    empirical result

    Analysis of the Alpaca dataset reveals distinct structural and semantic characteristics separating high-IFD and low-IFD samples:

    1. Embedding Space Clustering: 2D t-SNE visualizations of instruction embeddings show clear boundaries separating top 5% and bottom 5% IFD samples, contradicting the assumption that high-quality data must be uniformly distributed across embedding space.

    2. Syntactic Verb-Noun Patterns: Syntactic parsing with the Berkeley Neural Parser shows that top 5% IFD data is dominated by open-ended creative and explanatory tasks: Write Story (119 occurrences), Generate Story (98), Generate List (66), Explain Concept (48), and Create Story (44). In contrast, bottom 5% IFD data consists of rule-following, surface-level edits: Rewrite Sentence (155), Edit Sentence (89), Change Sentence (37), Classify Sentence (36), and Convert Sentence (27).

    3. Qualitative Dimension Ratings: Across six linguistic dimensions evaluated by ChatGPT, high-IFD instructions score higher on Scope, Complexity, Depth, and Knowledge Required, but substantially lower on Simplicity and Clarity compared to low-IFD instructions.

  9. Knowl 9 — Sub-Category Performance Variations and Domain Limitations of Low-Data Cherry Models

    empirical result

    Evaluating cherry models across individual task categories in Vicuna and WizardLM test sets reveals performance variations across task types:

    1. Strong Categories: Cherry models match or outperform full-dataset baselines on open-ended generation, reasoning, and conversational tasks, including Writing (win/tie/loss ratios of 0.60/0.20/0.20 on Vicuna), Roleplay (0.60/0.10/0.30), Knowledge (0.60/0.20/0.20), and Common-sense (0.60/0.20/0.30).

    2. Weak Categories: Cherry models lag behind full-dataset models in rule-bound, symbolic domains:

      • On Vicuna test tasks, the 10% Cherry WizardLM model underperforms full WizardLM in Math (0.00 wins / 0.33 ties / 0.67 losses) and Coding (0.00 / 0.29 / 0.71).
      • On WizardLM test tasks, 10% Cherry WizardLM underperforms in ComplexFormat (0.25 / 0.25 / 0.50) and Counterfactual reasoning (0.00 / 0.38 / 0.62).

    Because base 7B models have weak innate priors for math and coding, reducing total data volume deprives the model of the high sample counts required to learn alignment in these specialized domains.

  10. Knowl 10 — Operational Trade-Off Between Pre-Experienced Scoring and Direct Base Model Scoring

    limitation

    The primary practical limitation of the IFD selection framework is the computational overhead of training an intermediate pre-experienced model before dataset evaluation.

    While computing IFD scores directly on an un-finetuned base model (as demonstrated on LLaMA2-7B and LLaMA2-13B) simplifies deployment and still yields effective data selection, incorporating the brief pre-experience phase remains theoretically superior and achieves higher empirical performance by establishing a basic instruction-following prior before calculating Conditioned Answer Scores.

Coverage note — None was omitted.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  2. 2.Yihan Cao, Yanbin Kang, and Lichao Sun. 2023. Instruction mining: High-quality instruction data selection for large language models.
  3. 3.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2023. A survey on evaluation of large language models.
  4. 4.Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. 2023a. Alpagasus: Training a better alpaca with fewer data.
  5. 5.Pei Chen, Soumajyoti Sarkar, Leonard Lausen, Balasubramaniam Srinivasan, Sheng Zha, Ruihong Huang, and George Karypis. 2023b. Hytrel: Hypergraph-enhanced tabular data representation learning. In Advances in Neural Information Processing Systems, volume 36, pages 32173–32193. Curran Associates, Inc.
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  7. 7.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge.
  8. 8.Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems.
  9. 9.Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233.
  10. 10.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. GLM: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335, Dublin, Ireland. Association for Computational Linguistics.
  11. 11.Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpacafarm: A simulation framework for methods that learn from human feedback.
  12. 12.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. A framework for few-shot language model evaluation.
  13. 13.Sariel Har-Peled and Akash Kushal. 2005. Smaller coresets for k-median and k-means clustering. In Proceedings of the twenty-first annual symposium on Computational geometry, pages 126–134.
  14. 14.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In International Conference on Learning Representations.
  15. 15.Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn’t always right. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051.
  16. 16.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2023. Unnatural instructions: Tuning language models with (almost) no human labor. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14409–14428, Toronto, Canada. Association for Computational Linguistics.
  17. 17.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. UNIFIEDQA: Crossing format boundaries with a single QA system. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1896–1907, Online. Association for Computational Linguistics.
  18. 18.Diederik P. Kingma and Jimmy Ba. 2017. Adam: A method for stochastic optimization.
  19. 19.Miyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim, and Jaewoo Kang. 2020. Look at the first sentence: Position bias in question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1109–1121, Online. Association for Computational Linguistics.
  20. 20.Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Jiuxiang Gu, and Tianyi Zhou. 2024a. Selective reflection-tuning: Student-selected data recycling for llm instruction-tuning. ArXiv, abs/2402.10110.
  21. 21.Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Heng Huang, Jiuxiang Gu, and Tianyi Zhou. 2023a. Reflection-tuning: Data recycling improves llm instruction-tuning. ArXiv, abs/2310.11716.
  22. 22.Ming Li, Yong Zhang, Shwai He, Zhitao Li, Hongyu Zhao, Jianzong Wang, Ning Cheng, and Tianyi Zhou. 2024b. Superfiltering: Weak-to-strong data filtering for fast instruction-tuning. ArXiv, abs/2402.00530.
  23. 23.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023b. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval.
  24. 24.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  25. 25.Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. 2024a. Mitigating hallucination in large multi-modal models via robust instruction tuning. In The Twelfth International Conference on Learning Representations.
  26. 26.Fuxiao Liu, Xiaoyang Wang, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, and Dong Yu. 2023. Mmc: Advancing multimodal chart understanding with large-scale instruction tuning.
  27. 27.Xiaoyu Liu, Paiheng Xu, Junda Wu, Jiaxin Yuan, Yifan Yang, Yuhang Zhou, Fuxiao Liu, Tianrui Guan, Haoliang Wang, Tong Yu, Julian McAuley, Wei Ai, and Furong Huang. 2024b. Large language models and causal inference in collaboration: A comprehensive survey.
  28. 28.S. Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. 2023. The flan collection: Designing data and methods for effective instruction tuning. ArXiv, abs/2301.13688.
  29. 29.Sören Mindermann, Jan M Brauner, Muhammed T Razzak, Mrinank Sharma, Andreas Kirsch, Winnie Xu, Benedikt Höltgen, Aidan N Gomez, Adrien Morisot, Sebastian Farquhar, et al. 2022. Prioritized training on points that are learnable, worth learning, and not yet learnt. In International Conference on Machine Learning, pages 15630–15649.
  30. 30.Lili Mou, Yiping Song, Rui Yan, Ge Li, Lu Zhang, and Zhi Jin. 2016. Sequence to backward and forward sequences: A content-introducing approach to generative short-text conversation. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 3349–3358.
  31. 31.Alexander Munteanu, Chris Schwiegelshohn, Christian Sohler, and David Woodruff. 2018. On coresets for logistic regression. volume 31.
  32. 32.OpenAI. 2023. Gpt-4 technical report.
  33. 33.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  34. 34.Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite. 2021. Deep learning on a data diet: Finding important examples early in training. In Advances in Neural Information Processing Systems.
  35. 35.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only.
  36. 36.Teven Le Scao, Angela Fan, Christopher Akiki, Elizabeth-Jane Pavlick, Suzana Ili’c, Daniel Hesslow, Roman Castagn’e, Alexandra Sasha Luccioni, Franc¸cois Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Rose Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurenccon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa Etxabe, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris C. Emezue, Christopher Klamm, Colin Leong, Daniel Alexander van Strien, David Ifeoluwa Adelani, Dragomir R. Radev, Eduardo Gonz’alez Ponferrada, Efrat Levkovizh, Ethan Kim, Eyal Bar Natan, Francesco De Toni, Gérard Dupont, Germán Kruszewski, Giada Pistilli, Hady ElSahar, Hamza Benyamina, Hieu Trung Tran, Ian Yu, Idris Abdulmumin, Isaac Johnson, Itziar Gonzalez-Dios, Javier de la Rosa, Jenny Chim, Jesse Dodge, Jian Zhu, Jonathan Chang, Jorg Frohberg, Josephine L. Tobing, Joydeep Bhattacharjee, Khalid Almubarak, Kimbo Chen, Kyle Lo, Leandro von Werra, Leon Weber, Long Phan, Loubna Ben Allal, Ludovic Tanguy, Manan Dey, Manuel Romero Muñoz, Maraim Masoud, Mar’ia Grandury, Mario vSavsko, Max Huang, Maximin Coavoux, and Mayank Singh. 2022. Bloom: A 176b-parameter open-access multilingual language model. ArXiv, abs/2211.05100.
  37. 37.Wangtao Sun, Haotian Xu, Xuanqing Yu, Pei Chen, Shizhu He, Jun Zhao, and Kang Liu. 2024. Itd: Large language models can teach themselves induction through deduction.
  38. 38.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  39. 39.Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon. 2018. An empirical study of example forgetting during deep neural network learning. In International Conference on Learning Representations.
  40. 40.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models.
  41. 41.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models.
  42. 42.Ivor W Tsang, James T Kwok, Pak-Ming Cheung, and Nello Cristianini. 2005. Core vector machines: Fast svm training on very large data sets. Journal of Machine Learning Research, 6(4).
  43. 43.Thuy-Trang Vu, Xuanli He, Gholamreza Haffari, and Ehsan Shareghi. 2023. Koala: An index for quantifying overlaps with pre-training corpora.
  44. 44.Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023a. Large language models are not fair evaluators.
  45. 45.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023b. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13484–13508, Toronto, Canada. Association for Computational Linguistics.
  46. 46.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen. 2022. Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5085–5109, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  47. 47.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  48. 48.Sarah Wiegreffe, Matthew Finlayson, Oyvind Tafjord, Peter Clark, and Ashish Sabharwal. 2023. Increasing probability mass on answer choices does not always improve accuracy. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8392–8417.
  49. 49.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions.
  50. 50.Xiaohan Xu, Ming Li, Chongyang Tao, Tao Shen, Reynold Cheng, Jinyang Li, Can Xu, Dacheng Tao, and Tianyi Zhou. 2024. A survey on knowledge distillation of large language models.
  51. 51.Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021. CrossFit: A few-shot learning challenge for cross-task generalization in NLP. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7163–7189, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  52. 52.Seonghyeon Ye, Yongrae Jo, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, and Minjoon Seo. 2023. Selfee: Iterative self-revising llm empowered by self-feedback generation. Blog post.
  53. 53.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy. Association for Computational Linguistics.
  54. 54.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.
  55. 55.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. Lima: Less is more for alignment.
  56. 56.Kun Zhou, Kai Zhang, Yu Wu, Shujie Liu, and Jingsong Yu. 2019. Unsupervised context rewriting for open domain conversation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1834–1844.

Citation

MLA
Li, M., et al. “From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 7602–35, https://doi.org/10.18653/v1/2024.naacl-long.421.
APA
Li, M., Zhang, Y., Li, Z., Chen, J., Chen, L., Cheng, N., Wang, J., Zhou, T., & Xiao, J. (2024). From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7602–7635. https://doi.org/10.18653/v1/2024.naacl-long.421
Chicago
Li, M., Y. Zhang, Z. Li, et al. 2024. “From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7602–35. https://doi.org/10.18653/v1/2024.naacl-long.421.
Harvard
Li, M. et al. (2024) “From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7602–7635. Available at: https://doi.org/10.18653/v1/2024.naacl-long.421.
Vancouver
1. Li M, Zhang Y, Li Z, Chen J, Chen L, Cheng N, Wang J, Zhou T, Xiao J (2024) From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 7602–7635

BibTeX

@inproceedings{li-etal-2024-quantity,
    title = "From Quantity to Quality: Boosting {LLM} Performance with Self-Guided Data Selection for Instruction Tuning",
    author = "Li, Ming  and
      Zhang, Yong  and
      Li, Zhitao  and
      Chen, Jiuhai  and
      Chen, Lichang  and
      Cheng, Ning  and
      Wang, Jianzong  and
      Zhou, Tianyi  and
      Xiao, Jing",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.421/",
    doi = "10.18653/v1/2024.naacl-long.421",
    pages = "7602--7635"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/