LESS: Selecting Influential Data for Targeted Instruction Tuning

Mengzhou XiaSadhika MalladiSuchin GururanganSanjeev AroraDanqi Chen

article2024ICML449 citations

Proposes LESS, an efficient gradient-based data selection method for targeted instruction tuning that outperforms full-dataset training using only 5% of the data and transfers effectively across model sizes and families.

Listen

Instruction tuning is a vital process for aligning large language models to follow human directives. However, enterprise applications typically demand specialized competencies, such as mathematical reasoning or domain-specific question answering, rather than generic chat behavior. Training models on massive, mixed instruction pools frequently introduces irrelevant or conflicting data that can degrade targeted performance. Furthermore, high-quality, task-specific data is often scarce, making it critical to identify and extract the most relevant training instances from large, general-purpose datasets using only a handful of target examples.

The article introduces and evaluates LESS (Low-rank Gradient Similarity Search), an efficient algorithm designed to estimate data influence and select the most beneficial fine-tuning examples for specific target capabilities. The primary objective is to demonstrate that training on a small, carefully chosen subset of instruction data can match or exceed the performance of training on entire multi-source datasets.

To achieve this, the authors developed an optimizer-aware selection pipeline adapted to the Adam optimizer and variable-length text sequences, using cosine similarity to prevent biasing toward shorter inputs. The method builds an offline gradient datastore by applying parameter-efficient fine-tuning (LoRA) during a brief warmup stage and projecting gradient features into a low-dimensional space. The framework was evaluated across a diverse pool of approximately 270,000 instruction examples and tested on standard benchmarks—MMLU, TYDIQA, and BIG-Bench Hard—using base models including LLAMA-2 (7B and 13B), MISTRAL-7B, and the Pythia model family.

The key findings show that training models on just the top 5% of data selected by LESS consistently outperforms random selection by 2 to 5 percentage points across benchmarks. Remarkably, fine-tuning on this 5% subset frequently surpassed the performance achieved by training on 100% of the dataset, particularly on capable models like LLAMA-2-13B and MISTRAL-7B. In direct comparisons, LESS outperformed standard selection baselines based on keyword matching (BM25), n-gram distributions (DSIR), and hidden model representations (RDS). In addition, data selected using smaller models successfully transferred to train larger models and entirely different model families without requiring new gradient stores. Qualitative analysis confirmed that LESS identifies instances sharing the underlying reasoning structure required by the target task, rather than relying on superficial word overlap or shared language.

These results demonstrate that larger training volume does not guarantee superior capability; irrelevant data can introduce noise and trigger negative transfer. By reducing the required fine-tuning data to 5%, organizations can substantially lower model training compute costs, accelerate development timelines, and improve specialized performance. The high transferability of selected data means organizations can use small, lightweight models to curate datasets for larger enterprise deployments, further amortizing preparation costs.

Decision-makers should consider adopting targeted data selection workflows over brute-force fine-tuning on entire data pools. While computing the initial gradient datastore requires upfront computational investment (e.g., approximately 48 GPU hours for 270,000 examples), this represents a one-time cost that facilitates nearly instantaneous data curation for future downstream tasks. For specialized deployments, teams should pilot small-model data curation pipelines to build reusable gradient indices.

Several limitations warrant consideration. The framework requires a short warmup fine-tuning phase, as raw pre-trained models fail to generate accurate selection gradients. Furthermore, because sequence gradients are averaged across tokens, the method can experience edge cases with extremely long, open-ended generation tasks. While the empirical results provide high confidence in the method's effectiveness across evaluated benchmarks, users should recognize that optimizing validation cross-entropy loss does not always lead monotonically to higher generation accuracy.

Cover for LESS: Selecting Influential Data for Targeted Instruction Tuning

Abstract

Instruction tuning has unlocked powerful capabilities in large language models (LLMs), effectively using combined datasets to develop generalpurpose chatbots. However, real-world applications often require a specialized suite of skills (e.g., reasoning). The challenge lies in identifying the most relevant data from these extensive datasets to effectively develop specific capabilities, a setting we frame as targeted instruction tuning. We propose LESS, an optimizer-aware and practically efficient algorithm to effectively estimate data influences and perform Low-rank gradiEnt Similarity Search for instruction data selection. Crucially, LESS adapts existing influence formulations to work with the Adam optimizer and variable-length instruction data. LESS first constructs a highly reusable and transferable gradient datastore with low-dimensional gradient features and then selects examples based on their similarity to few-shot examples embodying a specific capability. Experiments show that training on a LESS-selected 5% of the data can often outperform training on the full dataset across diverse downstream tasks. Furthermore, the selected data is highly transferable: smaller models can be leveraged to select useful data for larger models and models from different families. Our qualitative analysis shows that our method goes beyond surface form cues to identify data that exemplifies the necessary reasoning skills for the intended downstream application.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries: Influence Formulation
  • 3 LESS: Estimating the Influence of Instructions
  • 3.1 Extension to Adam
  • 3.2 Addressing Abnormalities in Sequence-Level Gradients
  • 4 LESS: Estimating Influences Efficiently
  • 4.1 Efficient Gradient Feature Computation
  • 4.2 Data Selection Algorithm
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Baselines
  • 5.3 Main results
  • 6 Analysis
  • 6.1 Computational Complexity
  • 6.2 Qualitative Analysis
  • 7 Related Work
  • 8 Discussion and Future Work
  • References
  • A Training
  • A.1 Training Datasets
  • A.2 Training Details
  • B Evaluation Details
  • C Baseline Details
  • C.1 BM25
  • C.2 DSIR
  • D More Experiment Results
  • D.1 Comparing SGD and Adam Optimizers for Instruction Tuning
  • D.2 Selecting Data for SGD, SignGD, and Adam Gradients
  • D.3 Comparison to Base Models, and Training with Validation Data
  • D.4 Detailed Results on Using Different Projection Dimensions
  • D.5 Results on Pythia Models
  • D.6 Results on GSM8K and TruthfulQA
  • E Nuances in Algorithm Design
  • F Comparison to “Kernel Behavior” (Malladi et al., 2023; Jacot et al., 2018).
  • F.1 Does instruction tuning exhibit kernel behavior?
  • G Comparison to Data Models (Ilyas et al., 2022)
  • H Sequence Level Gradients in Instruction Tuning
  • H.1 Theoretical View
  • I Loss vs Accuracy
  • J Gradient Similarity Analysis
  • K Qualitative Analysis
  • K.1 Top Selected Examples for BBH and MMLU
  • K.2 Examples with Negative Gradient Similarity Scores
  • K.3 Subset Distribution of Selected Examples

Knowls

  1. Knowl 1 — Low-Rank Gradient Similarity Search (LESS) Algorithm

    algorithm

    LESS (Low-rank gradiEnt Similarity Search) selects a targeted subset of instruction-tuning data from a large candidate pool D\mathcal{D} that optimizes performance on a target evaluation task represented by a small validation set Dval\mathcal{D}_{\text{val}} partitioned into mm subtasks Dval(1),…,Dval(m)\mathcal{D}_{\text{val}}^{(1)}, \dots, \mathcal{D}_{\text{val}}^{(m)}.

    Input: Candidate instruction tuning dataset D\mathcal{D}, validation subtasks {Dval(j)}j=1m\{\mathcal{D}_{\text{val}}^{(j)}\}_{j=1}^m, selection model MSM_S parameterized by θ\theta, selection fraction kk (e.g., 5%5\%), warmup fraction (e.g., 5%5\%), warmup epochs N=4N = 4, projection dimension d=8192d = 8192
    Output: Selected training dataset Dtrain⊂D\mathcal{D}_{\text{train}} \subset \mathcal{D}
    1. Sample warmup subset Dwarmup⊂D\mathcal{D}_{\text{warmup}} \subset \mathcal{D} uniformly at random.
    2. Train MSM_S on Dwarmup\mathcal{D}_{\text{warmup}} using LoRA and the Adam optimizer for NN epochs, saving checkpoints {θi}i=1N\{\theta_i\}_{i=1}^N and average epoch learning rates {ηˉi}i=1N\{\bar{\eta}_i\}_{i=1}^N.
    3. Instantiate a random projection matrix Π∈RP×d\Pi \in \mathbb{R}^{P \times d} with Rademacher entries Πuv∼U({−1,1})\Pi_{uv} \sim \mathcal{U}(\{-1, 1\}), where PP is the number of trainable LoRA parameters.
    4. Build Gradient Datastore: For each candidate training point z∈Dz \in \mathcal{D} and each checkpoint θi\theta_i (i=1,…,Ni=1, \dots, N):
         a. Compute LoRA gradient ∇^ℓ(z;θi)\hat{\nabla} \ell(z; \theta_i) and Adam optimizer update vector Γ^(z,θi)\hat{\Gamma}(z, \theta_i).
         b. Project to low dimension: Γ~(z,θi)=Π⊤Γ^(z,θi)∈Rd\tilde{\Gamma}(z, \theta_i) = \Pi^\top \hat{\Gamma}(z, \theta_i) \in \mathbb{R}^d.
    5. For each validation subtask j∈{1,…,m}j \in \{1, \dots, m\} and each checkpoint θi\theta_i (i=1,…,Ni=1, \dots, N):
         a. Compute projected validation gradient for each z′∈Dval(j)z' \in \mathcal{D}_{\text{val}}^{(j)}: ∇~ℓ(z′;θi)=Π⊤∇^ℓ(z′;θi)\tilde{\nabla} \ell(z'; \theta_i) = \Pi^\top \hat{\nabla} \ell(z'; \theta_i).
         b. Compute subtask mean gradient feature: ∇ˉℓ(Dval(j);θi)=1∣Dval(j)∣∑z′∈Dval(j)∇~ℓ(z′;θi)\bar{\nabla} \ell(\mathcal{D}_{\text{val}}^{(j)}; \theta_i) = \frac{1}{|\mathcal{D}_{\text{val}}^{(j)}|} \sum_{z' \in \mathcal{D}_{\text{val}}^{(j)}} \tilde{\nabla} \ell(z'; \theta_i).
    6. For each candidate training point z∈Dz \in \mathcal{D}:
         a. For each subtask j∈{1,…,m}j \in \{1, \dots, m\}, compute multi-epoch cosine influence:
            InfAdam(z,Dval(j))=∑i=1Nηˉi⟨∇ˉℓ(Dval(j);θi),Γ~(z,θi)⟩∥∇ˉℓ(Dval(j);θi)∥∥Γ~(z,θi)∥\text{Inf}_{\text{Adam}}(z, \mathcal{D}_{\text{val}}^{(j)}) = \sum_{i=1}^N \bar{\eta}_i \frac{\langle \bar{\nabla} \ell(\mathcal{D}_{\text{val}}^{(j)}; \theta_i), \tilde{\Gamma}(z, \theta_i) \rangle}{\|\bar{\nabla} \ell(\mathcal{D}_{\text{val}}^{(j)}; \theta_i)\| \|\tilde{\Gamma}(z, \theta_i)\|}
         b. Assign aggregate influence score: S(z)=max⁡j∈{1,…,m}InfAdam(z,Dval(j))S(z) = \max_{j \in \{1, \dots, m\}} \text{Inf}_{\text{Adam}}(z, \mathcal{D}_{\text{val}}^{(j)}).
    7. Construct Dtrain\mathcal{D}_{\text{train}} by selecting the top k∣D∣k|\mathcal{D}| instances with the highest scores S(z)S(z).
    8. return Dtrain\mathcal{D}_{\text{train}}

    Steps 1 to 4 are performed offline once per candidate dataset D\mathcal{D}. The online data selection (Steps 5 to 7) takes less than one minute on modern GPU hardware.

  2. Knowl 2 — Adam Influence Formulation for Instruction Data

    definition

    Let a model parameterized by θ\theta be trained for NN epochs using the Adam optimizer, where θi\theta_i denotes the model checkpoint after the ii-th epoch and ηˉi\bar{\eta}_i is the average learning rate during the ii-th epoch. For an autoregressive instruction datapoint z=(s,c)z = (s, c) comprising an instruction sequence ss and a completion sequence cc, the loss ℓ(z;θ)\ell(z; \theta) is the average cross-entropy over tokens in cc.

    The Adam parameter update vector at step tt is: Γ(z,θt)≜mt+1vt+1+ϵ\Gamma(z, \theta^t) \triangleq \frac{m_{t+1}}{\sqrt{v_{t+1}} + \epsilon} where mt+1=β1mt+(1−β1)∇ℓ(z;θt)1−β1tm_{t+1} = \frac{\beta_1 m_t + (1 - \beta_1) \nabla \ell(z; \theta^t)}{1 - \beta_1^t} and vt+1=β2vt+(1−β2)(∇ℓ(z;θt))21−β2tv_{t+1} = \frac{\beta_2 v_t + (1 - \beta_2) (\nabla \ell(z; \theta^t))^2}{1 - \beta_2^t}, with operations applied elementwise.

    The trajectory influence of a candidate training instance zz on a validation instance z′z' under Adam optimization is defined as: InfAdam(z,z′)≜∑i=1Nηˉicos⁡(∇ℓ(z′;θi),Γ(z,θi))=∑i=1Nηˉi⟨∇ℓ(z′;θi),Γ(z,θi)⟩∥∇ℓ(z′;θi)∥∥Γ(z,θi)∥\text{Inf}_{\text{Adam}}(z, z') \triangleq \sum_{i=1}^N \bar{\eta}_i \cos(\nabla \ell(z'; \theta_i), \Gamma(z, \theta_i)) = \sum_{i=1}^N \bar{\eta}_i \frac{\langle \nabla \ell(z'; \theta_i), \Gamma(z, \theta_i) \rangle}{\|\nabla \ell(z'; \theta_i)\| \|\Gamma(z, \theta_i)\|} where cos⁡(u,v)=⟨u,v⟩∥u∥∥v∥\cos(u, v) = \frac{\langle u, v \rangle}{\|u\| \|v\|} is the cosine similarity between vectors uu and vv.

  3. Knowl 3 — Mitigating Sequence-Length Gradient Bias via Cosine Normalization

    empirical result

    In autoregressive instruction tuning, sequence-level gradients are computed by averaging token-level loss gradients across completion tokens. Consequently, the gradient norm ∥∇ℓ(z;θ)∥\|\nabla \ell(z; \theta)\| and the Adam update norm ∥Γ(z,θ)∥\|\Gamma(z, \theta)\| are strongly negatively correlated with the completion length of sequence zz.

    Using unnormalized inner products (standard first-order influence formulations) heavily biases selection toward examples with very short completions, degrading downstream performance. Replacing inner products with cosine similarity normalizes the gradient features and resolves this bias, as shown below on LLAMA-2-7B trained on the top 5% of data selected from a ~270K candidate pool:

    Random (5%) Vanilla Gradients Normalized Gradients
    Task Score Avg Len Score Avg Len Score Avg Len
    MMLU (5-shot acc) 46.5 105.4 47.3 11.3 50.2 15.1
    TYDIQA (1-shot F1) 52.7 105.4 52.3 19.7 56.2 44.4
    BBH (3-shot EM) 38.9 105.4 37.4 19.0 41.5 101.2
  4. Knowl 4 — Performance of LESS across Model Architectures and Tasks

    data/table

    Instruction tuning on a 5% subset of data selected by LESS consistently outperforms random 5% selection by 2 to 5 percentage points across diverse benchmarks (MMLU, TYDIQA, BBH) and model families (LLAMA-2-7B, LLAMA-2-13B, MISTRAL-7B). Furthermore, training on just 5% of LESS-selected data often matches or outperforms training on the full 100% dataset (~270k instances consisting of FLAN V2, COT, DOLLY, and OPEN ASSISTANT 1).

    MMLU TYDIQA BBH
    Model Full Rand. LESS-T LESS Full Rand. LESS-T LESS Full Rand. LESS-T LESS
    (100%) (5%) (5%) (5%) (100%) (5%) (5%) (5%) (100%) (5%) (5%) (5%)
    LLAMA-2-7B 51.6 46.5 - 50.2 54.0 52.7 - 56.2 43.2 38.9 - 41.5
    (0.5) (0.5) (0.4) (0.7) (0.5) (0.6)
    LLAMA-2-13B 54.5 53.4 54.6 54.0 54.3 53.0 57.5 54.6 50.8 47.0 49.9 50.6
    (0.1) (0.3) (0.7) (1.3) (0.8) (0.3) (1.6) (0.5) (0.6)
    MISTRAL-7B 60.4 60.0 60.6 61.8 57.7 56.9 61.7 60.3 53.0 54.5 56.0 56.0
    (0.1) (0.3) (0.4) (0.2) (1.7) (2.4) (0.1) (0.8) (1.0)

    LESS-T denotes transferring data selected using LLAMA-2-7B gradient features to train larger models (LLAMA-2-13B) or different model families (MISTRAL-7B). Standard deviations across three random seeds are in parentheses.

  5. Knowl 5 — Comparison of LESS with Surface and Representation-Based Selection Baselines

    data/table

    When selecting the top 5% of training data from a candidate pool of ~270K instruction examples for LLAMA-2-7B, LESS significantly outperforms lexical frequency matching (BM25), n-gram importance resampling (DSIR), and final-layer representation similarity (RDS).

    Benchmark Rand. (5%) BM25 DSIR RDS LESS (5%) Δ\Delta
    MMLU 46.5 (0.5) 47.6 46.1 (0.3) 45.0 (1.0) 50.2 (0.5) ↑2.6\uparrow 2.6
    TYDIQA 52.7 (0.4) 52.7 44.5 (1.7) 46.8 (1.3) 56.2 (0.7) ↑3.5\uparrow 3.5
    BBH 38.9 (0.5) 39.8 36.8 (0.1) 36.7 (1.3) 41.5 (0.6) ↑1.7\uparrow 1.7

    Δ\Delta denotes the improvement of LESS over the strongest non-random baseline on each evaluation benchmark. Baselines relying on surface-form word overlap or representations select data matching topic/language rather than reasoning mechanics, resulting in minimal or negative gains over random selection.

  6. Knowl 6 — Cross-Model and Cross-Scale Transferability of Gradient Stores

    empirical result

    Data selected using the low-dimensional gradient datastore of a smaller selection model MSM_S transfers effectively to train larger models or models from different families MTM_T (the LESS-T setting):

    1. Cross-Scale & Cross-Family Transfer: Using LLAMA-2-7B as MSM_S to select 5% of the data yields 54.6% on MMLU, 57.5% on TYDIQA, and 49.9% on BBH when training LLAMA-2-13B, matching or exceeding training LLAMA-2-13B with its own gradient store (54.0%, 54.6%, 50.6%). For MISTRAL-7B, LLAMA-2-7B selection yields 60.6% on MMLU, 61.7% on TYDIQA, and 56.0% on BBH.
    2. Small-to-Large Scaling within Pythia: In the Pythia model suite (14M, 410M, 1B, 6.9B, 12B) evaluated on TYDIQA, data selected using Pythia-14M (which scores 0.0% zero-shot accuracy itself) improves Pythia-12B performance from 14.9% (Random 5%) to 17.7%, and selection using Pythia-1B improves Pythia-12B to 32.6% (close to 33.2% on Full 100% data).
    3. Transfer Limits: Transferability across disparate architectures is asymmetric; while LLAMA-2-7B selects effectively for MISTRAL-7B, gradient stores built from Pythia models fail to outperform random selection when used to train LLAMA-2 models.
  7. Knowl 7 — Necessity of LoRA Warmup and Multi-Checkpoint Trajectory Aggregation

    empirical result

    The effectiveness of LESS depends on warmup fine-tuning on candidate training distribution samples and tracking gradient trajectories across multiple checkpoints:

    1. LoRA Warmup vs Off-the-Shelf Models: Computing influence using un-warmed base LLAMA-2-7B (0% warmup) achieves an average score of 46.2 across MMLU, TYDIQA, and BBH; LLAMA-2-7B-CHAT achieves 46.2. Training LoRA on 5% of the candidate data for warmup yields an average of 49.3 (rising to 49.9 with 25% warmup data, and 50.5 with 100%).
    2. Checkpoint Trajectory: Aggregating influence over N=4N=4 epoch checkpoints yields an average score of 49.3 (MMLU: 50.2, TYDIQA: 56.2, BBH: 41.5), outperforming a single checkpoint N=1N=1 at 47.8 (MMLU: 48.2, TYDIQA: 54.9, BBH: 40.2) and random selection at 46.0, indicating that early-stage optimization dynamics provide essential gradient signals.
    3. Optimizer Formulations: Adam-aware influence (InfAdam) achieves 50.5 average score compared to 49.7 for SGD influence (InfSGD) and 47.8 for SignGD influence when using 100% warmup training.
  8. Knowl 8 — Impact of Gradient Random Projection Dimension and Computational Cost

    empirical result

    LESS utilizes random Rademacher projections Π∈RP×d\Pi \in \mathbb{R}^{P \times d} to project high-dimensional LoRA gradients (P≈135MP \approx 135\text{M} parameters in LLAMA-2-7B, ∼1.95%\sim 1.95\% of total parameters) to dd dimensions.

    Empirical evaluation of projection dimensions d∈{1024,2048,4096,8192}d \in \{1024, 2048, 4096, 8192\} for LLAMA-2-7B data selection across MMLU, TYDIQA, and BBH shows that even small dimensions outperform random selection (45.2 average), with performance scaling monotonically with dd:

    • d=1024d=1024: 48.4 average score
    • d=2048d=2048: 48.8 average score
    • d=4096d=4096: 49.2 average score
    • d=8192d=8192: 49.7 average score

    Computational and storage requirements on a single A100 (80GB) GPU for candidate pool size ∣D∣≈270K|\mathcal{D}| \approx 270\text{K}, ∣Dwarmup∣=0.05∣D∣| \mathcal{D}_{\text{warmup}}| = 0.05 |\mathcal{D}|, N=4N=4 checkpoints, and d=8192d=8192 are:

    • Warmup LoRA Training: O(∣Dwarmup∣⋅N)\mathcal{O}(|\mathcal{D}_{\text{warmup}}| \cdot N), takes 6 GPU hours (one-time offline cost).
    • Gradient Feature Computation: O(∣D∣⋅N)\mathcal{O}(|\mathcal{D}| \cdot N), takes 48 GPU hours, requiring 17.7 GB storage for O(∣D∣⋅N⋅d)\mathcal{O}(|\mathcal{D}| \cdot N \cdot d) features (one-time offline cost).
    • Online Data Selection: O(∣D∣⋅∣Dval∣⋅d)\mathcal{O}(|\mathcal{D}| \cdot |\mathcal{D}_{\text{val}}| \cdot d), takes <1< 1 minute.
  9. Knowl 9 — Reasoning-Structure Alignment versus Surface-Form Matching

    empirical result

    Qualitative analysis reveals that LESS identifies training examples that share the underlying problem-solving and reasoning structures of the validation query, whereas baseline methods latch onto surface-level lexical or linguistic patterns:

    1. Multilingual QA (TYDIQA Bengali validation example): Given a single Bengali open-book question-answering validation instance, BM25 selects Bengali masked word prediction instances and RDS selects Bengali hate-speech classification instances due to language overlap. In contrast, LESS selects English open-book question-answering instances from DOLLY that embody the reading comprehension reasoning format, facilitating cross-lingual zero-shot task transfer.
    2. Dataset Selection Distributions: When selecting 5% subsets across benchmarks, LESS naturally aligns with task structures without manual heuristics:
      • For MMLU (multiple-choice academic QA), LESS predominantly selects multiple-choice instances from FLAN V2.
      • For TYDIQA (reading comprehension QA), LESS predominantly selects passage-based open-book QA instances from DOLLY.
      • For BBH (multi-step reasoning tasks), LESS predominantly selects step-by-step reasoning instances from COT.
  10. Knowl 10 — Limitations of the LESS Framework

    limitation

    The LESS data selection framework possesses four primary limitations:

    1. Mandatory Warmup Phase: Un-finetuned base LLMs and off-the-shelf instruction-tuned chat models produce ineffective gradient features; the selection model must undergo a preliminary warmup training run on a subset of the target instruction pool to align with its data distribution.
    2. Sequence-Level Gradient Averaging: Aggregating token gradients into a single sequence gradient serves as an imperfect surrogate for sequence-level influence, which can lead to ambiguities in long-sequence or open-ended generation settings.
    3. Loss-Accuracy Disconnect: LESS optimizes for minimizing validation cross-entropy loss, but cross-entropy loss reduction does not always correlate monotonically with task accuracy in complex long-form instruction tasks.
    4. Linearity Assumption in Data Scoring: First-order Taylor approximation scores candidate datapoints independently and sums their effects linearly, ignoring higher-order interactions and redundancy between co-selected training examples (e.g., scoring duplicate examples identically).

Coverage note — None was omitted; all core methodology, influence derivations, experimental results across models (Llama, Mistral, Pythia), baseline comparisons, scaling ablations, qualitative findings, and explicit limitations are captured in the knowls.

References

  1. 1.Anthropic. Introducing claude, 2023. URL https://www.anthropic.com/index/introducing-claude.
  2. 2.Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A. A latent variable model approach to pmi-based word embeddings. Transactions of the Association for Computational Linguistics, 4:385–399, 2016.
  3. 3.Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R. On exact computation with an infinitely wide neural net. Advances in neural information processing systems, 32, 2019.
  4. 4.Bae, J., Ng, N. H., Lo, A., Ghassemi, M., and Grosse, R. B. If influence functions are the answer, then what is the question? In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=hzbguA9zMJ.
  5. 5.Basu, S., Pope, P., and Feizi, S. Influence functions in deep learning are fragile. In International Conference on Learning Representations, 2020.
  6. 6.Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pp. 2397–2430. PMLR, 2023.
  7. 7.Bukharin, A. and Zhao, T. Data diversity matters for robust instruction tuning. arXiv preprint arXiv:2311.14736, 2023.
  8. 8.Cao, Y., Kang, Y., and Sun, L. Instruction mining: High-quality instruction data selection for large language models. arXiv preprint arXiv:2307.06290, 2023.
  9. 9.Chen, L., Li, S., Yan, J., Wang, H., Gunaratna, K., Yadav, V., Tang, Z., Srinivasan, V., Zhou, T., Huang, H., et al. Alpagasus: Training a better alpaca with fewer data. arXiv preprint arXiv:2307.08701, 2023a.
  10. 10.Chen, M. F., Roberts, N., Bhatia, K., WANG, J., Zhang, C., Sala, F., and Re, C. Skill-it! a data-driven skills framework for understanding and training language models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023b. URL https://openreview.net/forum?id=IoizwO1NLf.
  11. 11.Clark, J. H., Choi, E., Collins, M., Garrette, D., Kwiatkowski, T., Nikolaev, V., and Palomaki, J. TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 2020.
  12. 12.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  13. 13.Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M. Selection via proxy: Efficient data selection for deep learning. In International Conference on Learning Representations, 2019.
  14. 14.Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., and Xin, R. Free Dolly: Introducing the world’s first truly open instruction-tuned LLM, 2023.
  15. 15.Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B. Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233, 2023.
  16. 16.Du, Q., Zong, C., and Zhang, J. Mods: Model-oriented data selection for instruction tuning. arXiv preprint arXiv:2311.15653, 2023.
  17. 17.Engstrom, L., Feldmann, A., and Madry, A. Dsdm: Model-aware dataset selection with datamodels, 2024.
  18. 18.Feldman, V. and Zhang, C. What neural networks memorize and why: Discovering the long tail via influence estimation. In Advances in Neural Information Processing Systems, volume 33, pp. 2881–2891, 2020.
  19. 19.Google. An important next step on our ai journey, 2023. URL https://blog.google/technology/ai/bard-google-ai-search-updates/.
  20. 20.Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., Hubinger, E., Lukošiūtė, K., Nguyen, K., Joseph, N., McCandlish, S., Kaplan, J., and Bowman, S. R. Studying large language model generalization with influence functions, 2023.
  21. 21.Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 2020.
  22. 22.Guu, K., Webson, A., Pavlick, E., Dixon, L., Tenney, I., and Bolukbasi, T. Simfluence: Modeling the influence of individual training examples by simulating training runs, 2023.
  23. 23.Hampel, F. R. The influence curve and its role in robust estimation. Journal of the american statistical association, 69(346):383–393, 1974.
  24. 24.Han, X. In-context alignment: Chat with vanilla language models before fine-tuning. arXiv preprint arXiv:2308.04275, 2023.
  25. 25.Han, X., Simig, D., Mihaylov, T., Tsvetkov, Y., Celikyilmaz, A., and Wang, T. Understanding in-context learning via supportive pretraining data. In The 61st Annual Meeting Of The Association For Computational Linguistics, 2023.
  26. 26.Hanawa, K., Yokoi, S., Hara, S., and Inui, K. Evaluation of similarity-based explanations. In International Conference on Learning Representations, 2020.
  27. 27.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2020.
  28. 28.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  29. 29.Hu, J., Ruder, S., Siddhant, A., Neubig, G., Firat, O., and Johnson, M. XTREME: A massively multilingual multitask benchmark for evaluating cross-lingual generalisation. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp. 4411–4421. PMLR, 13–18 Jul 2020. URL https://proceedings.mlr.press/v119/hu20b.html.
  30. 30.Ilyas, A., Park, S. M., Engstrom, L., Leclerc, G., and Madry, A. Datamodels: Predicting predictions from training data. In Proceedings of the 39th International Conference on Machine Learning, 2022.
  31. 31.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. Advances in neural information processing systems, 31, 2018.
  32. 32.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  33. 33.Johnson, W. B. and Lindenstrauss, J. Extensions of lipschitz mappings into hilbert space. Contemporary mathematics, 26:189–206, 1984.
  34. 34.Kaushal, V., Iyer, R., Kothawade, S., Mahadev, R., Doctor, K., and Ramakrishnan, G. Learning from less data: A unified data subset selection and active learning framework for computer vision. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1289–1299. IEEE, 2019.
  35. 35.Killamsetty, K., Durga, S., Ramakrishnan, G., De, A., and Iyer, R. Grad-match: Gradient matching based data subset selection for efficient deep model training. In International Conference on Machine Learning, pp. 5464–5474. PMLR, 2021a.
  36. 36.Killamsetty, K., Zhao, X., Chen, F., and Iyer, R. Retrieve: Coreset selection for efficient and robust semi-supervised learning. In Advances in Neural Information Processing Systems, volume 34, pp. 14488–14501, 2021b.
  37. 37.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Bengio, Y. and LeCun, Y. (eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. URL http://arxiv.org/abs/1412.6980.
  38. 38.Koh, P. W. and Liang, P. Understanding black-box predictions via influence functions. In International conference on machine learning, pp. 1885–1894. PMLR, 2017.
  39. 39.Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., et al. OpenAssistant conversations–democratizing large language model alignment. In Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks), 2023.
  40. 40.Li, Y., Hui, B., Xia, X., Yang, J., Yang, M., Zhang, L., Si, S., Liu, J., Liu, T., Huang, F., et al. One shot learning as instruction data prospector for large language models. arXiv preprint arXiv:2312.10302, 2023.
  41. 41.Lin, B. Y., Ravichander, A., Lu, X., Dziri, N., Sclar, M., Chandu, K., Bhagavatula, C., and Choi, Y. The unlocking spell on base LLMs: Rethinking alignment via in-context learning. arXiv preprint arXiv:2312.01552, 2023a.
  42. 42.Lin, S., Hilton, J., and Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. In Association for Computational Linguistics (ACL), 2022.
  43. 43.Lin, Y., Tan, L., Lin, H., Zheng, Z., Pi, R., Zhang, J., Diao, S., Wang, H., Zhao, H., Yao, Y., et al. Speciality vs generality: An empirical study on catastrophic forgetting in fine-tuning foundation models. arXiv preprint arXiv:2309.06256, 2023b.
  44. 44.Liu, W., Zeng, W., He, K., Jiang, Y., and He, J. What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning. arXiv preprint arXiv:2312.15685, 2023.
  45. 45.Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688, 2023.
  46. 46.Madsen, A., Reddy, S., and Chandar, S. Post-hoc interpretability for neural nlp: A survey. ACM Computing Surveys, 55(8):1–42, 2022.
  47. 47.Malladi, S., Lyu, K., Panigrahi, A., and Arora, S. On the sdes and scaling rules for adaptive gradient algorithms. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp. 7697–7711. Curran Associates, Inc., 2022.
  48. 48.Malladi, S., Wettig, A., Yu, D., Chen, D., and Arora, S. A kernel-based view of language model fine-tuning. In International Conference on Machine Learning, pp. 23610–23641. PMLR, 2023.
  49. 49.Meding, K., Buschoff, L. M. S., Geirhos, R., and Wichmann, F. A. Trivial or impossible—dichotomous data difficulty masks model differences (on imagenet and beyond). In International Conference on Learning Representations, 2021.
  50. 50.Mindermann, S., Brauner, J. M., Razzak, M. T., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., et al. Prioritized training on points that are learnable, worth learning, and not yet learnt. In International Conference on Machine Learning, pp. 15630–15649. PMLR, 2022.
  51. 51.Mirzasoleiman, B., Bilmes, J., and Leskovec, J. Coresets for data-efficient training of machine learning models. In International Conference on Machine Learning, pp. 6950–6960. PMLR, 2020.
  52. 52.Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A. Orca: Progressive learning from complex explanation traces of gpt-4. arXiv preprint arXiv:2306.02707, 2023.
  53. 53.Nguyen, T., Chen, Z., and Lee, J. Dataset meta-learning from kernel ridge-regression. In International Conference on Learning Representations, 2021.
  54. 54.OpenAI. OpenAI: Introducing ChatGPT, 2022. URL https://openai.com/blog/chatgpt.
  55. 55.OpenAI. OpenAI: GPT-4, 2023. URL https://openai.com/research/gpt-4.
  56. 56.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  57. 57.Park, S. M., Georgiev, K., Ilyas, A., Leclerc, G., and Madry, A. Trak: Attributing model behavior at scale. In International Conference on Machine Learning (ICML), 2023.
  58. 58.Paul, M., Ganguli, S., and Dziugaite, G. K. Deep learning on a data diet: Finding important examples early in training. In Advances in Neural Information Processing Systems, volume 34, pp. 20596–20607, 2021.
  59. 59.Phillips, J. M. Coresets and sketches. In Handbook of discrete and computational geometry, pp. 1269–1288. Chapman and Hall/CRC, 2017.
  60. 60.Pruthi, G., Liu, F., Kale, S., and Sundararajan, M. Estimating training data influence by tracing gradient descent. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 19920–19930. Curran Associates, Inc., 2020.
  61. 61.Robertson, S., Zaragoza, H., et al. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389, 2009.
  62. 62.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Le Scao, T., Raja, A., et al. Multitask prompted training enables zero-shot task generalization. In ICLR 2022-Tenth International Conference on Learning Representations, 2022.
  63. 63.Saunshi, N., Gupta, A., Braverman, M., and Arora, S. Understanding influence functions and datamodels via harmonic analysis. In The Eleventh International Conference on Learning Representations, 2023.
  64. 64.Sener, O. and Savarese, S. Active learning for convolutional neural networks: A core-set approach. In International Conference on Learning Representations, 2018.
  65. 65.Søgaard, A. et al. Revisiting methods for finding influential examples. arXiv preprint arXiv:2111.04683, 2021.
  66. 66.Sorscher, B., Geirhos, R., Shekhar, S., Ganguli, S., and Morcos, A. Beyond neural scaling laws: beating power law scaling via data pruning. Advances in Neural Information Processing Systems, 35:19523–19536, 2022.
  67. 67.Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., et al. Challenging big-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, pp. 13003–13051, 2023.
  68. 68.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  69. 69.Tay, Y., Dehghani, M., Rao, J., Fedus, W., Abnar, S., Chung, H. W., Narang, S., Yogatama, D., Vaswani, A., and Metzler, D. Scale efficiently: Insights from pretraining and finetuning transformers. In International Conference on Learning Representations, 2021.
  70. 70.Toneva, M., Sordoni, A., des Combes, R. T., Trischler, A., Bengio, Y., and Gordon, G. J. An empirical study of example forgetting during deep neural network learning. In International Conference on Learning Representations, 2018.
  71. 71.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  72. 72.Wang, X., Pham, H., Michel, P., Anastasopoulos, A., Carbonell, J., and Neubig, G. Optimizing data usage via differentiable rewards. In International Conference on Machine Learning, pp. 9983–9995. PMLR, 2020.
  73. 73.Wang, X., Zhu, W., Saxon, M., Steyvers, M., and Wang, W. Y. Large language models are latent variable models: Explaining and finding good demonstrations for in-context learning. In Thirty-seventh Conference on Neural Information Processing Systems, 2023a.
  74. 74.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-Instruct: Aligning language models with self-generated instructions. In Association for Computational Linguistics (ACL), pp. 13484–13508.
  75. 75.Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Naik, A., Ashok, A., Dhanasekaran, A. S., Arunkumar, A., Stap, D., et al. Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 5085–5109, 2022.
  76. 76.Wang, Y., Ivison, H., Dasigi, P., Hessel, J., Khot, T., Chandu, K. R., Wadden, D., MacMillan, K., Smith, N. A., Beltagy, I., et al. How far can camels go? exploring the state of instruction tuning on open resources. In Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks), 2023b.
  77. 77.Wei, A., Hu, W., and Steinhardt, J. More than a toy: Random matrix models predict how real-world neural representations generalize. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 23549–23588. PMLR, 17–23 Jul 2022a. URL https://proceedings.mlr.press/v162/wei22a.html.
  78. 78.Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. In International Conference on Learning Representations (ICLR), 2022b.
  79. 79.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 24824–24837, 2022c.
  80. 80.Xia, M., Artetxe, M., Zhou, C., Lin, X. V., Pasunuru, R., Chen, D., Zettlemoyer, L., and Stoyanov, V. Training trajectories of language models across scales. In Association for Computational Linguistics (ACL), pp. 13711–13738.
  81. 81.Xia, M., Anastasopoulos, A., Xu, R., Yang, Y., and Neubig, G. Predicting performance for natural language processing tasks. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8625–8646, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.764. URL https://aclanthology.org/2020.acl-main.764.
  82. 82.Xia, X., Liu, J., Yu, J., Shen, X., Han, B., and Liu, T. Moderate coreset: A universal method of data selection for real-world data-efficient deep learning. In The Eleventh International Conference on Learning Representations, 2023.
  83. 83.Xie, S. M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q. V., Ma, T., and Yu, A. W. DoReMi: Optimizing data mixtures speeds up language model pretraining. In Thirty-seventh Conference on Neural Information Processing Systems, 2023a. URL https://openreview.net/forum?id=lXuByUeHhd.
  84. 84.Xie, S. M., Santurkar, S., Ma, T., and Liang, P. Data selection for language models via importance resampling. In Advances in Neural Information Processing Systems (NeurIPS), 2023b.
  85. 85.Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D. WizardLM: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244, 2023.
  86. 86.Yang, G. and Hu, E. J. Tensor programs iv: Feature learning in infinite-width neural networks. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 11727–11737. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/yang21c.html.
  87. 87.Yu, D., Wang, R., Li, Z., Hu, W., Salakhutdinov, R., Arora, S., and Du, S. S. Enhanced convolutional neural tangent kernels, 2020a. URL https://openreview.net/forum?id=BkgNqkHFPr.
  88. 88.Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C. Gradient surgery for multi-task learning. Advances in Neural Information Processing Systems, 33: 5824–5836, 2020b.
  89. 89.Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595, 2018.
  90. 90.Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al. LIMA: Less is more for alignment. In Advances in Neural Information Processing Systems (NeurIPS), 2023.

Citation

MLA
Xia, M., et al. “LESS: Selecting Influential Data for Targeted Instruction Tuning”. arXiv, 2024, http://arxiv.org/abs/2402.04333v3.
APA
Xia, M., Malladi, S., Gururangan, S., Arora, S., & Chen, D. (2024). LESS: Selecting Influential Data for Targeted Instruction Tuning. arXiv. http://arxiv.org/abs/2402.04333v3
Chicago
Xia, M., S. Malladi, S. Gururangan, S. Arora, and D. Chen. 2024. “LESS: Selecting Influential Data for Targeted Instruction Tuning”. arXiv. http://arxiv.org/abs/2402.04333v3.
Harvard
Xia, M. et al. (2024) “LESS: Selecting Influential Data for Targeted Instruction Tuning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.04333v3.
Vancouver
1. Xia M, Malladi S, Gururangan S, Arora S, Chen D (2024) LESS: Selecting Influential Data for Targeted Instruction Tuning. arXiv

BibTeX

@article{xia2024less,
  title = {LESS: Selecting Influential Data for Targeted Instruction Tuning},
  author = {Xia, Mengzhou and Malladi, Sadhika and Gururangan, Suchin and Arora, Sanjeev and Chen, Danqi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.04333v3},
  eprint = {2402.04333}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/