LESS: Selecting Influential Data for Targeted Instruction Tuning
Mengzhou XiaSadhika MalladiSuchin GururanganSanjeev AroraDanqi Chen
Proposes LESS, an efficient gradient-based data selection method for targeted instruction tuning that outperforms full-dataset training using only 5% of the data and transfers effectively across model sizes and families.
Instruction tuning is a vital process for aligning large language models to follow human directives. However, enterprise applications typically demand specialized competencies, such as mathematical reasoning or domain-specific question answering, rather than generic chat behavior. Training models on massive, mixed instruction pools frequently introduces irrelevant or conflicting data that can degrade targeted performance. Furthermore, high-quality, task-specific data is often scarce, making it critical to identify and extract the most relevant training instances from large, general-purpose datasets using only a handful of target examples.
The article introduces and evaluates LESS (Low-rank Gradient Similarity Search), an efficient algorithm designed to estimate data influence and select the most beneficial fine-tuning examples for specific target capabilities. The primary objective is to demonstrate that training on a small, carefully chosen subset of instruction data can match or exceed the performance of training on entire multi-source datasets.
To achieve this, the authors developed an optimizer-aware selection pipeline adapted to the Adam optimizer and variable-length text sequences, using cosine similarity to prevent biasing toward shorter inputs. The method builds an offline gradient datastore by applying parameter-efficient fine-tuning (LoRA) during a brief warmup stage and projecting gradient features into a low-dimensional space. The framework was evaluated across a diverse pool of approximately 270,000 instruction examples and tested on standard benchmarks—MMLU, TYDIQA, and BIG-Bench Hard—using base models including LLAMA-2 (7B and 13B), MISTRAL-7B, and the Pythia model family.
The key findings show that training models on just the top 5% of data selected by LESS consistently outperforms random selection by 2 to 5 percentage points across benchmarks. Remarkably, fine-tuning on this 5% subset frequently surpassed the performance achieved by training on 100% of the dataset, particularly on capable models like LLAMA-2-13B and MISTRAL-7B. In direct comparisons, LESS outperformed standard selection baselines based on keyword matching (BM25), n-gram distributions (DSIR), and hidden model representations (RDS). In addition, data selected using smaller models successfully transferred to train larger models and entirely different model families without requiring new gradient stores. Qualitative analysis confirmed that LESS identifies instances sharing the underlying reasoning structure required by the target task, rather than relying on superficial word overlap or shared language.
These results demonstrate that larger training volume does not guarantee superior capability; irrelevant data can introduce noise and trigger negative transfer. By reducing the required fine-tuning data to 5%, organizations can substantially lower model training compute costs, accelerate development timelines, and improve specialized performance. The high transferability of selected data means organizations can use small, lightweight models to curate datasets for larger enterprise deployments, further amortizing preparation costs.
Decision-makers should consider adopting targeted data selection workflows over brute-force fine-tuning on entire data pools. While computing the initial gradient datastore requires upfront computational investment (e.g., approximately 48 GPU hours for 270,000 examples), this represents a one-time cost that facilitates nearly instantaneous data curation for future downstream tasks. For specialized deployments, teams should pilot small-model data curation pipelines to build reusable gradient indices.
Several limitations warrant consideration. The framework requires a short warmup fine-tuning phase, as raw pre-trained models fail to generate accurate selection gradients. Furthermore, because sequence gradients are averaged across tokens, the method can experience edge cases with extremely long, open-ended generation tasks. While the empirical results provide high confidence in the method's effectiveness across evaluated benchmarks, users should recognize that optimizing validation cross-entropy loss does not always lead monotonically to higher generation accuracy.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). This work establishes the paradigm of instruction tuning across multi-task datasets, providing the core training framework that LESS aims to optimize via targeted data selection.
- Paper: Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt, Sören Mindermann et al. (2022). This paper introduces principled holdout loss estimation for training data selection, establishing the foundational concepts of influence-based sample efficiency that LESS adapts to gradient similarity in instruction tuning.
- Paper: Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks, Po-Nien Kung et al. (2023). This study formulates active task-level data selection for instruction tuning, directly preceding LESS's objective of curating targeted subsets from massive instruction repositories.
- Paper: Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor, Or Honovich et al. (2023). This paper presents the large-scale synthesis of varied instruction-tuning corpora, illustrating the data volume and redundancy challenges that motivate LESS's targeted selection algorithm.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). This work details the creation of synthetic instruction pools like Self-Instruct, providing the foundational datasets evaluated within LESS's data-selection benchmark.
- Paper: WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions, Can Xu et al. (2023). This paper introduces the WizardLM instruction benchmark, which forms a key source dataset from which LESS selects high-influence samples for complex reasoning tasks.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). This study analyzes low-rank and parameter-efficient adaptation methods, providing background for the low-rank gradient approximations utilized in LESS's feature datastore.
- Paper: DsDm: Model-Aware Dataset Selection with Datamodels, Logan Engstrom et al. (2024). This work extends model-aware data selection from targeted instruction tuning to pretraining by estimating document influence on downstream task losses using linear datamodels.
- Paper: From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning, Ming Li et al. (2024). This paper advances instruction data selection by introducing a self-guided difficulty metric to prune instruction datasets down to high-impact subsets without external models.
- Paper: InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions, Yifan Wang et al. (2024). This study applies targeted instruction data selection principles to continual learning by selecting replay samples based on instruction similarity and complexity metrics.
- Paper: From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning, Xuansheng Wu et al. (2024). This work explores the mechanistic and representational shifts that occur inside large language models during instruction tuning, offering deeper interpretability for why selective data tuning succeeds.
