Built independently by an author, for readers. Read the story and support ChapterPal

keyword

optimizer-aware influence

Optimizer-aware influence is a data attribution framework in machine learning that quantifies the impact of individual training examples on model parameters and predictions by explicitly incorporating the specific mathematical update rules and dynamics of the optimization algorithm used during training. While conventional influence functions generally rely on standard gradient descent approximations or assumptions of convergence under vanilla empirical risk minimization, optimizer-aware formulations adapt influence calculations to reflect the mechanics of advanced optimizers, such as adaptive learning rate scaling, momentum, and stateful gradient transformations like those found in Adam. By tracking how sample gradients translate into parameter changes under the actual optimization trajectory, this approach provides more faithful assessments of data relevance, which helps improve training data selection, model interpretability, and targeted fine-tuning across complex architectures.

1 item

LESS: Selecting Influential Data for Targeted Instruction Tuning

LESS: Selecting Influential Data for Targeted Instruction Tuning

Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, Danqi Chen

OrganizationsPrinceton UniversityUniversity of Washington

Why you should read this

Proposes LESS, an efficient gradient-based data selection method for targeted instruction tuning that outperforms full-dataset training using only 5% of the data and transfers effectively across model sizes and families.

Instruction tuning has unlocked powerful capabilities in large language models (LLMs), effectively using combined datasets to develop generalpurpose chatbots. However, real-world applications often require a specialized suite of skills (e.g., reasoning). The challenge lies in identifying the most relevant data from these extensive datasets to effectively develop specific capabilities, a setting we frame as targeted instruction tuning. We propose LESS, an optimizer-aware and practically efficient algorithm to effectively estimate data influences and perform Low-rank gradiEnt Similarity Search for instruction data selection. Crucially, LESS adapts existing influence formulations to work with the Adam optimizer and variable-length instruction data. LESS first constructs a highly reusable and transferable gradient datastore with low-dimensional gradient features and then selects examples based on their similarity to few-shot examples embodying a specific capability. Experiments show that training on a LESS-selected 5% of the data can often outperform training on the full dataset across diverse downstream tasks. Furthermore, the selected data is highly transferable: smaller models can be leveraged to select useful data for larger models and models from different families. Our qualitative analysis shows that our method goes beyond surface form cues to identify data that exemplifies the necessary reasoning skills for the intended downstream application.

Added

2026-09-28