Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks
Po-Nien KungFan YinDi WuKai-Wei ChangNanyun Peng
Proposes an active instruction tuning framework that quantifies prompt uncertainty to identify and train on the most informative tasks, achieving superior cross-task generalization with substantially fewer training datasets.
Training large language models on diverse collections of tasks using clear instructions improves their ability to generalize to new, unseen tasks. However, as available task repositories expand to tens of thousands of datasets, training on all tasks simultaneously requires prohibitive computing resources and massive financial costs. Standard shortcuts such as random sampling often select uninformative tasks and yield suboptimal model performance, while existing active learning techniques focus on individual data instances rather than ranking complete tasks.
The article develops and evaluates an active instruction tuning framework designed to identify the most informative tasks for training large language models. Its main objective is to establish a task-level selection metric based on prompt sensitivity that maximizes model generalization using fewer training tasks and lower computational overhead.
To achieve this, the article introduces prompt uncertainty, a technique that assesses how sensitive a model's predictions are to slight perturbations in task instructions, such as randomly omitting 20% of the instruction words. The framework iteratively selects the tasks exhibiting the highest prompt uncertainty—representing tasks the model has not yet reliably mastered—and adds them to the training set. The researchers validated this approach across two major benchmarks: the Natural Instructions V2 dataset across five randomized seeds using a 770-million parameter model, and the 52,000-task Self-Instruct dataset using a 7-billion parameter language model. Evaluation was conducted through automated linguistic scoring and blind pairwise comparisons evaluated by human annotators and advanced external AI models.
The article established several critical findings. First, selecting tasks via prompt uncertainty consistently outperformed standard random sampling and traditional complexity-based baselines across both benchmarks. Second, when training on fewer than half of the available tasks, the proposed method achieved superior cross-task generalization, approaching the performance of models trained on entire massive pools with significantly less data. Third, using a newly introduced diagnostic tool called Task Map, the evaluation revealed that training on prompt-uncertain, ambiguous tasks drives almost all generalization gains, whereas training on prompt-certain difficult tasks offers no performance benefit. Fourth, controlled experiments demonstrated that training on related tasks reduced prompt uncertainty by a factor of 21 compared to unrelated tasks, proving the metric directly reflects task novelty.
These findings indicate that targeted task selection can substantially reduce training timelines and hardware compute expenses while preserving or improving output quality on unseen instructions. Selecting tasks indiscriminately wastes compute budgets on uninformative or excessively difficult data that does not improve cross-task capabilities. Practitioners and engineering teams should implement prompt-uncertainty filtering to curate training pools and use the diagnostic task mapping technique to audit data quality and prune unhelpful, overly rigid tasks.
Decision-makers should note that the underlying evaluations were performed on well-curated open-source benchmarks and medium-sized foundation models in controlled settings. The findings do not account for reinforcement learning from human feedback or extreme continuous learning scenarios with severe data noise. Consequently, organizations should pilot active instruction tuning on internal task distributions to calibrate instruction perturbation rates before executing large-scale production training runs.
- Paper: Multitask Prompted Training Enables Zero-Shot Task Generalization, Victor Sanh et al. (2021). Introduces multitask prompted training on massive benchmark suites to enable zero-shot cross-task generalization, establishing the fundamental instruction fine-tuning baseline that active instruction tuning optimizes.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). Establishes instruction tuning across clustered datasets to drive zero-shot generalization, providing the foundational training paradigm evaluated and refined in the source.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). Introduces the 52,000-task synthetic instruction benchmark that serves as one of the primary evaluation environments for the source paper's active selection framework.
- Paper: Improving Generalization with Active Learning, David Cohn et al. (1994). Formalizes selective sampling and region-of-uncertainty active learning in neural networks, laying the foundational active learning theory that the source scales to task-level prompt selection.
- Paper: Active Learning with Statistical Models, David Cohn et al. (1996). Derives statistical variance-reduction criteria for querying the most informative data points, motivating the prompt-uncertainty selection principles utilized in the source.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). Analyzes language model sensitivity to minor prompt perturbations and formatting shifts, providing the foundational insights on prompt volatility that the source leverages to quantify task uncertainty.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Demonstrates automated evaluation and scoring of prompt variations using language models, informing the algorithmic perturbation mechanisms employed in the source.
- Paper: SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer, Tu Vu et al. (2022). Examines task transferability and cross-task generalization metrics across multi-task mixtures, offering theoretical background on inter-task relationships.
- Paper: DsDm: Model-Aware Dataset Selection with Datamodels, Logan Engstrom et al. (2024). Extends targeted subset selection by using model-aware datamodels to directly optimize target task performance, advancing beyond prompt uncertainty metrics for dataset curation.
- Paper: Scaling Instruction-Finetuned Language Models, Hyung Won Chung et al. (2024). Scales instruction tuning across massive task mixtures and model architectures, contextualizing the computational limits and generalization gains that task-selection methods seek to optimize.
- Paper: tinyBenchmarks: evaluating LLMs with fewer examples, Felipe Maia Polo et al. (2024). Applies efficient subset selection principles to model evaluation rather than training, drastically reducing the benchmark sample size required to assess multi-task ability.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). Analyzes the reliability of evaluating instruction-following capabilities, addressing key operational considerations for assessing the outputs generated by instruction-tuned models.
- Paper: Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision, Zhiqing Sun et al. (2024). Explores how models generalize from simpler supervision to difficult tasks, complementing the source's findings on the role of task difficulty and ambiguity in driving generalization.
