Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks

Po-Nien KungFan YinDi WuKai-Wei ChangNanyun Peng

article2023EMNLP64 citations

Proposes an active instruction tuning framework that quantifies prompt uncertainty to identify and train on the most informative tasks, achieving superior cross-task generalization with substantially fewer training datasets.

Listen

Training large language models on diverse collections of tasks using clear instructions improves their ability to generalize to new, unseen tasks. However, as available task repositories expand to tens of thousands of datasets, training on all tasks simultaneously requires prohibitive computing resources and massive financial costs. Standard shortcuts such as random sampling often select uninformative tasks and yield suboptimal model performance, while existing active learning techniques focus on individual data instances rather than ranking complete tasks.

The article develops and evaluates an active instruction tuning framework designed to identify the most informative tasks for training large language models. Its main objective is to establish a task-level selection metric based on prompt sensitivity that maximizes model generalization using fewer training tasks and lower computational overhead.

To achieve this, the article introduces prompt uncertainty, a technique that assesses how sensitive a model's predictions are to slight perturbations in task instructions, such as randomly omitting 20% of the instruction words. The framework iteratively selects the tasks exhibiting the highest prompt uncertainty—representing tasks the model has not yet reliably mastered—and adds them to the training set. The researchers validated this approach across two major benchmarks: the Natural Instructions V2 dataset across five randomized seeds using a 770-million parameter model, and the 52,000-task Self-Instruct dataset using a 7-billion parameter language model. Evaluation was conducted through automated linguistic scoring and blind pairwise comparisons evaluated by human annotators and advanced external AI models.

The article established several critical findings. First, selecting tasks via prompt uncertainty consistently outperformed standard random sampling and traditional complexity-based baselines across both benchmarks. Second, when training on fewer than half of the available tasks, the proposed method achieved superior cross-task generalization, approaching the performance of models trained on entire massive pools with significantly less data. Third, using a newly introduced diagnostic tool called Task Map, the evaluation revealed that training on prompt-uncertain, ambiguous tasks drives almost all generalization gains, whereas training on prompt-certain difficult tasks offers no performance benefit. Fourth, controlled experiments demonstrated that training on related tasks reduced prompt uncertainty by a factor of 21 compared to unrelated tasks, proving the metric directly reflects task novelty.

These findings indicate that targeted task selection can substantially reduce training timelines and hardware compute expenses while preserving or improving output quality on unseen instructions. Selecting tasks indiscriminately wastes compute budgets on uninformative or excessively difficult data that does not improve cross-task capabilities. Practitioners and engineering teams should implement prompt-uncertainty filtering to curate training pools and use the diagnostic task mapping technique to audit data quality and prune unhelpful, overly rigid tasks.

Decision-makers should note that the underlying evaluations were performed on well-curated open-source benchmarks and medium-sized foundation models in controlled settings. The findings do not account for reinforcement learning from human feedback or extreme continuous learning scenarios with severe data noise. Consequently, organizations should pilot active instruction tuning on internal task distributions to calibrate instruction perturbation rates before executing large-scale production training runs.

arXiv: 2311.00288
Cover for Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks

Abstract

Instruction tuning (IT) achieves impressive zero-shot generalization results by training large language models (LLMs) on a massive amount of diverse tasks with instructions. However, how to select new tasks to improve the performance and generalizability of IT models remains an open question. Training on all existing tasks is impractical due to prohibiting computation requirements, and randomly selecting tasks can lead to suboptimal performance. In this work, we propose active instruction tuning based on prompt uncertainty, a novel framework to identify informative tasks, and then actively tune the models on the selected tasks. We represent the informativeness of new tasks with the disagreement of the current model outputs over perturbed prompts. Our experiments on NIV2 and Self-Instruct datasets demonstrate that our method consistently outperforms other baseline strategies for task selection, achieving better out-of-distribution generalization with fewer training tasks. Additionally, we introduce a task map that categorizes and diagnoses tasks based on prompt uncertainty and prediction probability. We discover that training on ambiguous (prompt-uncertain) tasks improves generalization while training on difficult (prompt-certain and low-probability) tasks offers no benefit, underscoring the importance of task selection for instruction tuning.¹

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Active Instruction Tuning
  • 2.2 Prompt Uncertainty
  • 3 Experiment Setting
  • 3.1 Active Instruction Tuning Setting
  • 3.2 Task Selection Strategies
  • 3.3 Training and Evaluation
  • 4 Results
  • 4.1 NIV2 Results
  • 4.2 Self-Instruct Results
  • 5 Task Map
  • 6 Discussion
  • 6.1 Prompt Uncertainty Reflects Task Novelty
  • 6.2 Prompt Perturbation Methods
  • 7 Related Work
  • 7.1 Instruction Tuning Paradigm
  • 7.2 Uncertainty Estimation for LLMs
  • 7.3 Active Learning and Task Selection
  • 8 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Evaluation details
  • A.2 Experiment details

Knowls

  1. Knowl 1 — Active Instruction Tuning framework

    model/method

    Active Instruction Tuning is an iterative task-selection framework for improving an instruction-tuned language model's generalization to arbitrary unseen tasks without assuming a particular target task. A small initial set of tasks is sampled from a larger task pool and used to train an initial instruction-tuned model. At each subsequent iteration, every remaining task is scored for informativeness using the current model, a batch of high-scoring tasks is added to the training set, and a new model is trained on the enlarged set. The method evaluates the resulting model on unseen tasks after each iteration, thereby trading the cost of scoring and training selected tasks against the cost of training on the entire task pool.

  2. Knowl 2 — Prompt Uncertainty task score

    equation

    Prompt Uncertainty measures how much an instruction-tuned model's likelihood for its own prediction changes when the task instruction is mildly perturbed. For task tt, let I0tI^t_0 be the original instruction, IjtI^t_j for j=1,…,kj=1,\ldots,k be kk perturbed versions, XtX_t be the task's unlabeled instance set, WW be the model parameters, and xitx^t_i be nn randomly sampled instances from XtX_t. The model generates an output yity^t_i for each xitx^t_i under the original instruction and keeps that same output fixed when evaluating all perturbed instructions. Let pi,jt=P(yit∣xit,Ijt,W)p^t_{i,j}=P(y^t_i\mid x^t_i,I^t_j,W) denote the model's sentence-level likelihood of that output. The task-level score is

    Ut=1n∑i=1n1k∑j=1k∣pi,0t−pi,jt∣.U_t=\frac{1}{n}\sum_{i=1}^{n}\frac{1}{k}\sum_{j=1}^{k}\left|p^t_{i,0}-p^t_{i,j}\right|.

    A larger UtU_t means that the model is more sensitive to instruction wording for task tt. Because the score uses model-generated outputs and unlabeled instances rather than gold answers, it can be computed before task annotation and used to select tasks for either training or later annotation.

  3. Knowl 3 — Operational task-selection procedure

    algorithm

    The Prompt Uncertainty selection procedure takes an instruction-tuned model, a pool of candidate tasks, and a current training-task set as input, and returns an expanded training-task set.

    Input: current model parameters WW, candidate task pool, current training-task set, selection batch size BB, sample count nn, perturbation count kk.

    Output: updated training-task set and a newly trained instruction-tuned model.

    1. For every candidate task, sample nn unlabeled instances.
    2. For each sampled instance, generate the model's output under the original instruction and record its sentence likelihood.
    3. Create kk perturbed instructions by randomly dropping words from the original instruction; compute the likelihood of the fixed original output under every perturbed instruction.
    4. Average the absolute likelihood differences over perturbations and instances to obtain the task's Prompt Uncertainty score.
    5. Select the BB highest-scoring tasks, remove them from the candidate pool, and add them to the training-task set.
    6. Train a new instruction-tuned model on the updated training-task set.
    7. Repeat steps 1–6 until the desired training-task budget is reached.

    In the paper's experiments, k=20k=20 perturbations were used. NIV2 used n=10n=10 instances and selected batches of B=68B=68 tasks; Self-Instruct used n=1n=1 instance because each task had only one instance. Each word in an instruction was independently dropped with probability 0.20.2.

  4. Knowl 4 — Prompt uncertainty as task novelty

    assumption

    The paper hypothesizes that Prompt Uncertainty is an estimate of model-side, or epistemic, uncertainty about a task. If a model cannot robustly map varied forms of an instruction to the latent concept that defines the task, its predictions change substantially under small instruction perturbations and the task receives a high Prompt Uncertainty score. Training on such tasks is hypothesized to improve the model's association between instructions and latent task concepts, thereby improving zero-shot performance on unseen instructions. This interpretation treats prompt perturbations as an ensemble of nearby conditions and differs from ordinary prediction confidence: a model can be confident in a particular output while still being unstable with respect to the task instruction.

  5. Knowl 5 — Task Map diagnosis tool

    definition

    Task Map is a model-based diagnostic representation that places each task along two dimensions: Prompt Uncertainty and Prediction Probability. Prediction Probability is the model's confidence, measured through the likelihood assigned to its generated task output, whereas Prompt Uncertainty measures the consistency of mapping the instruction to a latent task concept. The paper uses these dimensions to distinguish three qualitative task groups: Ambiguous tasks have high Prompt Uncertainty and are difficult for the model to recognize; Easy tasks have low Prompt Uncertainty and high Prediction Probability; Difficult tasks have low Prompt Uncertainty but low Prediction Probability. Thus, Ambiguous tasks are instruction-recognition problems, while Difficult tasks are tasks the model recognizes but cannot perform confidently.

  6. Knowl 6 — NIV2 active-tuning experiment

    experimental setup

    The NIV2 experiment used 756 English training tasks and 119 held-out test tasks containing both classification and generative tasks. For each of five random seeds, 68 tasks were randomly assigned to the initial training set, 68 different tasks to validation, and the remaining 620 tasks to the candidate pool. The candidate pool was expanded in batches of 68 tasks, producing training-set sizes of 68, 136, 204, 272, and 340; the appendix maintained a fixed batch composition of 24 classification and 44 generative tasks. Prompt Uncertainty was compared with random sampling, high perplexity, and low perplexity selection. Perplexity-based selection used predicted sentence perplexity for generation tasks and entropy for classification tasks, aggregating scores over multiple instances. Each model was a T5-770M instruction-tuned model trained with learning rate 2×10−52\times10^{-5}, batch size 128, 200 instances per task, and eight epochs. The model received a task definition and two demonstrations, and performance was measured by Rouge-L on overall, classification, and generative subsets of both validation and unseen test tasks.

  7. Knowl 7 — Prompt Uncertainty improves NIV2 generalization

    data/table

    Across five random seeds, selecting high-Prompt-Uncertainty tasks produced the strongest overall Rouge-L scores at early and intermediate training budgets on both the unseen NIV2 test set and the validation set. The table reports mean overall Rouge-L; each value is averaged over the five seeds.

    Could not parse LaTeX table

    The advantage is especially consistent for classification tasks. At 340 training tasks, Prompt Uncertainty reached 53.15 test Rouge-L for classification versus 51.59 for random sampling, 52.32 for high perplexity, and 51.36 for low perplexity. On generative test tasks at the same budget, Prompt Uncertainty reached 43.95, compared with 43.91 for random sampling, 43.09 for high perplexity, and 43.91 for low perplexity. Training on all 680 available tasks achieved 49.11 overall test Rouge-L and 48.65 overall validation Rouge-L, so Prompt Uncertainty obtained much of the full-training performance with fewer tasks. When more than roughly half of the candidate pool was selected, the pool contained fewer high-uncertainty tasks and the advantage diminished.

  8. Knowl 8 — Self-Instruct active-tuning results

    empirical result

    The Self-Instruct experiment used a 52K-task pool, randomly initialized training with 500 tasks, and evaluated models after expanding to 1,000, 2,000, 4,000, 8,000, and 16,000 tasks. The evaluation set contained 252 user-oriented tasks. Models trained with Prompt Uncertainty were compared against random sampling using blind pairwise judgments from GPT-4, ChatGPT, and human annotators; the reported score is the number of tasks won minus the number lost.

    Could not parse LaTeX table

    Human evaluation of Prompt Uncertainty against random sampling produced net scores of 13, 8, 11, 9, and −2-2 at the five training budgets. Prompt Uncertainty was preferred by all three evaluator types through 8,000 tasks, while its benefit largely disappeared at 16,000 tasks as the remaining candidate pool became smaller. High- and low-perplexity selection were generally no better than, and often worse than, random sampling.

  9. Knowl 9 — Task-category training effects

    empirical result

    Training-task choice, rather than merely training-task quantity, affected cross-task generalization on NIV2. Tasks categorized as Ambiguous by Task Map—high Prompt Uncertainty tasks—consistently improved overall test and validation performance more than random task selection during the early active-tuning iterations. Training on Easy tasks was generally worse than random selection, although adding more Easy tasks could still yield a small improvement. Training on Difficult tasks—low Prompt Uncertainty combined with low Prediction Probability—provided essentially no additional benefit as more tasks were added. The paper hypothesizes that Difficult tasks may be overly specific and too hard to learn, making them poor sources of transferable cross-task knowledge.

  10. Knowl 10 — Prompt Uncertainty tracks relevant-task novelty

    empirical result

    A controlled NIV2 experiment tested whether Prompt Uncertainty changes when a model learns related tasks. Model M0M_0 was evaluated on 620 irrelevant tasks and four unseen Word Analogy tasks. A second model M1M_1 was obtained by further training M0M_0 on four other Word Analogy tasks, after which the same 620 irrelevant tasks and four held-out analogy tasks were evaluated again. The average Prompt Uncertainty decrease was 0.0390.039 for the four related analogy tasks but only 0.00180.0018 for the irrelevant tasks, a 21-fold larger shift for the related tasks. Prediction Probability did not increase for the held-out analogy tasks after training on related analogy tasks. These results support Prompt Uncertainty as a detector of task novelty and show that prediction confidence alone does not capture whether a task is novel to the model.

  11. Knowl 11 — Scope limitations of the evaluation

    limitation

    The evidence for Active Instruction Tuning was obtained with open-source instruction-tuned models and did not test the effects of reinforcement learning from human feedback. The datasets were relatively well-constructed, so the experiments do not establish how Prompt Uncertainty behaves with noisy, duplicated, or severely malformed tasks; such settings may require noise filtering or batch active learning. The experiments also compared task-selection strategies in a controlled retraining loop rather than studying continual learning, so they do not measure forgetting or the benefits of incrementally updating one persistent model. Finally, the perturbation assumption is not universally safe: randomly dropping 20% of words was judged suitable for the relatively redundant NIV2 and Self-Instruct instructions, but concise instructions may require different rates and may have their meanings changed by perturbation.

Coverage note — The appendix's detailed annotation prompts, inter-annotator agreement statistics, cost accounting, and GPU-resource measurements were omitted because they support evaluation reproducibility rather than adding a separate methodological or scientific finding.

References

  1. 1.Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul Fieguth, Xiaochun Cao, Abbas Khosravi, U Rajendra Acharya, et al. 2021. A review of uncertainty quantification in deep learning: Techniques, applications and challenges. Information Fusion, 76:243–297.
  2. 2.Stephen H Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, et al. 2022. Promptsource: An integrated development environment and repository for natural language prompts. arXiv preprint arXiv:2202.01279.
  3. 3.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  4. 4.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models.
  5. 5.Reuben Feinman, Ryan R Curtin, Saurabh Shintre, and Andrew B Gardner. 2017. Detecting adversarial samples from artifacts. arXiv preprint arXiv:1703.00410.
  6. 6.Matthew Finlayson, Kyle Richardson, Ashish Sabharwal, and Peter Clark. 2022. What makes instruction learning hard? an investigation and a new challenge in a synthetic environment. arXiv preprint arXiv:2204.09148.
  7. 7.Yarin Gal and Zoubin Ghahramani. 2016. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In international conference on machine learning, pages 1050–1059. PMLR.
  8. 8.Prakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri, Maxine Eskenazi, and Jeffrey P Bigham. 2022. Improving zero and few-shot generalization in dialogue through instruction tuning. arXiv preprint arXiv:2205.12673.
  9. 9.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022. Unnatural instructions: Tuning language models with (almost) no human labor. arXiv preprint arXiv:2212.09689.
  10. 10.Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel. 2011. Bayesian active learning for classification and preference learning. arXiv preprint arXiv:1112.5745.
  11. 11.Hamish Ivison, Noah A Smith, Hannaneh Hajishirzi, and Pradeep Dasigi. 2022. Data-efficient finetuning using cross-task nearest neighbors. arXiv preprint arXiv:2212.00196.
  12. 12.Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2023. Exploring the benefits of training expert language models over instruction tuning. arXiv preprint arXiv:2302.03202.
  13. 13.Wenxiang Jiao, Jen tse Huang, Wenxuan Wang, Xing Wang, Shuming Shi, and Zhaopeng Tu. 2023. Parrot: Translating during chat using large language models. ArXiv.
  14. 14.Po-Nien Kung and Nanyun Peng. 2023. Do models really learn to follow instructions? an empirical study of instruction tuning. arXiv preprint arXiv:2305.11383.
  15. 15.Po-Nien Kung, Sheng-Siang Yin, Yi-Cheng Chen, Tse-Hsuan Yang, and Yun-Nung Chen. 2021. Efficient multi-task auxiliary learning: selecting auxiliary data by feature similarity. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 416–428.
  16. 16.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in neural information processing systems, 30.
  17. 17.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  18. 18.Zekun Li, Baolin Peng, Pengcheng He, and Xifeng Yan. 2023. Evaluating the instruction-following robustness of large language models to prompt injection.
  19. 19.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. 2023. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688.
  20. 20.Andrey Malinin and Mark Gales. 2020. Uncertainty estimation in autoregressive structured prediction. arXiv preprint arXiv:2002.07650.
  21. 21.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  22. 22.Fredrik Olsson. 2009. A literature survey of active machine learning in the context of natural language processing.
  23. 23.OpenAI. 2023. Gpt-4 technical report.
  24. 24.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  25. 25.Jane Pan, Tianyu Gao, Howard Chen, and Danqi Chen. 2023. What in-context learning" learns" in-context: Disentangling task recognition and task learning. arXiv preprint arXiv:2305.09731.
  26. 26.Md Rizwan Parvez and Kai-Wei Chang. 2021. Evaluating the values of sources in transfer learning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5084–5116, Online. Association for Computational Linguistics.
  27. 27.Clifton Poth, Jonas Pfeiffer, Andreas Rücklé, and Iryna Gurevych. 2021. What to pre-train on? efficient intermediate task selection. arXiv preprint arXiv:2104.08247.
  28. 28.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  29. 29.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  30. 30.Thomas Scialom, Tuhin Chakrabarty, and Smaranda Muresan. 2022. Fine-tuned language models are continual learners. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6107–6122, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  31. 31.Burr Settles. 2009. Active learning literature survey.
  32. 32.Yanyao Shen, Hyokun Yun, Zachary Lipton, Yakov Kronrod, and Animashree Anandkumar. 2017. Deep active learning for named entity recognition. In Proceedings of the 2nd Workshop on Representation Learning for NLP, pages 252–256, Vancouver, Canada. Association for Computational Linguistics.
  33. 33.Aditya Siddhant and Zachary C. Lipton. 2018. Deep Bayesian active learning for natural language processing: Results of a large-scale empirical study. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2904–2909, Brussels, Belgium. Association for Computational Linguistics.
  34. 34.Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9275–9293, Online. Association for Computational Linguistics.
  35. 35.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  36. 36.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  37. 37.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023. How far can camels go? exploring the state of instruction tuning on open resources.
  38. 38.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022a. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  39. 39.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, A. Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Maitreya Patel, Kuntal Kumar Pal, M. Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Shailaja Keyur Sampat, Savan Doshi, Siddharth Deepak Mishra, Sujan Reddy, Sumanta Patro, Tanay Dixit, Xudong Shen, Chitta Baral, Yejin Choi, Noah A. Smith, Hanna Hajishirzi, and Daniel Khashabi. 2022b. Supernaturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks.
  40. 40.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  41. 41.Yijun Xiao and William Yang Wang. 2019. Quantifying uncertainties in natural language processing tasks. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 7322–7329.
  42. 42.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2021. An explanation of in-context learning as implicit bayesian inference. arXiv preprint arXiv:2111.02080.
  43. 43.Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao, Yanggang Wang, Haiyu Li, and Zhilin Yang. 2022. Zero-prompt: Scaling prompt-based pretraining to 1,000 tasks improves zero-shot generalization. arXiv preprint arXiv:2201.06910.
  44. 44.Tianci Xue, Ziqi Wang, Yixia Li, Yun Chen, and Guanhua Chen. 2023. Tadis: Steering models for deep-thinking about demonstration examples.
  45. 45.Cheng-Fu Yang, Yen-Chun Chen, Jianwei Yang, Xiyang Dai, Lu Yuan, Yu-Chiang Frank Wang, and Kai-Wei Chang. 2023. Lacma: Language-aligning contrastive learning with meta-actions for embodied instruction following.
  46. 46.Da Yin, Xiao Liu, Fan Yin, Ming Zhong, Hritik Bansal, Jiawei Han, and Kai-Wei Chang. 2023a. Dynosaur: A dynamic growth paradigm for instruction-tuning data curation. arXiv preprint arXiv:2305.14327.
  47. 47.Fan Yin, Yao Li, Cho-Jui Hsieh, and Kai-Wei Chang. 2022. Addmu: Detection of far-boundary adversarial examples with data and model uncertainty estimation. In Conference on Empirical Methods in Natural Language Processing.
  48. 48.Fan Yin, Jesse Vig, Philippe Laban, Shafiq Joty, Caiming Xiong, and Chien-Sheng Jason Wu. 2023b. Did you read the instructions? rethinking the effectiveness of task definitions in instruction learning. arXiv preprint arXiv:2306.01150.
  49. 49.Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and Guoyin Wang. 2023. Instruction tuning for large language models: A survey.
  50. 50.Zhisong Zhang, Emma Strubell, and Eduard Hovy. 2022. A survey of active learning for natural language processing. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6166–6190, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  51. 51.Jing Zhou, Zongyu Lin, Yanan Zheng, Jian Li, and Zhilin Yang. 2023. Not all tasks are born equal: Understanding zero-shot generalization. In The Eleventh International Conference on Learning Representations.

Citation

MLA
Kung, P.-N., et al. “Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1813–29, https://doi.org/10.18653/v1/2023.emnlp-main.112.
APA
Kung, P.-N., Yin, F., Wu, D., Chang, K.-W., & Peng, N. (2023). Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1813–1829. https://doi.org/10.18653/v1/2023.emnlp-main.112
Chicago
Kung, P.-N., F. Yin, D. Wu, K.-W. Chang, and N. Peng. 2023. “Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1813–29. https://doi.org/10.18653/v1/2023.emnlp-main.112.
Harvard
Kung, P.-N. et al. (2023) “Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1813–1829. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.112.
Vancouver
1. Kung P-N, Yin F, Wu D, Chang K-W, Peng N (2023) Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1813–1829

BibTeX

@inproceedings{kung-etal-2023-active,
    title = "Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks",
    author = "Kung, Po-Nien  and
      Yin, Fan  and
      Wu, Di  and
      Chang, Kai-Wei  and
      Peng, Nanyun",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.112/",
    doi = "10.18653/v1/2023.emnlp-main.112",
    pages = "1813--1829"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/