Symbol tuning improves in-context learning in language models

Jerry W. WeiLe HouAndrew K. LampinenXiangning ChenDa HuangYi TayXinyun ChenYifeng LuDenny ZhouTengyu Ma

article2023EMNLP108 citations

Demonstrates that finetuning language models on input-label pairs with arbitrary symbols forces them to genuinely reason over in-context mappings, leading to substantial gains on algorithmic tasks, underspecified prompts, and counterfactual examples where models must override prior knowledge.

Listen

Large language models frequently struggle to adapt reliably when given few-shot examples in prompts, often relying heavily on rigid instructions and pre-existing semantic knowledge rather than genuinely learning the task from the provided context. When prompts lack clear descriptions or natural language labels, these systems often fail to deduce the intended task. The article introduces and evaluates "symbol tuning," a finetuning method designed to force models to learn input–label mappings directly from in-context examples by replacing natural language labels with arbitrary symbols (such as random words, character combinations, or integers) and omitting instructions.

To evaluate this method, the authors finetuned instruction-tuned models across four sizes (from 8 billion to 540 billion parameters) on a balanced mixture of 22 classification datasets mapped to approximately 30,000 arbitrary symbols. They evaluated the models across 11 unseen natural language processing datasets in varying prompt formats, as well as on separate algorithmic reasoning benchmarks and prompts with intentionally flipped labels to test adaptability.

The findings show that symbol tuning substantially enhances in-context learning. First, for models with 62 billion parameters or more, the method improved performance across all prompt settings, delivering the largest gains—between 5.5% and 15.5%—when prompts lacked instructions and relevant labels. Notably, symbol-tuned 62-billion-parameter models matched or outperformed standard 540-billion-parameter models in these underspecified settings, effectively reducing the compute needed at inference time by roughly tenfold. Second, despite being trained solely on natural language text, symbol-tuned models demonstrated dramatic improvements on algorithmic reasoning benchmarks, gaining up to 18.2% on list manipulation tasks and 15.3% on Turing concept tasks. Third, symbol tuning restored the ability to follow flipped labels (such as inverted sentiment definitions), improving accuracy by 26.5% to 34.0% over standard instruction-tuned models, which routinely fail when in-context examples contradict their prior knowledge.

These results demonstrate that symbol tuning compels models to reason dynamically over exemplars rather than relying passively on memorized concepts. In practice, this reduces prompt brittleness and allows organizations to deploy smaller, more cost-effective models without sacrificing reasoning performance. The approach is directly actionable for teams developing or deploying large language models on complex reasoning pipelines, though further work is needed to explore scaling the training mixture beyond 22 datasets and adapting symbol tuning to open-ended text generation tasks. Overall, confidence in these findings is high for classification and algorithmic mappings across the tested model family, though practitioners should exercise caution when applying the technique to the smallest model sizes (such as 8 billion parameters), where symbol tuning caused minor performance drops on standard tasks due to overfitting.

No sufficiently relevant recommendations were found.

Cover for Symbol tuning improves in-context learning in language models

Abstract

We present symbol tuning—finetuning language models on in-context input–label pairs where natural language labels (e.g., “positive/negative sentiment”) are replaced with arbitrary symbols (e.g., “foo/bar”). Symbol tuning leverages the intuition that when a model cannot use instructions or natural language labels to figure out a task, it must instead do so by learning the input–label mappings.

We experiment with symbol tuning across PaLM models up to 540B parameters and observe benefits across various settings. First, symbol tuning boosts performance on unseen in-context learning tasks and is much more robust to underspecified prompts, such as those without instructions or without natural language labels. Second, symbol-tuned models are much stronger at algorithmic reasoning tasks, with up to 18.2% better performance on the List Functions benchmark and up to 15.3% better performance on the Simple Turing Concepts benchmark. Finally, symbol-tuned models show large improvements in following flipped-labels presented in-context, meaning that they are more capable of using in-context information to override prior knowledge.

Table of Contents

  • 1 Introduction
  • 2 Symbol tuning
  • 3 Experimental setup
  • 3.1 Tuning tasks & prompt formatting
  • 3.2 Evaluation tasks
  • 3.3 Models & finetuning procedure
  • 4 Symbol-tuned models are better in-context learners
  • 5 Symbol tuning improves algorithmic reasoning
  • 6 Symbol-tuned models can override priors via flipped labels
  • 7 Related work
  • 7.1 In-context learning via semantic prior knowledge
  • 7.2 In-context learning via in-context exemplars
  • 7.3 Tuning language models
  • 8 Limitations
  • 9 Conclusions
  • References

Knowls

  1. Knowl 1 — Symbol tuning trains models to infer tasks from arbitrary input–label mappings

    model/method

    Symbol tuning finetunes a language model on in-context classification prompts in which task instructions are removed and the original class labels are replaced by arbitrary, task-unrelated symbols. A prompt contains labeled examples and a query; because neither the instruction nor the symbol’s meaning identifies the task, the model must infer the input-to-label mapping from the examples. The intended effect is to strengthen use of in-context exemplars, rather than reliance on natural-language task descriptions or semantically meaningful labels.

  2. Knowl 2 — Tuning prompts use 22 classification datasets and multiple exemplars per class

    experimental setup

    The symbol-tuning mixture contains training examples from 22 publicly available NLP classification datasets spanning seven task groups: natural-language inference (RTE, WNLI, QNLI, MNLI, SNLI, CB); topic classification (AGN, TREC); miscellaneous classification (TEO, TEI, WIC, COLA); commonsense classification (COPA, PIQA); coreference (WSC, WINO); paraphrase detection (QQP, MRPC, PAWS); and sentiment analysis (RT, SST2, TES). For each dataset, prompts are built from its training split using a randomly selected input–label format and a randomly selected 2–10 in-context exemplars per class. The examples’ original labels are remapped to arbitrary symbols, and the tuning prompts do not provide task instructions.

  3. Knowl 3 — Tuning and evaluation use separate arbitrary-symbol inventories

    experimental setup

    The arbitrary labels come from approximately 300,000 symbols in three forms: integers, character combinations, and words. Approximately 30,000 symbols are used for finetuning: 1–4 digit integers, 1–3 letter combinations, and words from a 10,000-word list. Approximately 270,000 symbols are reserved for evaluation and excluded from the tuning inventory: 5-digit integers, 3–4 letter combinations, and words from a 100,000-word list. This separation tests use of symbols not encountered during finetuning.

  4. Knowl 4 — Symbol-tuning finetuning recipe for Flan-PaLM

    experimental setup

    Experiments finetune instruction-tuned PaLM models at 8B, 62B, and 540B parameters, as well as Flan-PaLM-62B-cont, a 62B model trained on 1.3T rather than 780B tokens. The training mixture samples among the 22 datasets and caps each dataset at 25,000 randomly selected training examples. Examples are packed into sequences, with inputs separated from labels by an end-of-sequence token. All models use batch size 32 and Adafactor; the learning rate is 3×10−33\times10^{-3} for 8B and 62B models and 1×10−31\times10^{-3} for the 540B model. Tuning uses input and target sequence lengths of 2,048 and 512 tokens. Reported checkpoints are after 4,000 steps for the 8B and 62B models (including 62B-cont) and 1,000 steps for the 540B model.

  5. Knowl 5 — Evaluation tests unseen tasks under four levels of prompt information

    experimental setup

    The in-context-learning evaluation uses the validation splits of 11 NLP datasets excluded from both symbol tuning and the original instruction tuning: SUBJ, TEH, TEAB, TEAT, TEFE, TEHI, ADEC, OR, SOT, TOS, and TC. At most 100 examples per dataset are evaluated. Prompts use a randomly selected input–label format and four in-context exemplars per class. The evaluation crosses two binary factors: whether task instructions are present and whether the exemplars use relevant natural-language labels. When relevant labels are absent, the original labels are remapped to randomly selected symbols from the held-out evaluation inventory. Accuracy is averaged across the 11 tasks.

  6. Knowl 6 — Symbol tuning improves unseen-task accuracy, especially without relevant labels

    data/table

    The table reports average accuracy (%) across 11 unseen classification tasks. Each cell gives the baseline Flan-PaLM result and, beneath it, the result after symbol tuning; the parenthesized change is the paper’s reported change from that baseline. Conditions are distinguished by availability of relevant natural-language labels and task instructions. The largest and most consistent gains occur when relevant labels are unavailable; the 8B model is an exception in the two conditions where relevant labels are available, where its accuracy falls.

    ModelRelevant labels ✓, instructions ✓Relevant labels ✓, instructions ✗Relevant labels ✗, instructions ✓Relevant labels ✗, instructions ✗
    Random guessing42.442.442.442.4
    Flan-PaLM-8B63.9 → 57.6 (−6.3)61.6 → 54.3 (−7.3)42.4 → 58.2 (+15.8)44.2 → 52.8 (+8.6)
    Flan-PaLM-62B74.3 → 75.5 (+1.2)70.0 → 70.8 (+0.8)57.0 → 71.4 (+14.4)50.5 → 60.3 (+9.8)
    Flan-PaLM-62B-cont77.3 → 78.9 (+1.6)70.3 → 74.5 (+4.2)56.3 → 71.8 (+15.5)51.0 → 62.1 (+11.1)
    Flan-PaLM-540B82.2 → 84.4 (+2.2)77.4 → 78.8 (+1.4)70.7 → 80.0 (+9.3)58.1 → 63.6 (+5.5)

    For models of 62B parameters and above, symbol tuning improves accuracy in all four conditions: gains are +0.8 to +4.2 when relevant labels are available and +5.5 to +15.5 when they are not. In the hardest condition, with neither instructions nor relevant labels, symbol-tuned 8B exceeds untuned 62B, and symbol-tuned 62B exceeds untuned 540B.

  7. Knowl 7 — Symbol tuning transfers to algorithmic reasoning tasks

    empirical result

    Although the finetuning mixture contains natural-language classification data and no numerical or algorithmic training tasks, symbol tuning improves performance on two unseen BIG-Bench task families. On the 20 list-function tasks with the highest human-accuracy baselines, grouped into five categories, the reported average accuracy gains are +18.2% for Flan-PaLM-8B, +11.1% for 62B, +15.5% for 62B-cont, and +3.6% for 540B. List-function evaluation is four-shot. On the Simple Turing Concepts tasks in the AS II subset with three or fewer instructions, where models infer a transformation of binary strings, gains are +15.3% for 8B, +15.3% for 62B, +14.1% for 62B-cont, and +4.7% for 540B. The paper reports that symbol-tuned 62B-cont exceeds untuned 540B on average across the list-function tasks; the authors do not directly compare these results with human baselines because the evaluation formats differ.

  8. Knowl 8 — Symbol tuning restores some ability to follow flipped in-context labels

    empirical result

    To test whether models can prioritize in-context mappings over prior knowledge, the evaluation flips the labels of both the exemplars and the evaluation examples. Tasks with more than two labels are excluded, and a model that follows the flipped mapping should score above random guessing. Instruction-tuned Flan-PaLM models generally score well below random guessing, whereas symbol-tuned models are substantially more capable of following the flipped labels. The reported average improvements over instruction-tuned models are 26.5% for 8B, 33.7% for 62B, and 34.0% for 540B. On some datasets, including OR, SUBJ, and TC, symbol-tuned models score substantially above random guessing; averaged across datasets, however, performance is not significantly above random guessing. Symbol-tuned models match or exceed the pretraining-only PaLM models in average performance, indicating a partial recovery of the flipped-label ability lost during instruction tuning.

  9. Knowl 9 — Evidence is limited to classification tuning on instruction-tuned PaLM models

    limitation

    The study remaps discrete classification labels and does not establish how symbol tuning should be applied to open-ended generation tasks. Experiments use 22 tuning datasets and one instruction-tuned model family, Flan-PaLM; they do not determine whether the effects generalize to other architectures, pretraining objectives, training procedures, or models without instruction tuning. The authors note that scaling the number of tuning tasks may change effectiveness but do not test beyond the 22 datasets. They also report that symbol tuning did not significantly affect chain-of-thought reasoning in their analysis, possibly because the tuning data contained no chain-of-thought prompts.

Coverage note — No substantial contributed material was deliberately omitted; the knowls cover the method, training and evaluation protocols, principal empirical findings, and stated limitations.

References

  1. 1.Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. 2023. What learning algorithm is in-context learning? Investigations with linear models. In International Conference on Learning Representations.
  2. 2.Neel Alex, Eli Lifland, Lewis Tunstall, Abhishek Thakur, Pegah Maham, C. Jess Riedel, Emmie Hine, Carolyn Ashurst, Paul Sedille, Alexis Carlier, Michael Noetel, and Andreas Stuhlmüller. 2021. RAFT: A real-world few-shot text classification benchmark. In Conference on Neural Information Processing Systems.
  3. 3.Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. 2019. SemEval-2019 Task 5: Multilingual detection of hate speech against immigrants and women in Twitter. In International Workshop on Semantic Evaluation.
  4. 4.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. PIQA: Reasoning about physical commonsense in natural language. In Conference of the Association for the Advancement of Artificial Intelligence.
  5. 5.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Conference on Empirical Methods in Natural Language Processing.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Conference on Neural Information Processing Systems.
  7. 7.Zihang Chen, Hongbo Zhang, Xiaoji Zhang, and Leqi Zhao. 2017. Quora question pairs.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, et al. 2022. PaLM: Scaling language modeling with Pathways.
  9. 9.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models.
  10. 10.Alexis Conneau and Douwe Kiela. 2018. SentEval: An evaluation toolkit for universal sentence representations. In Conference on Language Resources and Evaluation.
  11. 11.Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei. 2023. Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers. In Workshop on Understanding Foundation Models at the International Conference on Learning Representations.
  12. 12.Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant. 2022. What can transformers learn in-context? A case study of simple function classes. In Conference on Neural Information Processing Systems.
  13. 13.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In International Conference on Learning Representations.
  14. 14.Sakaguchi Keisuke, Le Bras Ronan, Bhagavatula Chandra, and Choi Yejin. 2021. WinoGrande: An adversarial winograd schema challenge at scale. Communications of the Association for Computing Machinery.
  15. 15.Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In International Conference on the Principles of Knowledge Representation and Reasoning.
  16. 16.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. Solving quantitative reasoning problems with language models. In Conference on Neural Information Processing Systems.
  17. 17.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A community library for natural language processing. In Conference on Empirical Methods in Natural Language Processing: System Demonstrations.
  18. 18.Xin Li and Dan Roth. 2002. Learning question classifiers. In Conference on Computational Linguistics.
  19. 19.Aman Madaan and Amir Yazdanbakhsh. 2022. Text and patterns: For effective chain of thought, it takes two to tango.
  20. 20.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022a. MetaICL: Learning to learn in context. In Conference of the North American Chapter of the Association for Computational Linguistics.
  21. 21.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022b. Rethinking the role of demonstrations: What makes in-context learning work? In Conference on Empirical Methods in Natural Language Processing.
  22. 22.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the Association for Computational Linguistics.
  23. 23.Saif Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu, and Colin Cherry. 2016. SemEval-2016 Task 6: Detecting stance in tweets. In International Workshop on Semantic Evaluation.
  24. 24.Allen Newell. 1980. Physical symbol systems. Cognitive Science.
  25. 25.Allen Newell and Herbert A. Simon. 1976. Computer science as empirical inquiry: Symbols and search. In Communications of the Association for Computing Machinery.
  26. 26.OpenAI. 2023. GPT-4 technical report.
  27. 27.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Conference on Neural Information Processing Systems.
  28. 28.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the Association for Computational Linguistics.
  29. 29.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research.
  30. 30.Aniruddh Raghu, Maithra Raghu, Samy Bengio, and Oriol Vinyals. 2020. Rapid learning or feature reuse? towards understanding the effectiveness of MAML. In International Conference on Learning Representations.
  31. 31.Janarthanan Rajendran, Alexander Irpan, and Eric Jang. 2020. Meta-learning requires meta-augmentation. In Conference on Neural Information Processing Systems.
  32. 32.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Conference on Empirical Methods in Natural Language Processing.
  33. 33.Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the Conference on Human Factors in Computing Systems.
  34. 34.Sara Rosenthal, Noura Farra, and Preslav Nakov. 2017. SemEval-2017 Task 4: Sentiment analysis in twitter. In International Workshop on Semantic Evaluation.
  35. 35.Joshua S. Rule, Joshua B. Tenenbaum, and Steven T. Piantadosi. 2020. The child as hacker. Trends in Cognitive Sciences.
  36. 36.Joshua Stewart Rule. 2020. The child as hacker: building more human-like models of learning. Ph.D. thesis, Massachusetts Institute of Technology.
  37. 37.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Tali Bers, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  38. 38.Adam Santoro, Andrew K. Lampinen, Kory W. Mathewson, Timothy P. Lillicrap, and David Raposo. 2021. Symbolic behaviour in artificial intelligence.
  39. 39.Noam Shazeer and Mitchell Stern. 2018. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning.
  40. 40.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Conference on Empirical Methods in Natural Language Processing.
  41. 41.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.
  42. 42.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. 2022. Challenging BIG-Bench tasks and whether chain-of-thought can solve them.
  43. 43.Jan Arne Telle, José Hernández-Orallo, and Cèsar Ferri. 2019. The teaching size: computable teachers and learners for universal languages. Machine Learning.
  44. 44.Cynthia Van Hee, Els Lefever, and Véronique Hoste. 2018. SemEval-2018 Task 3: Irony detection in english tweets. In International Workshop on Semantic Evaluation.
  45. 45.Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. 2022. Transformers learn in-context by gradient descent.
  46. 46.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. SuperGLUE: A stickier benchmark for general-purpose language understanding systems. In Conference on Neural Information Processing Systems.
  47. 47.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In BlackboxNLP Workshop at the Conference on Empirical Methods in Natural Language Processing.
  48. 48.Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun. 2022. Towards understanding chain-of-thought prompting: An empirical study of what matters.
  49. 49.Albert Webson and Ellie Pavlick. 2022. Do prompt-based models really understand the meaning of their prompts? In Conference of the North American Chapter of the Association for Computational Linguistics.
  50. 50.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022a. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  51. 51.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. In Conference on Neural Information Processing Systems.
  52. 52.Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma. 2023. Larger language models do in-context learning differently.
  53. 53.Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021. CrossFit: A few-shot learning challenge for cross-task generalization in NLP. In Conference on Empirical Methods in Natural Language Processing.
  54. 54.Mingzhang Yin, George Tucker, Mingyuan Zhou, Sergey Levine, and Chelsea Finn. 2020. Meta-learning without memorization. In International Conference on Learning Representations.
  55. 55.Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019. SemEval-2019 Task 6: Identifying and categorizing offensive language in social media (offenseval). In International Workshop on Semantic Evaluation.
  56. 56.Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Conference on Neural Information Processing Systems.
  57. 57.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase Adversaries from Word Scrambling. In Proceedings of the North American Chapter of the Association for Computational Linguistics.

Citation

MLA
Wei, J., et al. “Symbol Tuning Improves In-context Learning in Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 968–79, https://doi.org/10.18653/v1/2023.emnlp-main.61.
APA
Wei, J., Hou, L., Lampinen, A., Chen, X., Huang, D., Tay, Y., Chen, X., Lu, Y., Zhou, D., Ma, T., & Le, Q. (2023). Symbol tuning improves in-context learning in language models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 968–979. https://doi.org/10.18653/v1/2023.emnlp-main.61
Chicago
Wei, J., L. Hou, A. Lampinen, et al. 2023. “Symbol Tuning Improves In-context Learning in Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 968–79. https://doi.org/10.18653/v1/2023.emnlp-main.61.
Harvard
Wei, J. et al. (2023) “Symbol tuning improves in-context learning in language models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 968–979. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.61.
Vancouver
1. Wei J, Hou L, Lampinen A, et al (2023) Symbol tuning improves in-context learning in language models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 968–979

BibTeX

@inproceedings{wei-etal-2023-symbol,
    title = "Symbol tuning improves in-context learning in language models",
    author = "Wei, Jerry  and
      Hou, Le  and
      Lampinen, Andrew  and
      Chen, Xiangning  and
      Huang, Da  and
      Tay, Yi  and
      Chen, Xinyun  and
      Lu, Yifeng  and
      Zhou, Denny  and
      Ma, Tengyu  and
      Le, Quoc",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.61/",
    doi = "10.18653/v1/2023.emnlp-main.61",
    pages = "968--979"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/