OPT: Open Pre-trained Transformer Language Models

Susan ZhangStephen RollerNaman GoyalMikel ArtetxeMoya ChenShuohui ChenChristopher DewanMona T. DiabXian LiXi Victoria Lin

article2022arXiv4,957 citations

Presents an openly available suite of pre-trained language models scaling up to 175 billion parameters that match GPT-3 performance at one-seventh the carbon footprint, granting researchers full access to model weights, training logs, and execution code.

Listen

Meta AI developed the Open Pre-trained Transformer (OPT) suite of decoder-only language models, ranging from 125 million to 175 billion parameters, to make large-scale models more widely available for study. Prior models such as GPT-3 demonstrated strong zero- and few-shot performance but remained closed or costly to replicate, restricting independent research on robustness, bias, and toxicity. The project therefore set out to produce an openly shared family of models that roughly matches GPT-3 performance while applying current best practices in data curation and training efficiency.

The models were trained on a deduplicated English corpus of roughly 180 billion tokens drawn from established sources including BookCorpus, CC-News, selected Pile subsets, and Pushshift Reddit. Training of the 175B model ran on 992 A100 GPUs using fully sharded data parallelism and tensor parallelism, achieving high utilization and requiring only about one-seventh the carbon footprint estimated for GPT-3. The team released all model weights up to 66B parameters, granted research access to the 175B model, and published training logs and the metaseq codebase.

Evaluations show that OPT-175B matches or approaches GPT-3 averages on 14 standard NLP benchmarks in zero-shot settings and remains competitive in one- and few-shot regimes, although results vary substantially by task. In unsupervised dialogue tests it performs close to supervised BlenderBot models on several metrics. Bias and toxicity measurements reveal broadly comparable profiles to GPT-3, with OPT-175B sometimes producing higher toxicity rates and stronger stereotypical associations on certain benchmarks, likely reflecting differences in training data composition.

These outcomes indicate that high-performing large language models can be developed and shared at materially lower environmental cost, enabling wider participation in research on their limitations and societal effects. The release lowers barriers for academic, civil-society, and industry researchers to examine scaling behavior, safety mitigations, and ethical considerations without duplicating massive compute investments.

The authors recommend restricting OPT-175B to non-commercial research use until further mitigations address toxicity, repetition, and factual hallucination. Additional work is needed on data characterization, more consistent evaluation protocols, and techniques such as retrieval augmentation or instruction tuning before broader deployment. Main limitations include reliance on ad-hoc training interventions, incomplete coverage of all possible harms by current benchmarks, and the absence of standardized carbon-accounting methods across the field; readers should therefore treat reported performance figures as directional rather than definitive.

Cover for OPT: Open Pre-trained Transformer Language Models

Abstract

Large language models, which are often trained for hundreds of thousands of compute days, have shown remarkable capabilities for zero- and few-shot learning. Given their computational cost, these models are difficult to replicate without significant capital. For the few that are available through APIs, no access is granted to the full model weights, making them difficult to study. We present Open Pre-trained Transformers (OPT), a suite of decoder-only pre-trained transformers ranging from 125M to 175B parameters, which we aim to fully and responsibly share with interested researchers. We show that OPT-175B is comparable to GPT-3, while requiring only 1/7th the carbon footprint to develop. We are also releasing our logbook detailing the infrastructure challenges we faced, along with code for experimenting with all of the released models.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Models
  • 2.2 Training Setup
  • 2.3 Pre-training Corpus
  • 2.4 Training Efficiency
  • 2.5 Training Processes
  • 3 Evaluations
  • 3.1 Prompting & Few-Shot
  • 3.2 Dialogue
  • 4 Bias & Toxicity Evaluations
  • 4.1 Hate Speech Detection
  • 4.2 CrowS-Pairs
  • 4.3 StereoSet
  • 4.4 RealToxicityPrompts
  • 4.5 Dialogue Safety Evaluations
  • 5 Limitations
  • 6 Considerations for Release
  • 7 Related Work
  • 8 Conclusion
  • References
  • A Additional Evaluations
  • B Contributions
  • C Datasheet
  • C.1 Motivation
  • C.2 Composition
  • C.3 Collection Process
  • C.4 Preprocessing/cleaning/labeling
  • C.5 Uses
  • C.6 Distribution
  • C.7 Maintenance
  • D Model Card
  • D.1 Model Details
  • D.2 Intended Use
  • D.3 Data, Limitations, and Recommendations
  • E Sample Model Outputs

Knowls

  1. Knowl 1 — OPT Model Architecture Suite Specifications

    data/table

    The Open Pre-trained Transformer (OPT) family is a collection of decoder-only autoregressive language models ranging from 125 million to 175 billion parameters. The architectural parameters across the nine models in the suite are configured as follows:

    Model Layers (LL) Attention Heads (HH) Embedding Size (dmodeld_{\text{model}}) Peak LR Global Batch Size (tokens)
    125M 12 12 768 6.0×10−46.0 \times 10^{-4} 0.5M
    350M 24 16 1024 3.0×10−43.0 \times 10^{-4} 0.5M
    1.3B 24 32 2048 2.0×10−42.0 \times 10^{-4} 1.0M
    2.7B 32 32 2560 1.6×10−41.6 \times 10^{-4} 1.0M
    6.7B 32 32 4096 1.2×10−41.2 \times 10^{-4} 2.0M
    13B 40 40 5120 1.0×10−41.0 \times 10^{-4} 4.0M
    30B 48 56 7168 1.0×10−41.0 \times 10^{-4} 4.0M
    66B 64 72 9216 0.8×10−40.8 \times 10^{-4} 2.0M
    175B 96 96 12288 1.2×10−41.2 \times 10^{-4} 2.0M

    All models utilize standard Transformer decoder blocks with ReLU activations and evaluate inputs up to a fixed maximum sequence length of 2048 tokens. Batch sizes remain constant throughout the training duration of each model.

  2. Knowl 2 — OPT Pre-Training Configuration and Hyperparameters

    model/method

    The OPT models are initialized and trained using the following optimization setup:

    1. Weight Initialization: Model weights are sampled from a zero-mean normal distribution with standard deviation σ=0.006\sigma = 0.006. For output projection layers, the standard deviation is scaled by a factor of 1.02L\frac{1.0}{\sqrt{2L}}, where LL is the total number of layers. All bias parameters are initialized to 0.

    2. Optimization: Optimization is performed using AdamW with exponential decay rates (β1,β2)=(0.9,0.95)(\beta_1, \beta_2) = (0.9, 0.95) and weight decay set to 0.10.1. The learning rate follows a linear schedule, warming up from 0 to the model's peak learning rate over the first 2000 steps for OPT-175B (or over 375M tokens for models ≤66B\le 66\text{B}), before decaying down to 10%10\% of the peak learning rate over a total of 300B tokens.

    3. Regularization and Gradient Handling: A dropout rate of 0.10.1 is applied throughout intermediate layers, while embeddings use zero dropout (0.00.0). Gradient norm clipping is set at a threshold of 1.01.0 (adjusted down to 0.30.3 during periods of instability). To prevent numerical over/underflow during distributed all-reduce communication across NN parallel workers, a gradient predivide factor is used, dividing intermediate gradients twice by N\sqrt{N} rather than once by NN.

  3. Knowl 3 — OPT Pre-Training Corpus Composition and Preprocessing

    experimental setup

    The OPT pre-training corpus contains approximately 180 billion tokens (roughly 800 GB of text), tokenized using the GPT-2 byte-level Byte-Pair Encoding (BPE) tokenizer. The corpus is formed by concatenating three primary dataset sources:

    1. RoBERTa Corpora: Comprises BookCorpus (over 10,000 unpublished books), Stories (filtered CommonCrawl data matching Winograd story style), and CC-News v2 (an updated English CommonCrawl news crawl containing articles up to September 28, 2021).

    2. The Pile Subsets: A filtered selection from the Pile, specifically containing CommonCrawl (Pile-CC), OpenWebText2, USPTO, Project Gutenberg, OpenSubtitles, Wikipedia, DM Mathematics, and HackerNews. Subsets that exhibited high gradient norm spikes during early 1.3B exploratory runs were excluded.

    3. Pushshift.io Reddit: Conversational threads from Reddit. Documents were generated by extracting only the single longest comment chain from each dialogue tree while pruning all branching paths, reducing the original Reddit corpus size by approximately 66%66\%.

    Cross-dataset deduplication is performed across all sources using MinHash LSH with a Jaccard similarity threshold of ≥0.95\ge 0.95.

  4. Knowl 4 — Distributed Training Architecture and Efficiency of OPT-175B

    model/method

    OPT-175B was trained on 992 80GB NVIDIA A100 GPUs using the open-source metaseq training framework. Distributed scaling is achieved through a combination of Fully Sharded Data Parallel (FSDP) and Megatron-LM Tensor Parallelism without employing pipeline parallelism.

    Optimizer states for AdamW are stored and updated in FP32 format while being fully sharded across hosts, whereas model parameter weights are maintained in FP16 precision. Dynamic loss scaling is applied to prevent gradient underflow during half-precision backpropagation. The resulting execution achieves hardware utilization of up to 147 TFLOP/s per A100 GPU.

    The pre-training run of OPT-175B generated an estimated carbon footprint of 75 metric tons CO2eq\text{CO}_2\text{eq}, compared to an estimated 500 tons for GPT-3 and 380 tons for Gopher.

  5. Knowl 5 — Stability Interventions and Failure Recovery Protocol for OPT-175B Training

    model/method

    Pre-training OPT-175B across large GPU clusters encountered both infrastructure and numerical stability hurdles that required runtime interventions:

    • Hardware Faults: Over a two-month pre-training run, cluster hardware failures led to at least 35 manual restarts, the decommissioning/cycling of over 100 host nodes, and an estimated 70+ automatic restarts from saved checkpoints.
    • Loss Divergence Recovery: Numerical divergence events correlated with the dynamic loss scale factor collapsing to zero and the l2l_2-norm of final-layer activations spiking unboundedly. To recover from divergences, training was rolled back to an earlier checkpoint where the dynamic loss scalar was still in a healthy state (≥1.0\ge 1.0), combined with an immediate reduction in the peak learning rate.
    • Additional Stabilizations: Lowering the gradient norm clipping ceiling from 1.01.0 to 0.30.3, periodically resetting the dynamic loss scalar, and updating the tensor parallelism backend reduced activation norm pressure and stabilized optimization.
  6. Knowl 6 — Prompting and In-Context Few-Shot NLP Benchmark Performance

    empirical result

    Across 16 standard NLP benchmark tasks (including HellaSwag, StoryCloze, PIQA, ARC Easy, ARC Challenge, OpenBookQA, WinoGrad, Winogrande, and SuperGLUE subtasks BoolQ, CB, COPA, RTE, WSC, MultiRC, ReCoRD, WIC), OPT-175B demonstrates zero-shot, one-shot, and few-shot (32-shot) performance that broadly tracks the scaling curves of GPT-3.

    Key task-level dynamics include:

    1. Parity on Standard Tasks: On 10 of the standard benchmarks (e.g., HellaSwag, StoryCloze, PIQA, ARC Easy, OpenBookQA, Winogrande), OPT-175B matches GPT-3 performance within expected variance across scales.
    2. Underperformance: OPT-175B lags GPT-3 on ARC Challenge and consistently underperforms GPT-3 Davinci evaluations on MultiRC across zero-, one-, and few-shot configurations.
    3. Small Validation Set Instability: On tasks with small validation sets (CB with 56 examples, WSC with 104 examples, and BoolQ with 277 examples), both OPT and GPT model families show erratic, non-monotonic scaling behavior, hovering close to majority-class baselines.
  7. Knowl 7 — Zero-Shot Dialogue Performance of Unsupervised OPT-175B

    data/table

    OPT-175B was evaluated without fine-tuning on five multi-turn dialogue benchmarks: ConvAI2 (C2), Wizard of Wikipedia (WW), Empathetic Dialogues (ED), Blended Skill Talk (BST), and Wizard of Internet (WoI). All perplexity scores are normalized to the GPT-2 tokenizer space, and generation uses greedy decoding up to 32 tokens prompted solely with alternating "Person 1:" and "Person 2:" prefixes.

    Perplexity (↓\downarrow) Unigram F1 (↑\uparrow)
    Model Supervision C2 WW ED BST WoI C2 WW ED BST WoI
    Reddit 2.7B Unsupervised 18.9 21.0 11.6 17.4 18.0 0.126 0.133 0.135 0.133 0.124
    BlenderBot 1 Supervised 10.2 12.5 9.0 11.9 14.7 0.183 0.189 0.192 0.178 0.154
    R2C2 BlenderBot Supervised 10.5 12.4 9.1 11.7 14.6 0.205 0.198 0.197 0.186 0.160
    OPT-175B Unsupervised 10.8 13.3 10.3 12.1 12.0 0.185 0.152 0.149 0.162 0.147

    Unsupervised OPT-175B significantly outperforms the unsupervised Reddit 2.7B baseline across all benchmarks, achieves lower perplexity (12.0) on Wizard of Internet than supervised models, and matches the Unigram F1 performance of fully supervised BlenderBot 1 on ConvAI2 (0.185 vs 0.183).

  8. Knowl 8 — Hate Speech Detection Accuracy on ETHOS

    data/table

    Hate speech classification was evaluated on the ETHOS benchmark across zero-shot, one-shot, few-shot binary (classifying text as racist or sexist vs. neither), and few-shot multiclass settings. Scores reflect overall F1 performance comparing GPT-3 Davinci and OPT-175B:

    Prompting Configuration GPT-3 Davinci OPT-175B
    Zero-shot 0.628 0.667
    One-shot 0.616 0.713
    Few-shot (binary) 0.354 0.759
    Few-shot (multiclass) 0.672 0.812

    OPT-175B outperforms GPT-3 Davinci across all evaluation setups, which is attributed to stronger representation of social media conversational data (Pushshift.io Reddit) in the OPT pre-training corpus and potential API moderation artifacts in Davinci.

  9. Knowl 9 — Stereotype and Social Bias Evaluation on CrowS-Pairs and StereoSet

    data/table

    Social and demographic bias was assessed on CrowS-Pairs (measuring intrasentence stereotype preference; lower score indicates less bias) and StereoSet (measuring Language Modeling Score [LMS, ↑\uparrow], Stereotype Score [SS, ↓\downarrow toward 50], and Idealized Context Association Test [ICAT, ↑\uparrow]):

    CrowS-Pairs StereoSet
    Category GPT-3 OPT-175B Category Metric GPT-3 Davinci OPT-175B
    Gender 62.6 65.7 Profession LMS (↑\uparrow) 78.4 74.1
    Religion 73.3 68.6 SS (↓\downarrow) 63.4 62.6
    Race/Color 64.7 68.6 ICAT (↑\uparrow) 57.5 55.4
    Sexual orientation 76.2 78.6 Gender LMS (↑\uparrow) 75.6 74.0
    Age 64.4 67.8 SS (↓\downarrow) 66.5 63.6
    Nationality 61.6 62.9 ICAT (↑\uparrow) 50.6 53.8
    Disability 76.7 76.7 Religion LMS (↑\uparrow) 80.8 84.0
    Physical appearance 74.6 76.2 SS (↓\downarrow) 59.0 59.0
    Socioeconomic status 73.8 76.2 ICAT (↑\uparrow) 66.3 68.9
    Overall 67.2 69.5 Race LMS (↑\uparrow) 77.0 74.9
    SS (↓\downarrow) 57.4 56.8
    ICAT (↑\uparrow) 65.7 64.8
    Overall LMS (↑\uparrow) 77.6 74.8
    SS (↓\downarrow) 60.8 59.9
    ICAT (↑\uparrow) 60.8 60.0

    On CrowS-Pairs, OPT-175B exhibits higher stereotype bias overall (69.5 vs. 67.2) and in 7 of 9 subcategories relative to GPT-3. On StereoSet, OPT-175B and GPT-3 Davinci show nearly identical overall ICAT composite scores (60.0 vs. 60.8), with OPT-175B yielding marginally better (lower) Stereotype Scores (59.9 vs. 60.8) but lower language modeling scores.

  10. Knowl 10 — Toxicity Generation and Conversational Safety Unit Test Profiles

    data/table

    Toxicity and dialogue safety were measured on RealToxicityPrompts (generating 25 samples of 20 tokens per prompt with nucleus sampling p=0.9p=0.9), SaferDialogues (evaluating recovery after failure via PPL and F1), and Safety Bench Unit Tests (measuring unsafe response rates across Safe, Realistic, Unsafe, and Adversarial settings):

    SaferDialogues Safety Unit Tests (% unsafe ↓\downarrow)
    Model PPL (↓\downarrow) F1 (↑\uparrow) Safe Realistic Unsafe Adversarial
    Reddit 2.7B 16.2 0.140 0.300 0.261 0.450 0.439
    BlenderBot 1 12.4 0.161 0.028 0.150 0.250 0.194
    R2C2 BlenderBot 13.8 0.160 0.022 0.133 0.289 0.222
    OPT-175B 14.7 0.141 0.033 0.261 0.567 0.283

    On RealToxicityPrompts, OPT-175B generates higher average toxicity across all prompt toxicity tiers than either PaLM or GPT-3 Davinci. On Safety Bench Unit Tests, unsupervised OPT-175B achieves low failure rates in the Safe setting (0.033) but has a high failure rate in the Unsafe prompt setting (0.567), significantly exceeding models fine-tuned on curated dialogue data.

  11. Knowl 11 — Behavioral Limitations and Generation Vulnerabilities in OPT-175B

    limitation

    Qualitative evaluation of OPT-175B identifies several recurring operational limitations:

    1. Instruction Execution Failure: OPT-175B performs poorly when prompted with direct declarative instructions or point-blank interrogatives, tending to simulate a continuing dialogue transcript that opens with the prompt rather than executing the requested instruction.
    2. Repetitive Looping: Unsupervised generation frequently falls into repetitive token cycles and loops, a degeneration failure mode that nucleus sampling reduces but does not completely eliminate for single generations.
    3. Factual Hallucination: The model produces confident but factually incorrect assertions, particularly in technical, mathematical, or multi-step reasoning tasks such as non-additive arithmetic or programming token variations.
    4. Toxicity and Stereotyping Susceptibility: Pre-training on unmoderated social web corpora increases susceptibility to eliciting toxic continuations and reinforcing social stereotypes when given adversarial or open-ended prompts.

Coverage note — Omitted qualitative code and poetry text samples from the appendix (Figures 8-13) and administrative release documentation (datasheet and model card administrative details) as these duplicate the quantitative benchmark findings and methodological specifications covered in the extracted knowls.

References

  1. 1.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977.
  2. 2.Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giri Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Mona T. Diab, Zornitsa Kozareva, and Ves Stoyanov. 2021. Efficient large scale language modeling with mixtures of experts. CoRR, abs/2112.10684.
  3. 3.Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. The pushshift reddit dataset. CoRR, abs/2001.08435.
  4. 4.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  5. 5.Yonatan Bisk, Rowan Zellers, Ronan Le bras, Jianfeng Gao, and Yejin Choi. 2020. Piqa: Reasoning about physical commonsense in natural language. Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432–7439.
  6. 6.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. Gpt-neox-20b: An open-source autoregressive language model.
  7. 7.Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1004–1015, Online. Association for Computational Linguistics.
  8. 8.Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie S. Chen, Kathleen Creel, Jared Quincy Davis, Dorottya Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, and et al. 2021. On the opportunities and risks of foundation models. CoRR, abs/2108.07258.
  9. 9.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2021. Improving language models by retrieving from trillions of tokens. arXiv preprint arXiv:2112.04426.
  10. 10.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  11. 11.Nathanael Chambers and Dan Jurafsky. 2008. Unsupervised learning of narrative event chains. In Proceedings of ACL-08: HLT, pages 789–797, Columbus, Ohio. Association for Computational Linguistics.
  12. 12.Ke-Li Chiu and Rohan Alexander. 2021. Detecting hate speech with gpt-3. arXiv preprint arXiv:2103.12407.
  13. 13.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  14. 14.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the AI2 reasoning challenge. CoRR, abs/1803.05457.
  15. 15.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019. Plug and play language models: A simple approach to controlled text generation. arXiv preprint arXiv:1912.02164.
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In North American Association for Computational Linguistics (NAACL).
  17. 17.Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. Bold: Dataset and metrics for measuring biases in open-ended language generation. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 862–872.
  18. 18.Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2021. Anticipating safety issues in e2e conversational ai: Framework and tooling. arXiv preprint arXiv:2107.03451.
  19. 19.Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2020a. Queens are powerful too: Mitigating gender bias in dialogue generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8173–8188, Online. Association for Computational Linguistics.
  20. 20.Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019a. Build it break it fix it for dialogue safety: Robustness from adversarial human attack. arXiv preprint arXiv:1908.06083.
  21. 21.Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, Shrimai Prabhumoye, Alan W. Black, Alexander Rudnicky, Jason Williams, Joelle Pineau, Mikhail Burtsev, and Jason Weston. 2020b. The second conversational intelligence challenge (ConvAI2). In The NeurIPS '18 Competition, pages 187–208, Cham. Springer International Publishing.
  22. 22.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019b. Wizard of Wikipedia: Knowledge-powered conversational agents. In Proceedings of the International Conference on Learning Representations.
  23. 23.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2021a. The pile: An 800gb dataset of diverse text for language modeling. CoRR, abs/2101.00027.
  24. 24.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021b. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 3816–3830. Association for Computational Linguistics.
  25. 25.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daume III, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM, 64(12):86–92.
  26. 26.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online. Association for Computational Linguistics.
  27. 27.Udit Gupta, Young Geun Kim, Sylvia Lee, Jordan Tse, Hsien-Hsin S Lee, Gu-Yeon Wei, David Brooks, and Carole-Jean Wu. 2021. Chasing carbon: The elusive environmental footprint of computing. IEEE International Symposium on High-Performance Computer Architecture (HPCA 2021).
  28. 28.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778.
  29. 29.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022. Training compute-optimal large language models.
  30. 30.Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. ArXiv, abs/1904.09751.
  31. 31.Abigail Z. Jacobs and Hanna Wallach. 2021. Measurement and fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT '21, page 375–385, New York, NY, USA. Association for Computing Machinery.
  32. 32.Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving. 2021. Alignment of language agents. CoRR, abs/2103.14659.
  33. 33.Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2021. Internet-augmented dialogue generation. CoRR, abs/2107.07566.
  34. 34.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2020. GEDI: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367.
  35. 35.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. CoRR, abs/2104.08691.
  36. 36.Hector J Levesque, Ernest Davis, and Leora Morgenstern. 2011. The Winograd schema challenge. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning, volume 46, page 47.
  37. 37.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen-tau Yih, Tim Rocktaschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  38. 38.Xiang Lisa Li and Percy Liang. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. pages 4582–4597.
  39. 39.Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021. Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning, pages 6565–6576. PMLR.
  40. 40.Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham. 2021. Jurassic-1: Technical details and evaluation. Technical report, AI21 Labs.
  41. 41.Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2019a. Does gender matter? towards fairness in dialogue systems. arXiv preprint arXiv:1910.10486.
  42. 42.Haokun Liu, William Huang, Dhara Mungra, and Samuel R. Bowman. 2020. Precise task formalization matters in Winograd schema evaluations. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8275–8280, Online. Association for Computational Linguistics.
  43. 43.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021. What makes good in-context examples for gpt-3? CoRR, abs/2101.06804.
  44. 44.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  45. 45.Ilya Loshchilov and Frank Hutter. 2017. Fixing weight decay regularization in adam. CoRR, abs/1711.05101.
  46. 46.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity.
  47. 47.Clara Meister, Tim Vieira, and Ryan Cotterell. 2020. Best-first beam search. Transactions of the Association for Computational Linguistics, 8:795–809.
  48. 48.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. 2017. Mixed precision training. arXiv preprint arXiv:1710.03740.
  49. 49.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? A new dataset for open book question answering. CoRR, abs/1809.02789.
  50. 50.Tomas Mikolov, Jiri Kopecky, Lukas Burget, Ondrej Glembek, et al. 2009. Neural network based language models for highly inflective languages. In 2009 IEEE international conference on acoustics, speech and signal processing, pages 4725–4728. IEEE.
  51. 51.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2021. Metaicl: Learning to learn in context.
  52. 52.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  53. 53.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2018. Model cards for model reporting. CoRR, abs/1810.03993.
  54. 54.Ioannis Mollas, Zoe Chrysopoulou, Stamatis Karlos, and Grigorios Tsoumakas. 2020. ETHOS: an online hate speech detection dataset. CoRR, abs/2006.08328.
  55. 55.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James F. Allen. 2016. A corpus and evaluation framework for deeper understanding of commonsense stories. CoRR, abs/1604.01696.
  56. 56.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pretrained language models. In Association for Computational Linguistics (ACL).
  57. 57.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  58. 58.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman. 2020. Crows-pairs: A challenge dataset for measuring social biases in masked language models. arXiv preprint arXiv:2010.00133.
  59. 59.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2022. A conversational paradigm for program synthesis. arXiv preprint.
  60. 60.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  61. 61.David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350.
  62. 62.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. Advances in Neural Information Processing Systems, 34.
  63. 63.Fabio Petroni, Tim Rocktaschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  64. 64.Alec Radford, Karthik Narasimhan, Time Salimans, and Ilya Sutskever. 2018. Improving language understanding with unsupervised learning. Technical report, OpenAI.
  65. 65.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. Technical report, OpenAI.
  66. 66.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William S. Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs/2112.11446.
  67. 67.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research (JMLR), 21:1–67.
  68. 68.Anand Rajaraman and Jeffrey David Ullman. 2011. Mining of massive datasets. Cambridge University Press.
  69. 69.Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019. Towards empathetic open-domain conversation models: A new benchmark and dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5370–5381, Florence, Italy. Association for Computational Linguistics.
  70. 70.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325, Online. Association for Computational Linguistics.
  71. 71.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Winogrande: An adversarial winograd schema challenge at scale. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 8732–8740. AAAI Press.
  72. 72.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021. Multitask prompted training enables zero-shot task generalization.
  73. 73.Teven Le Scao and Alexander M. Rush. 2021. How many data points is a prompt worth? pages 2627–2636.
  74. 74.Timo Schick and Hinrich Schutze. 2020. It's not just size that matters: Small language models are also few-shot learners. CoRR, abs/2009.07118.
  75. 75.Timo Schick, Sahana Udupa, and Hinrich Schutze. 2021. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp. Transactions of the Association for Computational Linguistics, 9:1408–1424.
  76. 76.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  77. 77.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation. arXiv preprint arXiv:1909.01326.
  78. 78.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. pages 4222–4235.
  79. 79.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053.
  80. 80.Kurt Shuster, Mojtaba Komeili, Leonard Adolphs, Stephen Roller, Arthur Szlam, and Jason Weston. 2022. Language models that seek for knowledge: Modular search & generation for dialogue and prompt completion. arXiv preprint arXiv:2203.13224.
  81. 81.Eric Smith, Mary Williamson, Kurt Shuster, Jason Weston, and Y-Lan Boureau. 2020. Can you put it all together: Evaluating conversational agents' ability to blend skills. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. ACL.
  82. 82.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zheng, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro. 2022. Using deepspeed and megatron to train megatron-turing NLG 530b, A large-scale generative language model. CoRR, abs/2201.11990.
  83. 83.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  84. 84.Trieu H. Trinh and Quoc V. Le. 2018. A simple method for commonsense reasoning. CoRR, abs/1806.02847.
  85. 85.Megan Ung, Jing Xu, and Y-Lan Boureau. 2021. Saferdialogues: Taking feedback gracefully after conversational safety failures. ArXiv, abs/2110.07518.
  86. 86.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems.
  87. 87.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. SuperGLUE: A stickier benchmark for general-purpose language understanding systems. arXiv preprint 1905.00537.
  88. 88.Albert Webson and Ellie Pavlick. 2021. Do prompt-based models really understand the meaning of their prompts? arXiv preprint arXiv:2109.01247.
  89. 89.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. Finetuned language models are zero-shot learners. CoRR, abs/2109.01652.
  90. 90.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2021a. Ethical and social risks of harm from language models.
  91. 91.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021b. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359.
  92. 92.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020. Neural text generation with unlikelihood training. In International Conference on Learning Representations.
  93. 93.Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga Behram, James Huang, Charles Bai, Michael Gschwind, Anurag Gupta, Myle Ott, Anastasia Melnikov, Salvatore Candido, David Brooks, Geeta Chauhan, Benjamin Lee, Hsien-Hsin S. Lee, Bugra Akyildiz, Maximilian Balandat, Joe Spisak, Ravi Jain, Mike Rabbat, and Kim Hazelwood. 2022. Sustainable AI: environmental implications, challenges and opportunities. In Proceedings of the Conference on Machine Learning and Systems.
  94. 94.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020. Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079.
  95. 95.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021a. Bot-adversarial dialogue for safe conversational agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2950–2968, Online. Association for Computational Linguistics.
  96. 96.Jing Xu, Arthur Szlam, and Jason Weston. 2021b. Beyond goldfish memory: Long-term open-domain conversation. arXiv preprint arXiv:2107.07567.
  97. 97.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 4791–4800. Association for Computational Linguistics.
  98. 98.Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. CoRR, abs/1506.06724.

Citation

MLA
Zhang, S., et al. “OPT: Open Pre-trained Transformer Language Models”. arXiv, 2022, http://arxiv.org/abs/2205.01068v4.
APA
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., & Zettlemoyer, L. (2022). OPT: Open Pre-trained Transformer Language Models. arXiv. http://arxiv.org/abs/2205.01068v4
Chicago
Zhang, S., S. Roller, N. Goyal, et al. 2022. “OPT: Open Pre-trained Transformer Language Models”. arXiv. http://arxiv.org/abs/2205.01068v4.
Harvard
Zhang, S. et al. (2022) “OPT: Open Pre-trained Transformer Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.01068v4.
Vancouver
1. Zhang S, Roller S, Goyal N, et al (2022) OPT: Open Pre-trained Transformer Language Models. arXiv

BibTeX

@article{zhang2022opt,
  title = {OPT: Open Pre-trained Transformer Language Models},
  author = {Zhang, Susan and Roller, Stephen and Goyal, Naman and Artetxe, Mikel and Chen, Moya and Chen, Shuohui and Dewan, Christopher and Diab, Mona and Li, Xian and Lin, Xi Victoria and Mihaylov, Todor and Ott, Myle and Shleifer, Sam and Shuster, Kurt and Simig, Daniel and Koura, Punit Singh and Sridhar, Anjali and Wang, Tianlu and Zettlemoyer, Luke},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.01068v4},
  eprint = {2205.01068}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/