OPT: Open Pre-trained Transformer Language Models
Susan ZhangStephen RollerNaman GoyalMikel ArtetxeMoya ChenShuohui ChenChristopher DewanMona T. DiabXian LiXi Victoria Lin
Presents an openly available suite of pre-trained language models scaling up to 175 billion parameters that match GPT-3 performance at one-seventh the carbon footprint, granting researchers full access to model weights, training logs, and execution code.
Meta AI developed the Open Pre-trained Transformer (OPT) suite of decoder-only language models, ranging from 125 million to 175 billion parameters, to make large-scale models more widely available for study. Prior models such as GPT-3 demonstrated strong zero- and few-shot performance but remained closed or costly to replicate, restricting independent research on robustness, bias, and toxicity. The project therefore set out to produce an openly shared family of models that roughly matches GPT-3 performance while applying current best practices in data curation and training efficiency.
The models were trained on a deduplicated English corpus of roughly 180 billion tokens drawn from established sources including BookCorpus, CC-News, selected Pile subsets, and Pushshift Reddit. Training of the 175B model ran on 992 A100 GPUs using fully sharded data parallelism and tensor parallelism, achieving high utilization and requiring only about one-seventh the carbon footprint estimated for GPT-3. The team released all model weights up to 66B parameters, granted research access to the 175B model, and published training logs and the metaseq codebase.
Evaluations show that OPT-175B matches or approaches GPT-3 averages on 14 standard NLP benchmarks in zero-shot settings and remains competitive in one- and few-shot regimes, although results vary substantially by task. In unsupervised dialogue tests it performs close to supervised BlenderBot models on several metrics. Bias and toxicity measurements reveal broadly comparable profiles to GPT-3, with OPT-175B sometimes producing higher toxicity rates and stronger stereotypical associations on certain benchmarks, likely reflecting differences in training data composition.
These outcomes indicate that high-performing large language models can be developed and shared at materially lower environmental cost, enabling wider participation in research on their limitations and societal effects. The release lowers barriers for academic, civil-society, and industry researchers to examine scaling behavior, safety mitigations, and ethical considerations without duplicating massive compute investments.
The authors recommend restricting OPT-175B to non-commercial research use until further mitigations address toxicity, repetition, and factual hallucination. Additional work is needed on data characterization, more consistent evaluation protocols, and techniques such as retrieval augmentation or instruction tuning before broader deployment. Main limitations include reliance on ad-hoc training interventions, incomplete coverage of all possible harms by current benchmarks, and the absence of standardized carbon-accounting methods across the field; readers should therefore treat reported performance figures as directional rather than definitive.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). OPT directly builds upon and evaluates its open model suite by comparing its performance and architecture against the pioneering GPT-3 model.
- Paper: Scaling Laws for Neural Language Models, Jared Kaplan et al. (2020). The scaling laws established in this foundational work provide the theoretical basis for sizing and compute allocation used during the training of the OPT model suite.
- Paper: ZeRO: Memory optimizations Toward Training Trillion Parameter Models, Samyam Rajbhandari et al. (2020). OPT leverages ZeRO memory optimization techniques to enable the efficient distributed training of its largest 175-billion-parameter model.
- Paper: Efficient large-scale language model training on GPU clusters using megatron-LM, Deepak Narayanan et al. (2021). The Megatron-LM parallelization framework provides the foundational infrastructure and tensor-parallel methods essential for training OPT at scale.
- Paper: LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, Tim Dettmers et al. (2022). LLM.int8() builds directly upon models like OPT-175B to demonstrate 8-bit matrix multiplication for efficient inference without accuracy degradation.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). SmoothQuant extends the evaluation of post-training quantization methods to large models like OPT-175B, enabling efficient 8-bit weight and activation processing.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). GPTQ directly applies its advanced post-training quantization technique to compress massive transformer models like OPT-175B to low-bit representations.
- Paper: SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification, Xupeng Miao et al. (2024). SpecInfer uses OPT model variants as target and speculative models to accelerate generative serving performance through tree-based verification.
