BloombergGPT: A Large Language Model for Finance

Shijie WuOzan IrsoySteven LuVadim DabravolskiMark DredzeSebastian GehrmannPrabhanjan KambadurDavid RosenbergGideon Mann

article2023arXiv1,508 citations

Introduces BloombergGPT, a 50-billion-parameter model trained on over 700 billion financial and general tokens that demonstrates how domain-specific pretraining significantly improves financial task performance while maintaining general language capabilities.

Listen

Financial technology increasingly relies on natural language processing for critical operations such as sentiment analysis, market news classification, and information extraction. While general-purpose large language models have demonstrated powerful few-shot learning capabilities, they often struggle with the nuanced terminology and specialized formats typical of financial data. At the same time, existing domain-specific models have predominantly focused on smaller architectures or isolated domains like biomedicine. The objective of the article is to demonstrate the design, training, and evaluation of BloombergGPT, a 50-billion-parameter language model purpose-built for the financial industry that delivers best-in-class financial capability without sacrificing general-purpose performance.

To achieve this, the authors constructed a 709-billion-token training corpus by combining a massive, 363-billion-token proprietary financial dataset called FinPile—spanning news, company filings, press releases, and financial web documents collected over decades—with 345 billion tokens from high-quality public text sources. Using 512 compute accelerators, the team trained a 50.6-billion-parameter decoder-only transformer model on 569 billion tokens, guided by compute-optimal scaling principles and using an expressive custom tokenizer. The resulting model was evaluated across standard public benchmarks and proprietary internal financial benchmarks against established open models such as GPT-NeoX, OPT-66B, and BLOOM-176B.

Across the evaluations, BloombergGPT demonstrated substantial performance advantages. On financial sentiment analysis and question-answering benchmarks, it outperformed all comparable models by significant margins, including achieving a 100% win rate across internal aspect-specific sentiment tasks and a 43.41% accuracy on conversational financial reasoning compared to 27.88%–36.31% for alternative models. In exploratory joint entity extraction and ticker disambiguation tasks, the model scored 64.83 average F1, surpassing the second-best model by over 6 points. Crucially, on general-purpose NLP suites—including reading comprehension, linguistic tasks, and challenging reasoning benchmarks—BloombergGPT remained highly competitive, often matching or exceeding the performance of similarly sized or much larger general-purpose models.

These findings prove that training on a balanced mix of specialized internal data and broad public text avoids the degradation of general reasoning abilities while unlocking superior domain-specific performance. For organizations, this approach delivers lower operational risk and higher accuracy for sensitive workflows like automated query generation, market intelligence, and document analysis. Rather than choosing between brittle specialized models and less accurate general models, enterprises can leverage domain-adapted foundation models for multi-task applications.

Looking forward, the authors recommend exploring task-specific fine-tuning and alignment techniques to further tailor domain-specific models to complex business processes. Organizations should also evaluate the effect of cleaner, curated datasets on mitigating toxic language and model bias. Finally, while BloombergGPT exhibits strong empirical results, the authors note that the proprietary model weights cannot be publicly shared due to data security and privacy constraints associated with commercial financial archives, meaning future external adopters should focus on applying these mixed-dataset pretraining practices within their own secure environments.

arXiv: 2303.17564
Cover for BloombergGPT: A Large Language Model for Finance

Abstract

The use of NLP in the realm of financial technology is broad and complex, with applications ranging from sentiment analysis and named entity recognition to question answering. Large Language Models (LLMs) have been shown to be effective on a variety of tasks; however, no LLM specialized for the financial domain has been reported in literature. In this work, we present BloombergGPT, a 50 billion parameter language model that is trained on a wide range of financial data. We construct a 363 billion token dataset based on Bloomberg's extensive data sources, perhaps the largest domain-specific dataset yet, augmented with 345 billion tokens from general purpose datasets. We validate BloombergGPT on standard LLM benchmarks, open financial benchmarks, and a suite of internal benchmarks that most accurately reflect our intended usage. Our mixed dataset training leads to a model that outperforms existing models on financial tasks by significant margins without sacrificing performance on general LLM benchmarks. Additionally, we explain our modeling choices, training process, and evaluation methodology. We release Training Chronicles (Appendix C) detailing our experience in training BloombergGPT.

Table of Contents

  • 1 Introduction
  • 1.1 BloombergGPT
  • 1.2 Broader Contributions
  • 2 Dataset
  • 2.1 Financial Datasets (363B tokens – 51.27% of training)
  • 2.1.1 Web (298B tokens – 42.01% of training)
  • 2.1.2 News (38B tokens – 5.31% of training)
  • 2.1.3 Filings (14B tokens – 2.04% of training)
  • 2.1.4 Press (9B tokens – 1.21% of training)
  • 2.1.5 Bloomberg (5B tokens – 0.70% of training)
  • 2.2 Public Datasets (345B tokens – 48.73% of training)
  • 2.2.1 The Pile (184B tokens – 25.9% of training)
  • 2.2.2 C4 (138B tokens – 19.48% of training)
  • 2.2.3 Wikipedia (24B tokens – 3.35% of training)
  • 2.3 Tokenization
  • 3 Model
  • 3.1 Architecture
  • 3.2 Model Scaling
  • 3.3 Training Configuration
  • 3.4 Large-scale Optimization
  • 4 Training Run
  • 5 Evaluation
  • 5.1 Few-shot Methodology
  • 5.2 Heldout Loss
  • 5.3 Financial Tasks
  • 5.3.1 External Financial Tasks
  • 5.3.2 Internal Task: Sentiment Analysis
  • 5.3.3 Exploratory Task: NER
  • 5.4 BIG-bench Hard
  • 5.5 Knowledge Assessments
  • 5.6 Reading Comprehension
  • 5.7 Linguistic Tasks
  • 5.8 Summary
  • 6 Qualitative Samples
  • 7 Related Work
  • 8 Ethics, Limitations, and Implications
  • 8.1 Ethical Use
  • 8.2 Openness
  • 9 Conclusion
  • References
  • A Architecture
  • A.0 Notation
  • A.1 Full Architecture
  • A.2 SelfAttention with ALiBi (SA\mathop{\mathrm{SA}}\nolimits)
  • A.3 LayerNorm (LN\mathop{\mathrm{LN}}\nolimits)
  • A.4 FeedForwardNetwork (FFN\mathop{\mathrm{FFN}}\nolimits)
  • A.5 List of All Trainable Parameters
  • B Details on external financial tasks
  • C Training Chronicles
  • C.0 Still
  • C.1 Elbow
  • C.2 Slide
  • C.3 Suspense

Knowls

  1. Knowl 1 — BloombergGPT Architecture and Parameter Specification

    model/method

    BloombergGPT is a 50.6-billion-parameter (50,558,868,480 trainable parameters) causal, decoder-only transformer language model based on the BLOOM architecture with ALiBi positional biases and embedding layer normalization.

    Model Dimensions and Structural Parameters

    • Number of transformer decoder layers: L=70L = 70
    • Number of attention heads per layer: N=40N = 40
    • Head dimension: Dn=192D^n = 192 for each head n∈{1,…,N}n \in \{1, \dots, N\}
    • Hidden state dimension: D=N⋅Dn=7,680D = N \cdot D^n = 7{,}680
    • Feed-forward intermediate hidden dimension: D′=4D=30,720D' = 4D = 30{,}720
    • Vocabulary size: ∣V∣=131,072|V| = 131{,}072 (2172^{17})
    • Training sequence length: T=2,048T = 2{,}048

    Forward Pass Equations

    Given an input sequence of token indices (x1,…,xT)(x_1, \dots, x_T) with xt∈{1,…,∣V∣}x_t \in \{1, \dots, |V|\} represented by standard basis vectors ext∈R∣V∣e_{x_t} \in \mathbb{R}^{|V|}, the initial input representation matrix H0=[h10,…,hT0]∈RD×TH^0 = [h^0_1, \dots, h^0_T] \in \mathbb{R}^{D \times T} is computed via tied token embedding matrix Wem∈RD×∣V∣W^{\text{em}} \in \mathbb{R}^{D \times |V|} followed by an initial embedding layer normalization LNem\text{LN}^{\text{em}}: hˉt0=Wemext,ht0=LNem(hˉt0)∀t∈{1,…,T}\bar{h}^0_t = W^{\text{em}} e_{x_t}, \quad h^0_t = \text{LN}^{\text{em}}(\bar{h}^0_t) \quad \forall t \in \{1, \dots, T\}

    For layers ℓ=1,…,L\ell = 1, \dots, L, layer activations Hℓ∈RD×TH^\ell \in \mathbb{R}^{D \times T} are updated sequentially via pre-layer-normalized self-attention SAℓ\text{SA}_\ell and feed-forward sub-layers FFNℓ\text{FFN}_\ell: Hˉℓ=Hℓ−1+SAℓ(LNℓin(Hℓ−1))\bar{H}^\ell = H^{\ell-1} + \text{SA}_\ell(\text{LN}^{\text{in}}_\ell(H^{\ell-1})) Hℓ=Hˉℓ+FFNℓ(LNℓat(Hˉℓ))H^\ell = \bar{H}^\ell + \text{FFN}_\ell(\text{LN}^{\text{at}}_\ell(\bar{H}^\ell)) where LNℓin\text{LN}^{\text{in}}_\ell and LNℓat\text{LN}^{\text{at}}_\ell have distinct trainable gain γ∈RD\gamma \in \mathbb{R}^D and bias β∈RD\beta \in \mathbb{R}^D parameters.

    Self-Attention with ALiBi

    For layer input X∈RD×TX \in \mathbb{R}^{D \times T} and head n∈{1,…,N}n \in \{1, \dots, N\}, query, key, and value matrices are projected via Wℓn,q,Wℓn,k,Wℓn,v∈RDn×DW_\ell^{n,q}, W_\ell^{n,k}, W_\ell^{n,v} \in \mathbb{R}^{D^n \times D} and biases bℓn,q,bℓn,k,bℓn,v∈RDnb_\ell^{n,q}, b_\ell^{n,k}, b_\ell^{n,v} \in \mathbb{R}^{D^n}: Qn=Wℓn,qX+bℓn,q,Kn=Wℓn,kX+bℓn,k,Vn=Wℓn,vX+bℓn,vQ^n = W_\ell^{n,q} X + b_\ell^{n,q}, \quad K^n = W_\ell^{n,k} X + b_\ell^{n,k}, \quad V^n = W_\ell^{n,v} X + b_\ell^{n,v}

    ALiBi adds fixed, non-learned linear bias matrix An∈RT×TA^n \in \mathbb{R}^{T \times T} directly to the attention logits. Setting N~=2⌊log⁡2N⌋\tilde{N} = 2^{\lfloor \log_2 N \rfloor} and n~=1+((n−1) mod N~)−0.5⌊(n−1)/N~⌋\tilde{n} = 1 + ((n-1) \bmod \tilde{N}) - 0.5 \lfloor (n-1)/\tilde{N} \rfloor, the ALiBi bias elements are: ai,jn=2−8Nn~(i−j)⋅I(i<j)∀i,j∈{1,…,T}a^n_{i,j} = 2^{-\frac{8}{N}\tilde{n}} (i - j) \cdot \mathbb{I}(i < j) \quad \forall i, j \in \{1, \dots, T\}

    With causal mask M∈RT×TM \in \mathbb{R}^{T \times T} (mi,j=0m_{i,j} = 0 for i≤ji \le j and −∞-\infty for i>ji > j): Sˉn=An+(Kn)⊤QnDn,Sn=droppat(softmax(Sˉn+M))\bar{S}^n = A^n + \frac{(K^n)^\top Q^n}{\sqrt{D^n}}, \quad S^n = \text{drop}_{p_{\text{at}}}(\text{softmax}(\bar{S}^n + M)) Yˉn=VnSn∈RDn×T,SAℓ(X)=dropph(∑n=1NUℓnYˉn+cℓ)\bar{Y}^n = V^n S^n \in \mathbb{R}^{D^n \times T}, \quad \text{SA}_\ell(X) = \text{drop}_{p_{\text{h}}}\left( \sum_{n=1}^N U_\ell^n \bar{Y}^n + c_\ell \right) where Uℓn∈RD×DnU_\ell^n \in \mathbb{R}^{D \times D^n} and cℓ∈RDc_\ell \in \mathbb{R}^D.

    Feed-Forward Network

    FFNℓ(x)=droppf(UℓfGELU(Wℓfx+bℓf)+cℓf)\text{FFN}_\ell(x) = \text{drop}_{p_{\text{f}}}\left( U_\ell^f \text{GELU}(W_\ell^f x + b_\ell^f) + c_\ell^f \right) where Wℓf∈RD′×DW_\ell^f \in \mathbb{R}^{D' \times D}, Uℓf∈RD×D′U_\ell^f \in \mathbb{R}^{D \times D'}, bℓf∈RD′b_\ell^f \in \mathbb{R}^{D'}, and cℓf∈RDc_\ell^f \in \mathbb{R}^D.

    Output Logits

    Output logits Y∈R∣V∣×TY \in \mathbb{R}^{|V| \times T} are produced by applying a final layer normalization LNf\text{LN}^f to HLH^L and multiplying by the transpose of the untied input embedding matrix WemW^{\text{em}} without an output bias term: Y=(Wem)⊤LNf(HL)Y = (W^{\text{em}})^\top \text{LN}^f(H^L)

    Initialization

    Parameters are sampled from N(0,z)\mathcal{N}(0, z) where z=1/(3D)≈0.006588z = \sqrt{1/(3D)} \approx 0.006588. For the second FFN projection UℓfU_\ell^f and attention output projection UℓnU_\ell^n, the initialization range is rescaled by 1/2L1/\sqrt{2L} to N(0,z′)\mathcal{N}(0, z') with z′=z/2Lz' = z/\sqrt{2L}.

  2. Knowl 2 — FinPile and Public Pretraining Dataset Composition

    data/table

    BloombergGPT was trained on a 709-billion-token pretraining corpus consisting of roughly half domain-specific financial text (FinPile, 51.27%) and half general-purpose public datasets (48.73%). All sub-corpora were deduplicated at the document level.

    Dataset Docs (10410^4) Chars/Doc Chars (10810^8) Chars/Token Tokens (%)
    FinPile 175,886 1,017 17,883 4.92 51.27%
    Web 158,250 933 14,768 4.96 42.01%
    News 10,040 1,665 1,672 4.44 5.31%
    Filings 3,335 2,340 780 5.39 2.04%
    Press 1,265 3,443 435 5.06 1.21%
    Bloomberg Archives 2,996 758 227 4.60 0.70%
    Public 50,744 3,314 16,818 4.87 48.73%
    C4 34,832 2,206 7,683 5.56 19.48%
    The Pile 13,694 6,105 8,409 4.57 25.90%
    Wikipedia (07/01/2022) 2,218 3,271 726 3.06 3.35%
    Total 226,631 1,531 34,701 4.89 100.00%

    FinPile Data Subsets

    • Web (298B tokens, 42.01%): High-quality, financially relevant web pages identified from specific regional domains (US 15.95%, Asia-Pacific 4.72%, UK 1.98%).
    • News (38B tokens, 5.31%): Hundreds of English financial news providers, including transcripts from Bloomberg TV (0.41%).
    • Filings (14B tokens, 2.04%): Company financial disclosures and regulatory filings (majority from SEC EDGAR, 1.90% of total), normalized from original PDF documents.
    • Press Releases (9B tokens, 1.21%): Public corporate communications and financial press releases.
    • Bloomberg Internal (5B tokens, 0.70%): Bloomberg News (0.44%) and Bloomberg First Word real-time news wire (0.13%). Documents in FinPile span timestamps from March 1, 2007 to July 31, 2022.
  3. Knowl 3 — Parallel Hierarchical Unigram Tokenizer Construction

    algorithm

    BloombergGPT utilizes a byte-based Unigram tokenizer with a vocabulary of 217=131,0722^{17} = 131{,}072 tokens trained on The Pile. To overcome memory constraints and the computational complexity of Unigram optimization over large datasets, the tokenizer is trained in parallel using a hierarchical split-and-merge approach.

    Input: Training corpus consisting of 22 domains D1,…,D22D_1, \dots, D_{22} from The Pile; target vocabulary size Vtarget=131,072V_{\text{target}} = 131{,}072
    Output: Unified Unigram tokenizer vocabulary VfinalV_{\text{final}} with assigned unigram token probabilities pp
    1. Pre-tokenize raw text into chunks using greedy regular expression matching:
         [ A-Za-z]+|[0-9]|[^A-Za-z0-9]+
       (letters include leading spaces to build multi-word tokens; individual digits are isolated into single-character tokens).
    2. for each domain k∈{1,…,22}k \in \{1, \dots, 22\}:
         Split domain DkD_k into 256 equal-sized byte shards Sk,1,…,Sk,256S_{k,1}, \dots, S_{k,256}
         for each shard j∈{1,…,256}j \in \{1, \dots, 256\}:
           Train an independent Unigram tokenizer Tk,jT_{k,j} with vocabulary size 65,536 (2162^{16})
         Merge the 256 shard tokenizers into domain tokenizer TkT_k by computing weighted average token probabilities proportional to shard byte sizes
    3. Combine the 22 domain tokenizers into a single global candidate tokenizer TglobalT_{\text{global}} (approx7,000,000\\approx 7{,}000{,}000 tokens) by computing weighted average token probabilities proportional to domain byte sizes
    4. Drop tokens with the lowest probability mass until Vtarget=131,072V_{\text{target}} = 131{,}072 tokens remain, and re-normalize probabilities
    5. Identify any of the 256 possible single-byte tokens not present in the vocabulary (36 missing bytes) and add them, along with the <|endoftext|> token
    6. return VfinalV_{\text{final}}

    Vocabulary Size Selection Heuristic

    The vocabulary size was selected by evaluating candidate sizes from 25,00025{,}000 to 550,000550{,}000 on the C4 dataset. Each candidate size VV computes the total compressed size (in bits) of C4 when representing each token in log⁡2(V)\log_2(V) bits. The minimum compressed size occurred at V=125,000V = 125{,}000, which was rounded up to the nearest power of 2 (217=131,0722^{17} = 131{,}072 tokens).

  4. Knowl 4 — Compute-Optimal Model Scaling and Layer-Width Allocation

    equation

    Under a fixed compute budget of 1.3×1061.3 \times 10^6 GPU hours on NVIDIA 40GB A100 GPUs, the model size and training tokens were determined using Chinchilla scaling laws adjusted for activation checkpointing overhead, combined with depth-to-width allocation laws.

    Compute Scaling with Checkpointing Adjustment

    Activation checkpointing recomputes forward passes during backpropagation, adding approximately 0.33×0.33\times FLOPs per step (1.33×1.33\times total forward computation). To account for this overhead, the effective compute budget plugged into Chinchilla equations is scaled by 0.75×1.3M0.75 \times 1.3\text{M} hours, corresponding to total compute C=2.36×1023C = 2.36 \times 10^{23} FLOPs.

    Fitting log-scaled regression lines to Chinchilla empirical evaluations yields the optimal parameter count N∗N^* and token count D∗D^*:

    • Chinchilla Approach 1: log⁡10(N∗)=0.498⋅log⁡10(C)−1.004  ⟹  N∗=52.993B parameters\log_{10}(N^*) = 0.498 \cdot \log_{10}(C) - 1.004 \implies N^* = 52.993\text{B parameters} log⁡10(D∗)=0.502⋅log⁡10(C)+0.229  ⟹  D∗=1111.112B tokens\log_{10}(D^*) = 0.502 \cdot \log_{10}(C) + 0.229 \implies D^* = 1111.112\text{B tokens}

    • Chinchilla Approach 2: log⁡10(N∗)=0.490⋅log⁡10(C)−0.839  ⟹  N∗=49.753B parameters\log_{10}(N^*) = 0.490 \cdot \log_{10}(C) - 0.839 \implies N^* = 49.753\text{B parameters} log⁡10(D∗)=0.510⋅log⁡10(C)+0.062  ⟹  D∗=1175.766B tokens\log_{10}(D^*) = 0.510 \cdot \log_{10}(C) + 0.062 \implies D^* = 1175.766\text{B tokens}

    Because domain-specific tokens in FinPile were limited to ≈363B\approx 363\text{B} tokens and required to constitute at least 50% of the training mixture, the token count was capped around 569B569\text{B} tokens, leaving a ∼30%\sim 30\% compute buffer for retries and restarts.

    Width-to-Depth Allocation

    To allocate the ≈50B\approx 50\text{B} parameters across layers LL and hidden dimension DD, the depth-to-width scaling relationship from Levine et al. is applied: D=exp⁡(5.039)⋅exp⁡(0.0555⋅L)D = \exp(5.039) \cdot \exp(0.0555 \cdot L) Sweeping integer values of LL to match ≈50B\approx 50\text{B} parameters gives L=70L = 70 and D=7,510D = 7{,}510. Enforcing divisibility by N=40N = 40 attention heads (giving head dimension Dn=192D^n = 192) and alignment to multiples of 8 for Tensor Core efficiency yields final shape L=70,D=7,680L = 70, D = 7{,}680 (50.6B50.6\text{B} total parameters).

  5. Knowl 5 — Distributed Optimization and Hardware Infrastructure

    experimental setup

    Hardware Stack

    • Compute Cluster: 64 AWS p4d.24xlarge instances totaling 512 NVIDIA 40GB A100 GPUs.
    • Interconnect: NVIDIA NVSwitch intra-node (600 GB/s bandwidth) and GPUDirect over AWS Elastic Fabric Adapter (EFA) inter-node (400 Gb/s bandwidth).
    • Storage: Amazon FSx for Lustre file system with throughput up to 1000 MB/s per TiB.

    Distributed Strategy and Parallelism

    • Model Distribution: AWS SageMaker Model Parallelism (SMP) library implementing ZeRO Stage 3 optimization.
    • Sharding: Model state (parameters, gradients, and optimizer state) sharded across 128 GPUs, maintaining 4 data-parallel replicas across the 512 GPUs.
    • Communication: Scale-aware model partitioning and 2-hop gradient aggregation from MiCS.
    • Memory Management: Activation checkpointing enabled on all 70 transformer layers.
    • Throughput: Average computation rate of 102 TFLOPs per GPU with a step duration of 32.5 seconds.

    Mixed Precision and Kernel Optimizations

    • Forward and backward passes executed in BF16.
    • Master weights, optimizer states, ALiBi matrices, and loss softmax computed in FP32.
    • Self-attention uses a fused masked-causal-softmax kernel in SMP, saving GPU memory and contributing a 4–5 TFLOPs speedup.

    Optimization Hyperparameters

    • Optimizer: AdamW with β1=0.9\beta_1 = 0.9, β2=0.95\beta_2 = 0.95, weight decay λ=0.1\lambda = 0.1 applied strictly to non-bias, non-LayerNorm parameters.
    • Learning Rate Schedule: Linear warmup over the first 1,800 steps to LRmax=6×10−5\text{LR}_{\text{max}} = 6 \times 10^{-5}, followed by cosine decay down to LRfinal=6×10−6\text{LR}_{\text{final}} = 6 \times 10^{-6} (0.1×LRmax0.1 \times \text{LR}_{\text{max}}).
    • Batch Size Warmup: Batch size of 1,024 sequences (2.1M tokens) for the initial 7,200 steps, transitioning to 2,048 sequences (4.2M tokens) for subsequent steps.
    • Gradient Clipping: Maximum gradient L2L_2 norm threshold of 0.3.
  6. Knowl 6 — Training Dynamics, Failure Diagnoses, and Stabilization Interventions

    empirical result

    Training BloombergGPT involved several failed configurations and runtime interventions before arriving at the stable final run:

    1. Curriculum Learning Failure (Run v0)

    Ordering training data strictly by chronological timestamp (2007 to 2022) resulted in stagnant loss around step 20k, with training loss diverging from validation loss. Chronological ordering was abandoned in favor of uniform shard shuffling.

    2. Gradient Explosion and LayerNorm Weight Growth (Run v1)

    In run v1.0, gradient norms exploded after step 12k with sharp validation loss spikes. Weight norm tracking identified the Layer 1 input LayerNorm multiplier (gamma1in\\gamma_1^{\text{in}}) growing unbounded. Investigation uncovered two root causes:

    1. An inherited implementation bug applying AdamW weight decay to LayerNorm scale parameters (γ\gamma) initialized to 1.
    2. The omission of embedding layer normalization (LNemLN^{\text{em}}).

    Stabilization was achieved in run v2 by reintroducing LNemLN^{\text{em}}, removing weight decay from all γ\gamma multipliers, adopting Megatron standard deviation rescaling (1/2L1/\sqrt{2L}), applying query-key layer scaling, reducing LRmax\text{LR}_{\text{max}} from 1×10−41 \times 10^{-4} to 6×10−56 \times 10^{-5}, and lowering gradient clipping from 1.0 to 0.3.

    3. Mid-Run Validation Plateaus and Interventions (Run v2)

    Training progressed stably for 42 days (115,500 steps) until validation loss plateaued at 2.116 and downstream performance stalled. Sequential adjustments were executed:

    • Step 115,500 (v2.1): Rolled back to step 115,500, re-shuffled subsequent training data, and reduced max learning rate from 6×10−56 \times 10^{-5} to 4×10−54 \times 10^{-5}.
    • Step 129,900 (v2.3): Halved learning rate to 2×10−52 \times 10^{-5} and introduced dropout (p=0.1p = 0.1).
    • Step 137,100 (v2.5): Halved learning rate to 1×10−51 \times 10^{-5} with dropout p=0.1p = 0.1.
    • Termination: Training concluded at step 139,200 (569B tokens processed, ≈80%\approx 80\% of one epoch) when validation loss and downstream benchmarks (MMLU and BIG-bench Hard) showed no further progress.
  7. Knowl 7 — External Financial NLP Benchmark Performance

    data/table

    BloombergGPT was evaluated against peer open models on standard financial NLP benchmarks under standardized few-shot prompting setups.

    Task BloombergGPT (50.6B) GPT-NeoX (20B) OPT (66B) BLOOM (176B)
    ConvFinQA 43.41 30.06 27.88 36.31
    FiQA SA 75.07 50.59 51.60 53.12
    FPB 51.07 44.64 48.67 50.25
    Headline 82.20 73.22 79.41 76.51
    NER 60.82 60.98 57.49 55.56
    All Tasks (avg) 62.51 51.90 53.01 54.35
    Win Rate 0.93 0.27 0.33 0.47

    Task Formulations and Metrics

    • ConvFinQA (1-shot exact match accuracy): Multi-turn conversational numerical reasoning over S&P 500 financial earnings reports containing text and Markdown tables.
    • FiQA SA (5-shot weighted F1): Aspect-specific sentiment classification (positive, neutral, negative) discretized from continuous sentiment scores [−1,1][-1, 1] on financial news headlines and microblogs.
    • FPB (5-shot weighted F1): Financial PhraseBank sentiment classification on sentences from financial news with ≥50%\ge 50\% annotator agreement.
    • Headline (5-shot weighted F1): Binary classification across 9 sub-categories regarding gold commodity price and market movements.
    • NER (20-shot entity-level F1): Named entity extraction (PER, LOC, ORG) on financial agreements filed with the SEC.

    BloombergGPT achieved the highest performance on 4 of the 5 tasks, obtaining an overall win rate of 0.93 against evaluated models.

  8. Knowl 8 — Internal Financial Domain Evaluation: Sentiment, NER, and Ticker Disambiguation

    empirical result

    BloombergGPT was evaluated on proprietary, expert-curated Bloomberg financial benchmarks covering aspect sentiment analysis, named entity recognition (NER), and joint named entity disambiguation (NER+NED).

    1. Aspect-Specific Sentiment Analysis (5-shot, Weighted F1)

    Across internal sentiment tasks, BloombergGPT outperformed comparison models by over 25 to 60 percentage points on news and transcripts:

    • Equity News Sentiment: BloombergGPT 79.63 vs. GPT-NeoX 14.17, OPT66B\text{OPT}_{66\text{B}} 20.98, BLOOM176B\text{BLOOM}_{176\text{B}} 19.96.
    • Equity Social Media Sentiment: BloombergGPT 72.40 vs. GPT-NeoX 66.48, OPT66B\text{OPT}_{66\text{B}} 71.36, BLOOM176B\text{BLOOM}_{176\text{B}} 68.04.
    • Equity Transcript Sentiment: BloombergGPT 65.06 vs. GPT-NeoX 25.08, OPT66B\text{OPT}_{66\text{B}} 37.58, BLOOM176B\text{BLOOM}_{176\text{B}} 34.82.
    • ES News (ESG) Sentiment: BloombergGPT 46.12 vs. GPT-NeoX 26.99, OPT66B\text{OPT}_{66\text{B}} 31.44, BLOOM176B\text{BLOOM}_{176\text{B}} 28.07.
    • Country News Sentiment: BloombergGPT 49.14 vs. GPT-NeoX 13.45, OPT66B\text{OPT}_{66\text{B}} 17.41, BLOOM176B\text{BLOOM}_{176\text{B}} 16.06.
    • Average Across All Sentiment Tasks: BloombergGPT 62.47 (win rate 1.00) vs. GPT-NeoX 29.23, OPT66B\text{OPT}_{66\text{B}} 35.76, BLOOM176B\text{BLOOM}_{176\text{B}} 33.39.

    2. Internal NER and Joint NER+NED (20-shot, Entity F1)

    Evaluated across 7 source types (Bloomberg First Word, Bloomberg News, SEC Filings, Headlines, Premium News, Press Conference Transcripts, Social Media):

    • Standard NER: BLOOM176B\text{BLOOM}_{176\text{B}} achieved an average F1 of 64.83 (win rate 0.95); BloombergGPT achieved 62.63 (win rate 0.57), outperforming similarly sized models (GPT-NeoX 59.75, OPT66B\text{OPT}_{66\text{B}} 58.09).
    • NER+NED (Company Mention to Stock Ticker Disambiguation): The model identifies company mentions and directly outputs corresponding exchange ticker symbols (e.g., mapping Apple and Intel to AAPL, INTC). BloombergGPT achieved an average F1 of 64.83 (win rate 0.95) compared to GPT-NeoX (39.26), OPT66B\text{OPT}_{66\text{B}} (58.79), and BLOOM176B\text{BLOOM}_{176\text{B}} (45.43).
  9. Knowl 9 — General LLM Benchmark Performance

    data/table

    To evaluate general capabilities, BloombergGPT was evaluated on standard general-purpose LLM benchmarks. Across evaluated models in its parameter class, BloombergGPT achieved the highest average scores without performance degradation from domain specialization.

    Benchmark Suite BloombergGPT GPT-NeoX OPT (66B) BLOOM (176B) GPT-3 (175B)
    BIG-bench Hard (3-shot avg) 41.97 40.25 39.58 44.91 52.30
    NLP Tasks avg 54.39 51.63 52.60 54.96 62.70
    Algorithmic Tasks avg 28.42 27.84 25.37 33.95 40.90
    Win Rate 0.57 0.45 0.39 0.75 –
    MMLU (5-shot accuracy) 39.18 35.95 35.99 39.13 43.90
    Humanities 36.26 32.75 33.28 34.05 40.80
    STEM 35.12 33.43 30.72 36.75 36.70
    Social Sciences 40.04 36.63 38.32 41.50 50.40
    Other (Finance/Accounting) 46.36 42.29 42.63 46.48 48.80
    Knowledge Tasks (1-shot avg) 66.50 63.10 64.95 67.01 –
    ARC (easy) 73.99 70.79 71.25 75.93 71.20
    ARC (challenging) 48.63 45.39 44.54 50.85 53.20
    CommonsenseQA 65.52 60.36 66.42 64.21 –
    PiQA 77.86 75.84 77.58 77.04 80.50
    Win Rate 0.75 0.08 0.33 0.67 –
    Reading Comprehension (1-shot avg) 61.22 42.81 50.21 49.37 67.00
    BoolQ 74.59 46.36 57.46 52.94 76.70
    OpenBookQA 51.60 44.20 58.00 47.20 58.80
    RACE (middle) 54.32 41.23 47.42 52.30 57.40
    RACE (high) 41.74 34.33 37.02 39.14 45.90
    MultiRC 62.29 22.86 18.80 26.65 72.90
    ReCoRD 82.79 67.86 82.53 78.01 90.20
    Win Rate 0.94 0.06 0.50 0.50 –
    Linguistic Tasks (1-shot avg) 60.63 57.18 58.59 58.26 63.40
    Win Rate 0.85 0.27 0.58 0.42 –
  10. Knowl 10 — Few-Shot Scoring and Prompt Formulation Methodology

    model/method

    Downstream tasks are evaluated via standard few-shot prompting without task descriptions, prompt parameter fine-tuning, or Chain-of-Thought prompting. Few-shot exemplars are randomly sampled from training splits per test instance.

    Classification Decision Rules

    For classification tasks where a fixed set of candidate output strings A\mathcal{A} is given, the model selects a candidate α^∈A\hat{\alpha} \in \mathcal{A} based on conditional sequence probabilities p(α∣s)p(\alpha \mid s) given input context ss according to one of three scoring strategies:

    1. Regular Likelihood Scoring: α^reg=arg⁡max⁡α∈Ap(α∣s)\hat{\alpha}_{\text{reg}} = \arg\max_{\alpha \in \mathcal{A}} p(\alpha \mid s)

    2. Calibrated Likelihood Scoring: Normalizes candidate likelihood by the unconditional candidate probability conditioned only on the static prompt prefix (e.g., "Answer:"): α^cal=arg⁡max⁡α∈Ap(α∣s)p(α∣"Answer:")\hat{\alpha}_{\text{cal}} = \arg\max_{\alpha \in \mathcal{A}} \frac{p(\alpha \mid s)}{p(\alpha \mid \text{"Answer:"})}

    3. Length-Normalized Likelihood Scoring: Normalizes candidate log-likelihood by the number of sub-word tokens len(α)\text{len}(\alpha) in the candidate string: α^norm=arg⁡max⁡α∈Ap(α∣s)len(α)\hat{\alpha}_{\text{norm}} = \arg\max_{\alpha \in \mathcal{A}} \frac{p(\alpha \mid s)}{\text{len}(\alpha)}

    For each model and evaluation task, the scoring method that yields the highest accuracy on the validation split is selected. For open-ended generation tasks, autoregressive outputs are generated via greedy decoding.

Coverage note — None was omitted; all key architectural components, dataset breakdowns, tokenization details, training stability dynamics, and financial/general benchmark results are fully captured.

References

  1. 1.Dogu Araci. Finbert: Financial sentiment analysis with pre-trained language models. arXiV preprint arXiV:1908.10063, 2019.
  2. 2.Siqi Bao, Huang He, Fan Wang, Hua Wu, Haifeng Wang, Wenquan Wu, Zhihua Wu, Zhen Guo, Hua Lu, Xinxian Huang, Xin Tian, Xinchao Xu, Yingzhan Lin, and Zheng-Yu Niu. PLATO-XL: Exploring the large-scale pre-training of dialogue generation. In Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, pages 107–118, Online only, November 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.findings-aacl.10.
  3. 3.Iz Beltagy, Kyle Lo, and Arman Cohan. SciBERT: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3615–3620, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1371. URL https://aclanthology.org/D19-1371.
  4. 4.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623, 2021.
  5. 5.Luisa Bentivogli, Bernardo Magnini, Ido Dagan, Hoa Trang Dang, and Danilo Giampiccolo. The fifth PASCAL recognizing textual entailment challenge. In Proceedings of the Second Text Analysis Conference, TAC 2009, Gaithersburg, Maryland, USA, November 16-17, 2009. NIST, 2009. URL https://tac.nist.gov/publications/2009/additional.papers/RTE5_overview.proceedings.pdf.
  6. 6.Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. The values encoded in machine learning research. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 173–184, 2022.
  7. 7.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. PIQA: reasoning about physical commonsense in natural language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 7432–7439. AAAI Press, 2020. URL https://ojs.aaai.org/index.php/AAAI/article/view/6239.
  8. 8.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021. URL https://doi.org/10.5281/zenodo.5297715. If you use this software, please cite it using these metadata.
  9. 9.Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of BigScience Episode #5 – Workshop on Challenges & Perspectives in Creating Large Language Models, pages 95–136, virtual+Dublin, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.bigscience-1.9. URL https://aclanthology.org/2022.bigscience-1.9.
  10. 10.Elliot Bolton, David Hall, Michihiro Yasunaga, Tony Lee, Chris Manning, and Percy Liang. BioMedLM. https://github.com/stanford-crfm/BioMedLM, 2023.
  11. 11.Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, S. Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen A. Creel, Jared Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren E. Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas F. Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, O. Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Benjamin Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, J. F. Nyarko, Giray Ogut, Laurel J. Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Robert Reich, Hongyu Ren, Frieda Rong, Yusuf H. Roohani, Camilo Ruiz, Jack Ryan, Christopher R’e, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishna Parasuram Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tram`er, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei A. Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang. On the opportunities and risks of foundation models. ArXiV, abs/2108.07258, 2021.
  12. 12.Kaj Bostrom and Greg Durrett. Byte pair encoding is suboptimal for language model pretraining. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4617–4624, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.414. URL https://aclanthology.org/2020.findings-emnlp.414.
  13. 13.Thorsten Brants, Ashok C. Popat, Peng Xu, Franz J. Och, and Jeffrey Dean. Large language models in machine translation. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), pages 858–867, Prague, Czech Republic, June 2007. Association for Computational Linguistics. URL https://aclanthology.org/D07-1090.
  14. 14.Peter F Brown, Vincent J Della Pietra, Peter V Desouza, Jennifer C Lai, and Robert L Mercer. Class-based n-gram models of natural language. Computational linguistics, 18(4):467–480, 1992.
  15. 15.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  16. 16.Nicholas Carlini, Florian Tram`er, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, Ulfar Erlingsson, ´ Alina Oprea, and Colin Raffel. Extracting training data from large language models. In USENIX Security Symposium, 2020.
  17. 17.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models, 2022. URL https://arxiv.org/abs/2202.07646.
  18. 18.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harrison Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, David W. Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William H. Guss, Alex Nichol, Igor Babuschkin, S. Arun Balaji, Shantanu Jain, Andrew Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew M. Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. arXiV, abs/2107.03374, 2021a.
  19. 19.Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost. arXiV preprint arXiV:1604.06174, 2016.
  20. 20.Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. FinQA: A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3697–3711, Online and Punta Cana, Dominican Republic, November 2021b. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.300. URL https://aclanthology.org/2021.emnlp-main.300.
  21. 21.Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. ConvFinQA: Exploring the chain of numerical reasoning in conversational finance question answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6279–6292, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.421.
  22. 22.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garc´ıa, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark D´ıaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways. arXiV, abs/2204.02311, 2022.
  23. 23.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2924–2936, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1300. URL https://aclanthology.org/N19-1300.
  24. 24.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiV, abs/1803.05457, 2018.
  25. 25.Ido Dagan, Oren Glickman, and Bernardo Magnini. The pascal recognising textual entailment challenge. In Machine Learning Challenges Workshop, 2007.
  26. 26.Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung, pages 107–124, 2019.
  27. 27.Alexandra DeLucia, Shijie Wu, Aaron Mueller, Carlos Aguirre, Philip Resnik, and Mark Dredze. Bernice: A multilingual pre-trained encoder for Twitter. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6191–6205, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.415.
  28. 28.Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 8-bit optimizers via blockwise quantization. In International Conference on Learning Representations, 2022.
  29. 29.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  30. 30.Jesse Dodge, Maarten Sap, Ana Marasovi´c, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1286–1305, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.98. URL https://aclanthology.org/2021.emnlp-main.98.
  31. 31.Mark Dredze, Prabhanjan Kambadur, Gary Kazantsev, Gideon Mann, and Miles Osborne. How twitter is changing the nature of financial news discovery. In proceedings of the second international workshop on data science for macro-modeling, pages 1–5, 2016.
  32. 32.Ingrid E Fisher, Margaret R Garnsey, and Mark E Hughes. Natural language processing in accounting, auditing and finance: A synthesis of the literature with a roadmap for future research. Intelligent Systems in Accounting, Finance and Management, 23(3):157–214, 2016.
  33. 33.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling, 2021. URL https://arxiv.org/abs/2101.00027.
  34. 34.Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text, 2022. URL https://arxiv.org/abs/2202.06935.
  35. 35.Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. The third PASCAL recognizing textual entailment challenge. In Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, pages 1–9, Prague, June 2007. Association for Computational Linguistics. URL https://aclanthology.org/W07-1401.
  36. 36.Amelia Glaese, Nat McAleese, Maja Trebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Soˇna Mokr´a, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving. Improving alignment of dialogue agents via targeted human judgements, 2022. URL https://arxiv.org/abs/2209.14375.
  37. 37.Andrew S. Gordon, Zornitsa Kozareva, and Melissa Roemmele. Semeval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In International Workshop on Semantic Evaluation, 2011.
  38. 38.Tanya Goyal, Junyi Jessy Li, and Greg Durrett. News summarization and evaluation in the era of gpt-3, 2022. URL https://arxiv.org/abs/2209.12356.
  39. 39.Suchin Gururangan, Ana Marasovi´c, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.740. URL https://aclanthology.org/2020.acl-main.740.
  40. 40.R Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. The second pascal recognising textual entailment challenge. In Proceedings of the Second PASCAL Challenges Workshop on Recognising Textual Entailment, volume 7, 2006.
  41. 41.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiV preprint arXiV:1606.08415, 2016.
  42. 42.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ.
  43. 43.Alex Henry, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen. Query-key normalization for transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4246–4253, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.379. URL https://aclanthology.org/2020.findings-emnlp.379.
  44. 44.Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish. Scaling laws for transfer. arXiV preprint arXiV:2102.01293, 2021.
  45. 45.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack William Rae, and Laurent Sifre. An empirical analysis of compute-optimal large language model training. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=iBBcRUlOAPR.
  46. 46.Jeremy Howard and Sebastian Ruder. Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328–339, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1031. URL https://aclanthology.org/P18-1031.
  47. 47.Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. Clinicalbert: Modeling clinical notes and predicting hospital readmission. arXiV, 4 2019. URL http://arxiv.org/abs/1904.05342.
  48. 48.Frederick Jelinek. Continuous speech recognition by statistical methods. Proceedings of the IEEE, 64(4):532–556, 1976.
  49. 49.Yacine Jernite, Huu Nguyen, Stella Biderman, Anna Rogers, Maraim Masoud, Valentin Danchev, Samson Tan, Alexandra Sasha Luccioni, Nishant Subramani, Isaac Johnson, Gerard Dupont, Jesse Dodge, Kyle Lo, Zeerak Talat, Dragomir Radev, Aaron Gokaslan, Somaieh Nikpoor, Peter Henderson, Rishi Bommasani, and Margaret Mitchell. Data governance in the age of large-scale data-driven language technology. In 2022 ACM Conference on Fairness, Accountability, and Transparency. ACM, jun 2022. doi: 10.1145/3531146.3534637. URL https://doi.org/10.1145%2F3531146.3534637.
  50. 50.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiV, 1 2020. URL http://arxiv.org/abs/2001.08361.
  51. 51.Can Karakus, Rahul Huilgol, Fei Wu, Anirudh Subramanian, Cade Daniel, Derya Cavdar, Teng Xu, Haohan Chen, Arash Rahnama, and Luis Quintela. Amazon sagemaker model parallelism: A general and flexible framework for large model training. arXiV preprint arXiV:2111.05972, 2021.
  52. 52.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 252–262, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-1023. URL https://aclanthology.org/N18-1023.
  53. 53.Vijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. Reducing activation recomputation in large transformer models, 2022. URL https://arxiv.org/abs/2205.05198.
  54. 54.Taku Kudo. Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66–75, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1007. URL https://aclanthology.org/P18-1007.
  55. 55.Taku Kudo and John Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium, November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-2012. URL https://aclanthology.org/D18-2012.
  56. 56.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. RACE: Large-scale ReAding comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785–794, Copenhagen, Denmark, September 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1082. URL https://aclanthology.org/D17-1082.
  57. 57.Teven Le Scao, Thomas Wang, Daniel Hesslow, Stas Bekman, M Saiful Bari, Stella Biderman, Hady Elsahar, Niklas Muennighoff, Jason Phang, Ofir Press, Colin Raffel, Victor Sanh, Sheng Shen, Lintang Sutawika, Jaesung Tae, Zheng Xin Yong, Julien Launay, and Iz Beltagy. What language model to train if you have one million GPU hours? In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 765–782, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.findings-emnlp.54.
  58. 58.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36:1234–1240, 2 2020. ISSN 14602059. doi: 10.1093/bioinformatics/btz682.
  59. 59.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445, Dublin, Ireland, May 2022a. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.577. URL https://aclanthology.org/2022.acl-long.577.
  60. 60.Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, Rose E. Wang, Minae Kwon, Joon Sung Park, Hancheng Cao, Tony Lee, Rishi Bommasani, Michael S. Bernstein, and Percy Liang. Evaluating human-language model interaction. CoRR, abs/2212.09746, 2022b. doi: 10.48550/arXiv.2212.09746. URL https://doi.org/10.48550/arXiv.2212.09746.
  61. 61.Eric Lehman, Evan Hernandez, Diwakar Mahajan, Jonas Wulff, Micah J. Smith, Zachary Ziegler, Daniel Nadler, Peter Szolovits, Alistair Johnson, and Emily Alsentzer. Do we still need clinical language models?, 2023. URL https://arxiv.org/abs/2302.08091.
  62. 62.Hector J. Levesque, Ernest Davis, and L. Morgenstern. The winograd schema challenge. In International Conference on Principles of Knowledge Representation and Reasoning, 2011.
  63. 63.Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua. Limits to depth efficiencies of self-attention. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 22640–22651. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/ff4dfdf5904e920ce52b48c1cef97829-Paper.pdf.
  64. 64.Patrick Lewis, Myle Ott, Jingfei Du, and Veselin Stoyanov. Pretrained language models for biomedical and clinical tasks: Understanding and extending the state-of-the-art. In Proceedings of the 3rd Clinical Natural Language Processing Workshop, pages 146–157, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.clinicalnlp-1.17. URL https://aclanthology.org/2020.clinicalnlp-1.17.
  65. 65.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving quantitative reasoning problems with language models, 2022. URL https://arxiv.org/abs/2206.14858.
  66. 66.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher R´e, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel J. Orr, Lucia Zheng, Mert Y¨uksekg¨on¨ul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. Holistic evaluation of language models. CoRR, abs/2211.09110, 2022. doi: 10.48550/arXiv.2211.09110. URL https://doi.org/10.48550/arXiv.2211.09110.
  67. 67.Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham. Jurassic-1: Technical details and evaluation. White Paper. AI21 Labs, 1, 2021.
  68. 68.Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Sal Candido, and Alexander Rives. Language models of protein sequences at the scale of evolution enable accurate structure prediction. bioRxiv, 2022. doi: 10.1101/2022.07.20.500902. URL https://www.biorxiv.org/content/early/2022/07/21/2022.07.20.500902.
  69. 69.Tianyu Liu, Yuchen Eleanor Jiang, Nicholas Monath, Ryan Cotterell, and Mrinmaya Sachan. Autoregressive structured prediction with language models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 993–1005, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.findings-emnlp.70.
  70. 70.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7.
  71. 71.Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. BioGPT: generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23(6), sep 2022. doi: 10.1093/bib/bbac409. URL https://doi.org/10.1093%2Fbib%2Fbbac409.
  72. 72.Jouni Luoma and Sampo Pyysalo. Exploring cross-sentence contexts for named entity recognition with BERT. In Proceedings of the 28th International Conference on Computational Linguistics, pages 904–914, Barcelona, Spain (Online), December 2020. International Committee on Computational Linguistics. doi: 10.18653/v1/2020.coling-main.78. URL https://aclanthology.org/2020.coling-main.78.
  73. 73.Macedo Maia, Siegfried Handschuh, Andr´e Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. Www’18 open challenge: Financial opinion mining and question answering. In Pierre-Antoine Champin, Fabien Gandon, Mounia Lalmas, and Panagiotis G. Ipeirotis, editors, Companion of the The Web Conference 2018 on The Web Conference 2018, WWW 2018, Lyon , France, April 23-27, 2018, pages 1941–1942. ACM, 2018. doi: 10.1145/3184558.3192301. URL https://doi.org/10.1145/3184558.3192301.
  74. 74.Pekka Malo, Ankur Sinha, Pekka J. Korhonen, Jyrki Wallenius, and Pyry Takala. Good debt or bad debt: Detecting semantic orientations in economic texts. J. Assoc. Inf. Sci. Technol., 65(4):782–796, 2014. doi: 10.1002/asi.23062. URL https://doi.org/10.1002/asi.23062.
  75. 75.Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gall´e, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoˆıt Sagot, and Samson Tan. Between words and characters: A brief history of open-vocabulary modeling and tokenization in nlp, 2021. URL https://arxiv.org/abs/2112.10508.
  76. 76.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1260. URL https://aclanthology.org/D18-1260.
  77. 77.Tomas Mikolov, Martin Karafi´at, Lukas Burget, Jan Cernock`y, and Sanjeev Khudanpur. Recurrent neural network based language model. In Interspeech, pages 1045–1048. Makuhari, 2010.
  78. 78.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 839–849, San Diego, California, June 2016. Association for Computational Linguistics. doi: 10.18653/v1/N16-1098. URL https://aclanthology.org/N16-1098.
  79. 79.Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. BERTweet: A pre-trained language model for English tweets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 9–14, Online, October 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.2. URL https://aclanthology.org/2020.emnlp-demos.2.
  80. 80.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. Adversarial NLI: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4885–4901, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.441. URL https://aclanthology.org/2020.acl-main.441.
  81. 81.Erik Nijkamp, Jeffrey Ruffolo, Eli N. Weinstein, Nikhil Naik, and Ali Madani. Progen2: Exploring the boundaries of protein language models. CoRR, abs/2206.13517, 2022. doi: 10.48550/arXiv.2206.13517. URL https://doi.org/10.48550/arXiv.2206.13517.
  82. 82.NVIDIA. Train with mixed precision, 2023. URL https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html.
  83. 83.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=TG8KACxEON.
  84. 84.Baolin Peng, Michel Galley, Pengcheng He, Chris Brockett, Lars Liden, Elnaz Nouri, Zhou Yu, Bill Dolan, and Jianfeng Gao. Godel: Large-scale pre-training for goal-directed dialog. arXiV preprint arXiV:2206.11309, 2022.
  85. 85.Mohammad Taher Pilehvar and Jose Camacho-Collados. WiC: the word-in-context dataset for evaluating context-sensitive meaning representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1267–1273, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1128. URL https://aclanthology.org/N19-1128.
  86. 86.Ofir Press, Noah Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=R8sQPpGCv0.
  87. 87.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training, 2018. URL https://gluebenchmark.com/leaderboard.
  88. 88.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners, 2019. URL https://github.com/codelucas/newspaper.
  89. 89.Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. Compressive transformers for long-range sequence modelling. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=SylKikSYDH.
  90. 90.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. Scaling language models: Methods, analysis & insights from training gopher. arXiV, 12 2021. URL http://arxiv.org/abs/2112.11446.
  91. 91.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  92. 92.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16. IEEE, 2020.
  93. 93.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325, Online, April 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.24. URL https://aclanthology.org/2021.eacl-main.24.
  94. 94.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WINOGRANDE: An adversarial winograd schema challenge at scale. Commun. ACM, 64:99–106, 2019.
  95. 95.Julio Cesar Salinas Alvarado, Karin Verspoor, and Timothy Baldwin. Domain adaption of named entity recognition to support credit risk assessment. In Proceedings of the Australasian Language Technology Association Workshop 2015, pages 84–90, Parramatta, Australia, December 2015. URL https://aclanthology.org/U15-1010.
  96. 96.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=9Vrb9D0WI4.
  97. 97.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili´c, Daniel Hesslow, Roman Castagn´e, Alexandra Sasha Luccioni, Fran¸cois Yvon, Matthias Gall´e, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoˆıt Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Lauren¸con, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, Dragomir Radev, Eduardo Gonz´alez Ponferrada, Efrat Levkovizh, Ethan Kim, Eyal Bar Natan, Francesco De Toni, G´erard Dupont, Germ´an Kruszewski, Giada Pistilli, Hady Elsahar, Hamza Benyamina, Hieu Tran, Ian Yu, Idris Abdulmumin, Isaac Johnson, Itziar Gonzalez-Dios, Javier de la Rosa, Jenny Chim, Jesse Dodge, Jian Zhu, Jonathan Chang, J¨org Frohberg, Joseph Tobing, Joydeep Bhattacharjee, Khalid Almubarak, Kimbo Chen, Kyle Lo, Leandro Von Werra, Leon Weber, Long Phan, Loubna Ben allal, Ludovic Tanguy, Manan Dey, Manuel Romero Mu˜noz, Maraim Masoud, Mar´ıa Grandury, Mario Saˇsko, Max Huang, Maximin Coavoux, Mayank ˇ Singh, Mike Tian-Jian Jiang, Minh Chien Vu, Mohammad A. Jauhar, Mustafa Ghaleb, Nishant Subramani, Nora Kassner, Nurulaqilla Khamis, Olivier Nguyen, Omar Espejel, Ona de Gibert, Paulo Villegas, Peter Henderson, Pierre Colombo, Priscilla Amuok, Quentin Lhoest, Rheza Harliman, Rishi Bommasani, Roberto Luis L´opez, Rui Ribeiro, Salomey Osei, Sampo Pyysalo, Sebastian Nagel, Shamik Bose, Shamsuddeen Hassan Muhammad, Shanya Sharma, Shayne Longpre, Somaieh Nikpoor, Stanislav Silberberg, Suhas Pai, Sydney Zink, Tiago Timponi Torrent, Timo Schick, Tristan Thrush, Valentin Danchev, Vassilina Nikoulina, Veronika Laippala, Violette Lepercq, Vrinda Prabhu, Zaid Alyafeai, Zeerak Talat, Arun Raja, Benjamin Heinzerling, Chenglei Si, Davut Emre Ta¸sar, Elizabeth Salesky, Sabrina J. Mielke, Wilson Y. Lee, Abheesht Sharma, Andrea Santilli, Antoine Chaffin, Arnaud Stiegler, Debajyoti Datta, Eliza Szczechla, Gunjan Chhablani, Han Wang, Harshit Pandey, Hendrik Strobelt, Jason Alan Fries, Jos Rozen, Leo Gao, Lintang Sutawika, M Saiful Bari, Maged S. Al-shaibani, Matteo Manica, Nihal Nayak, Ryan Teehan, Samuel Albanie, Sheng Shen, Srulik Ben-David, Stephen H. Bach, Taewoon Kim, Tali Bers, Thibault Fevry, Trishala Neeraj, Urmish Thakker, Vikas Raunak, Xiangru Tang, Zheng-Xin Yong, Zhiqing Sun, Shaked Brody, Yallow Uri, Hadar Tojarieh, Adam Roberts, Hyung Won Chung, Jaesung Tae, Jason Phang, Ofir Press, Conglong Li, Deepak Narayanan, Hatim Bourfoune, Jared Casper, Jeff Rasley, Max Ryabinin, Mayank Mishra, Minjia Zhang, Mohammad Shoeybi, Myriam Peyrounette, Nicolas Patry, Nouamane Tazi, Omar Sanseviero, Patrick von Platen, Pierre Cornette, Pierre Fran¸cois Lavall´ee, R´emi Lacroix, Samyam Rajbhandari, Sanchit Gandhi, Shaden Smith, St´ephane Requena, Suraj Patil, Tim Dettmers, Ahmed Baruwa, Amanpreet Singh, Anastasia Cheveleva, Anne-Laure Ligozat, Arjun Subramonian, Aur´elie N´ev´eol, Charles Lovering, Dan Garrette, Deepak Tunuguntla, Ehud Reiter, Ekaterina Taktasheva, Ekaterina Voloshina, Eli Bogdanov, Genta Indra Winata, Hailey Schoelkopf, Jan-Christoph Kalo, Jekaterina Novikova, Jessica Zosa Forde, Jordan Clive, Jungo Kasai, Ken Kawamura, Liam Hazan, Marine Carpuat, Miruna Clinciu, Najoung Kim, Newton Cheng, Oleg Serikov, Omer Antverg, Oskar van der Wal, Rui Zhang, Ruochen Zhang, Sebastian Gehrmann, Shachar Mirkin, Shani Pais, Tatiana Shavrina, Thomas Scialom, Tian Yun, Tomasz Limisiewicz, Verena Rieser, Vitaly Protasov, Vladislav Mikhailov, Yada Pruksachatkun, Yonatan Belinkov, Zachary Bamberger, Zdenˇek Kasner, Alice Rueda, Amanda Pestana, Amir Feizpour, Ammar Khan, Amy Faranak, Ana Santos, Anthony Hevia, Antigona Unldreaj, Arash Aghagol, Arezoo Abdollahi, Aycha Tammour, Azadeh HajiHosseini, Bahareh Behroozi, Benjamin Ajibade, Bharat Saxena, Carlos Mu˜noz Ferrandis, Danish Contractor, David Lansky, Davis David, Douwe Kiela, Duong A. Nguyen, Edward Tan, Emi Baylor, Ezinwanne Ozoani, Fatima Mirza, Frankline Ononiwu, Habib Rezanejad, Hessie Jones, Indrani Bhattacharya, Irene Solaiman, Irina Sedenko, Isar Nejadgholi, Jesse Passmore, Josh Seltzer, Julio Bonis Sanz, Livia Dutra, Mairon Samagaio, Maraim Elbadri, Margot Mieskes, Marissa Gerchick, Martha Akinlolu, Michael McKenna, Mike Qiu, Muhammed Ghauri, Mykola Burynok, Nafis Abrar, Nazneen Rajani, Nour Elkott, Nour Fahmy, Olanrewaju Samuel, Ran An, Rasmus Kromann, Ryan Hao, Samira Alizadeh, Sarmad Shubber, Silas Wang, Sourav Roy, Sylvain Viguier, Thanh Le, Tobi Oyebade, Trieu Le, Yoyo Yang, Zach Nguyen, Abhinav Ramesh Kashyap, Alfredo Palasciano, Alison Callahan, Anima Shukla, Antonio Miranda-Escalada, Ayush Singh, Benjamin Beilharz, Bo Wang, Caio Brito, Chenxi Zhou, Chirag Jain, Chuxin Xu, Cl´ementine Fourrier, Daniel Le´on Peri˜n´an, Daniel Molano, Dian Yu, Enrique Manjavacas, Fabio Barth, Florian Fuhrimann, Gabriel Altay, Giyaseddin Bayrak, Gully Burns, Helena U. Vrabec, Imane Bello, Ishani Dash, Jihyun Kang, John Giorgi, Jonas Golde, Jose David Posada, Karthik Rangasai Sivaraman, Lokesh Bulchandani, Lu Liu, Luisa Shinzato, Madeleine Hahn de Bykhovetz, Maiko Takeuchi, Marc P`amies, Maria A Castillo, Marianna Nezhurina, Mario S¨anger, Matthias Samwald, Michael Cullan, Michael Weinberg, Michiel De Wolf, Mina Mihaljcic, Minna Liu, Moritz Freidank, Myungsun Kang, Natasha Seelam, Nathan Dahlberg, Nicholas Michio Broad, Nikolaus Muellner, Pascale Fung, Patrick Haller, Ramya Chandrasekhar, Renata Eisenberg, Robert Martin, Rodrigo Canalli, Rosaline Su, Ruisi Su, Samuel Cahyawijaya, Samuele Garda, Shlok S Deshmukh, Shubhanshu Mishra, Sid Kiblawi, Simon Ott, Sinee Sang-aroonsiri, Srishti Kumar, Stefan Schweter, Sushil Bharati, Tanmay Laud, Th´eo Gigant, Tomoya Kainuma, Wojciech Kusa, Yanis Labrak, Yash Shailesh Bajaj, Yash Venkatraman, Yifan Xu, Yingxin Xu, Yu Xu, Zhe Tan, Zhongli Xie, Zifan Ye, Mathilde Bras, Younes Belkada, and Thomas Wolf. Bloom: A 176b-parameter open-access multilingual language model. arXiV, 11 2022. URL http://arxiv.org/abs/2211.05100.
  98. 98.Mike Schuster and Kaisuke Nakajima. Japanese and korean voice search. In 2012 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5149–5152. IEEE, 2012.
  99. 99.Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/P16-1162. URL https://aclanthology.org/P16-1162.
  100. 100.Raj Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah, Wendi Du, Sudheer Chava, Natraj Raman, Charese Smiley, Jiaao Chen, and Diyi Yang. When FLUE meets FLANG: Benchmarks and large pretrained language model for financial domain. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2322–2335, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.148.
  101. 101.Noam Shazeer. GLU variants improve transformer. arXiV preprint arXiV:2002.05202, 2020.
  102. 102.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiV preprint arXiV:1909.08053, 2019.
  103. 103.Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Scharli, Aakanksha Chowdhery, Philip Mansfield, Blaise Aguera y Arcas, Dale Webster, Greg S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle Barral, Christopher Semturs, Alan Karthikesalingam, and Vivek Natarajan. Large language models encode clinical knowledge, 2022. URL https://arxiv.org/abs/2212.13138.
  104. 104.Ankur Sinha and Tanmay Khandait. Impact of news on the commodity market: Dataset and results. CoRR, abs/2009.04202, 2020. URL https://arxiv.org/abs/2009.04202.
  105. 105.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model, 2022. URL https://arxiv.org/abs/2201.11990.
  106. 106.Saleh Soltan, Shankar Ananthakrishnan, Jack G. M. FitzGerald, Rahul Gupta, Wael Hamza, Haidar Khan, Charith S. Peris, Stephen Rawls, Andrew Rosenbaum, Anna Rumshisky, Chandan Prakash, Mukund Sridhar, Fabian Triefenbach, Apurv Verma, Gokhan Tur, and Premkumar Natarajan. Alexatm 20b: Few-shot learning using a large-scale multilingual seq2seq model. arXiV, abs/2208.01448, 2022.
  107. 107.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adria Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ameet Annasaheb Rahane, Anantharaman S. Iyer, Anders Andreassen, Andrea Santilli, Andreas Stuhlmuller, Andrew M. Dai, Andrew D. La, Andrew Kyle Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karakacs, Bridget R. Roberts, Bao Sheng Loe, Barret Zoph, Bartlomiej Bojanowski, Batuhan Ozyurt, Behnam Hedayatnia, Behnam Neyshabur, Benjamin Inden, Benno Stein, Berk Ekmekci, Bill Yuchen Lin, Blake Stephen Howald, Cameron Diao, Cameron Dour, Catherine Stinson, Cedrick Argueta, C’esar Ferri Ram’irez, Chandan Singh, Charles Rathkopf, Chenlin Meng, Chitta Baral, Chiyu Wu, Chris Callison-Burch, Chris Waites, Christian Voigt, Christopher D. Manning, Christopher Potts, Cindy Tatiana Ramirez, Clara Rivera, Clemencia Siro, Colin Raffel, Courtney Ashcraft, Cristina Garbacea, Damien Sileo, Daniel H Garrette, Dan Hendrycks, Dan Kilman, Dan Roth, Daniel Freeman, Daniel Khashabi, Daniel Levy, Daniel Gonz’alez, Danny Hernandez, Danqi Chen, Daphne Ippolito, Dar Gilboa, David Dohan, D. Drakard, David Jurgens, Debajyoti Datta, Deep Ganguli, Denis Emelin, Denis Kleyko, Deniz Yuret, Derek Chen, Derek Tam, Dieuwke Hupkes, Diganta Misra, Dilyar Buzan, Dimitri Coelho Mollo, Diyi Yang, Dong-Ho Lee, Ekaterina Shutova, Ekin Dogus Cubuk, Elad Segal, Eleanor Hagerman, Elizabeth Barnes, Elizabeth P. Donoway, Ellie Pavlick, Emanuele Rodola, Emma FC Lam, Eric Chu, Eric Tang, Erkut Erdem, Ernie Chang, Ethan A. Chi, Ethan Dyer, Ethan J. Jerzak, Ethan Kim, Eunice Engefu Manyasi, Evgenii Zheltonozhskii, Fan Xia, Fatemeh Siar, Fernando Mart’inez-Plumed, Francesca Happ’e, Fran¸cois Chollet, Frieda Rong, Gaurav Mishra, Genta Indra Winata, Gerard de Melo, Germ´an Kruszewski, Giambattista Parascandolo, Giorgio Mariani, Gloria Wang, Gonzalo Jaimovitch-L’opez, Gregor Betz, Guy Gur-Ari, Hana Galijasevic, Han Sol Kim, Hannah Rashkin, Hanna Hajishirzi, Harsh Mehta, Hayden Bogar, Henry Shevlin, Hinrich Sch¨utze, Hiromu Yakura, Hongming Zhang, Hubert Wong, Ian Aik-Soon Ng, Isaac Noble, Jaap Jumelet, Jack Geissinger, John Kernion, Jacob Hilton, Jaehoon Lee, Jaime Fern´andez Fisac, J. Brooker Simon, James Koppel, James Zheng, James Zou, Jan Koco’n, Jana Thompson, Jared Kaplan, Jarema Radom, Jascha Narain Sohl-Dickstein, Jason Phang, Jason Wei, Jason Yosinski, Jekaterina Novikova, Jelle Bosscher, Jenni Marsh, Jeremy Kim, Jeroen Taal, Jesse Engel, Jesujoba Oluwadara Alabi, Jiacheng Xu, Jiaming Song, Jillian Tang, Jane W Waweru, John Burden, John Miller, John U. Balis, Jonathan Berant, Jorg Frohberg, Jos Rozen, Jos´e Hern´andez-Orallo, Joseph Boudeman, Joseph Jones, Joshua B. Tenenbaum, Joshua S. Rule, Joyce Chua, Kamil Kanclerz, Karen Livescu, Karl Krauth, Karthik Gopalakrishnan, Katerina Ignatyeva, Katja Markert, Kaustubh D. Dhole, Kevin Gimpel, Kevin Ochieng’ Omondi, Kory Wallace Mathewson, Kristen Chiafullo, Ksenia Shkaruta, Kumar Shridhar, Kyle McDonell, Kyle Richardson, Laria Reynolds, Leo Gao, Li Zhang, Liam Dugan, Lianhui Qin, Lidia Contreras-Ochando, Louis-Philippe Morency, Luca Moschella, Luca Lam, Lucy Noble, Ludwig Schmidt, Luheng He, Luis Oliveros Col’on, Luke Metz, Lutfi Kerem cSenel, Maarten Bosma, Maarten Sap, Maartje ter Hoeve, Madotto Andrea, Maheen Saleem Farooqi, Manaal Faruqui, Mantas Mazeika, Marco Baturan, Marco Marelli, Marco Maru, M Quintana, Marie Tolkiehn, Mario Giulianelli, Martha Lewis, Martin Potthast, Matthew Leavitt, Matthias Hagen, M’aty’as Schubert, Medina Baitemirova, Melissa Arnaud, Melvin Andrew McElrath, Michael A. Yee, Michael Cohen, Mi Gu, Michael I. Ivanitskiy, Michael Starritt, Michael Strube, Michal Swkedrowski, Michele Bevilacqua, Michihiro Yasunaga, Mihir Kale, Mike Cain, Mimee Xu, Mirac Suzgun, Monica Tiwari, Mohit Bansal, Moin Aminnaseri, Mor Geva, Mozhdeh Gheini, T MukundVarma, Nanyun Peng, Nathan Chi, Nayeon Lee, Neta Gur-Ari Krakover, Nicholas Cameron, Nicholas S. Roberts, Nicholas Doiron, Nikita Nangia, Niklas Deckers, Niklas Muennighoff, Nitish Shirish Keskar, Niveditha Iyer, Noah Constant, Noah Fiedel, Nuan Wen, Oliver Zhang, Omar Agha, Omar Elbaghdadi, Omer Levy, Owain Evans, Pablo Antonio Moreno Casares, Parth Doshi, Pascale Fung, Paul Pu Liang, Paul Vicol, Pegah Alipoormolabashi, Peiyuan Liao, Percy Liang, Peter W. Chang, Peter Eckersley, Phu Mon Htut, Pi-Bei Hwang, P. Milkowski, Piyush S. Patil, Pouya Pezeshkpour, Priti Oli, Qiaozhu Mei, QING LYU, Qinlang Chen, Rabin Banjade, Rachel Etta Rudolph, Raefer Gabriel, Rahel Habacker, Ram’on Risco Delgado, Rapha¨el Milli`ere, Rhythm Garg, Richard Barnes, Rif A. Saurous, Riku Arakawa, Robbe Raymaekers, Robert Frank, Rohan Sikand, Roman Novak, Roman Sitelew, Ronan Le Bras, Rosanne Liu, Rowan Jacobs, Rui Zhang, Ruslan Salakhutdinov, Ryan Chi, Ryan Lee, Ryan Stovall, Ryan Teehan, Rylan Yang, Sahib J. Singh, Saif M. Mohammad, Sajant Anand, Sam Dillavou, Sam Shleifer, Sam Wiseman, Samuel Gruetter, Sam Bowman, Samuel S. Schoenholz, Sanghyun Han, Sanjeev Kwatra, Sarah A. Rous, Sarik Ghazarian, Sayan Ghosh, Sean Casey, Sebastian Bischoff, Sebastian Gehrmann, Sebastian Schuster, Sepideh Sadeghi, Shadi S. Hamdan, Sharon Zhou, Shashank Srivastava, Sherry Shi, Shikhar Singh, Shima Asaadi, Shixiang Shane Gu, Shubh Pachchigar, Shubham Toshniwal, Shyam Upadhyay, Shyamolima Debnath, Siamak Shakeri, Simon Thormeyer, Simone Melzi, Siva Reddy, Sneha Priscilla Makini, Soo hwan Lee, Spencer Bradley Torene, Sriharsha Hatwar, Stanislas Dehaene, Stefan Divic, Stefano Ermon, Stella Rose Biderman, Stephanie C. Lin, S. Prasad, Steven T. Piantadosi, Stuart M. Shieber, Summer Misherghi, Svetlana Kiritchenko, Swaroop Mishra, Tal Linzen, Tal Schuster, Tao Li, Tao Yu, Tariq A. Ali, Tatsuo Hashimoto, Te-Lin Wu, Theo Desbordes, Theodore Rothschild, Thomas Phan, Tianle Wang, Tiberius Nkinyili, Timo Schick, T. N. Kornev, Timothy Telleen-Lawton, Titus Tunduny, Tobias Gerstenberg, Trenton Chang, Trishala Neeraj, Tushar Khot, Tyler O’Brien Shultz, Uri Shaham, Vedant Misra, Vera Demberg, Victoria Nyamai, Vikas Raunak, Vinay Venkatesh Ramasesh, Vinay Uday Prabhu, Vishakh Padmakumar, Vivek Srikumar, William Fedus, William Saunders, William Zhang, W Vossen, Xiang Ren, Xiaoyu Tong, Xinyi Wu, Xudong Shen, Yadollah Yaghoobzadeh, Yair Lakretz, Yang Song, Yasaman Bahri, Ye Ji Choi, Yichi Yang, Yiding Hao, Yifu Chen, Yonatan Belinkov, Yu Hou, Yu Hou, Yuntao Bai, Zachary Seid, Zhao Xinran, Zhuoye Zhao, Zi Fu Wang, Zijie J. Wang, Zirui Wang, Ziyi Wu, Sahib Singh, and Uri Shaham. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv, abs/2206.04615, 2022.
  108. 108.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. arXiV preprint arXiV:2104.09864, 2021a.
  109. 109.Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. CoRR, abs/2104.09864, 2021b. URL https://arxiv.org/abs/2104.09864.
  110. 110.Ilya Sutskever, James Martens, and Geoffrey E Hinton. Generating text with recurrent neural networks. In Proceedings of the 28th international conference on machine learning (ICML-11), pages 1017–1024, 2011.
  111. 111.Mirac Suzgun, Nathan Scales, Nathanael Sch¨arli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. CoRR, abs/2210.09261, 2022. doi: 10.48550/arXiv.2210.09261. URL https://doi.org/10.48550/arXiv.2210.09261.
  112. 112.Oyvind Tafjord and Peter Clark. General-purpose question-answering with macaw. arXiV preprint arXiV:2109.02593, 2021.
  113. 113.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1421. URL https://aclanthology.org/N19-1421.
  114. 114.Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. Scale efficiently: Insights from pre-training and fine-tuning transformers. arXiV preprint arXiV:2109.10686, 2021.
  115. 115.Yi Tay, Mostafa Dehghani, Samira Abnar, Hyung Won Chung, William Fedus, Jinfeng Rao, Sharan Narang, Vinh Q Tran, Dani Yogatama, and Donald Metzler. Scaling laws vs model architectures: How does inductive bias influence scaling? arXiV preprint arXiV:2207.10551, 2022a.
  116. 116.Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Siamak Shakeri, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. Ul2: Unifying language learning paradigms, 2022b. URL https://arxiv.org/abs/2205.05131.
  117. 117.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. Galactica: A large language model for science. arXiV, 11 2022. URL http://arxiv.org/abs/2211.09085.
  118. 118.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. Lamda: Language models for dialog applications, 2022. URL https://arxiv.org/abs/2201.08239.
  119. 119.Erik F. Tjong Kim Sang and Fien De Meulder. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pages 142–147, 2003. URL https://aclanthology.org/W03-0419.
  120. 120.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: Open and efficient foundation language models, 2023. URL https://arxiv.org/abs/2302.13971.
  121. 121.Tina Tseng, Amanda Stent, and Domenic Maida. Best practices for managing data annotation projects, 2020. URL http://rgdoi.net/10.13140/RG.2.2.34497.58727.
  122. 122.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, L ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
  123. 123.Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  124. 124.Yue Wang, Weishi Wang, Shafiq Joty, and Steven C.H. Hoi. CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8696–8708, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.685. URL https://aclanthology.org/2021.emnlp-main.685.
  125. 125.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. Finetuned language models are zero-shot learners, 2021. URL https://arxiv.org/abs/2109.01652.
  126. 126.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models. Transactions on Machine Learning Research (TMLR), 2022a. doi: 10.48550/ARXIV.2206.07682. URL https://arxiv.org/abs/2206.07682.
  127. 127.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022b. URL https://openreview.net/forum?id=_VjQlMeSB_J.
  128. 128.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel. Ethical and social risks of harm from language models, 2021. URL https://arxiv.org/abs/2112.04359.
  129. 129.Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John F. J. Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sande Minnich Brown, Zachary Kenton, William T. Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William S. Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel. Taxonomy of risks posed by language models. 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022.
  130. 130.Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. Challenges in detoxifying language models. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2447–2469, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp.210. URL https://aclanthology.org/2021.findings-emnlp.210.
  131. 131.Shijie Wu and Mark Dredze. Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 833–844, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1077. URL https://aclanthology.org/D19-1077.
  132. 132.Yonghui Wu, Mike Schuster, Z. Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason R. Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Gregory S. Corrado, Macduff Hughes, and Jeffrey Dean. Google’s neural machine translation system: Bridging the gap between human and machine translation. ArXiV, abs/1609.08144, 2016.
  133. 133.Yijia Xiao, Jiezhong Qiu, Ziang Li, Chang-Yu Hsieh, and Jie Tang. Modeling protein using large-scale pretrain language model. CoRR, abs/2108.07435, 2021. URL https://arxiv.org/abs/2108.07435.
  134. 134.Frank Z Xing, Erik Cambria, and Roy E Welsch. Natural language based financial forecasting: a survey. Artificial Intelligence Review, 50(1):49–73, 2018.
  135. 135.Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein. Detoxifying language models risks marginalizing minority voices. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2390–2397, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.190. URL https://aclanthology.org/2021.naacl-main.190.
  136. 136.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1472. URL https://aclanthology.org/P19-1472.
  137. 137.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Peng Zhang, Yuxiao Dong, and Jie Tang. Glm-130b: An open bilingual pre-trained model. arXiV, 10 2022. URL http://arxiv.org/abs/2210.02414.
  138. 138.Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. Record: Bridging the gap between human and machine commonsense reading comprehension. arXiV, abs/1810.12885, 2018.
  139. 139.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. Opt: Open pre-trained transformer language models. arXiV, 5 2022a. URL http://arxiv.org/abs/2205.01068.
  140. 140.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 270–278, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-demos.30. URL https://aclanthology.org/2020.acl-demos.30.
  141. 141.Zhen Zhang, Shuai Zheng, Yida Wang, Justin Chiu, George Karypis, Trishul Chilimbi, Mu Li, and Xin Jin. Mics: Near-linear scaling for training gigantic model on public cloud, 2022b. URL https://arxiv.org/abs/2205.00119.

Citation

MLA
Wu, S., et al. “BloombergGPT: A Large Language Model for Finance”. arXiv, 2023, http://arxiv.org/abs/2303.17564v3.
APA
Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., & Mann, G. (2023). BloombergGPT: A Large Language Model for Finance. arXiv. http://arxiv.org/abs/2303.17564v3
Chicago
Wu, S., O. Irsoy, S. Lu, et al. 2023. “BloombergGPT: A Large Language Model for Finance”. arXiv. http://arxiv.org/abs/2303.17564v3.
Harvard
Wu, S. et al. (2023) “BloombergGPT: A Large Language Model for Finance”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.17564v3.
Vancouver
1. Wu S, Irsoy O, Lu S, Dabravolski V, Dredze M, Gehrmann S, Kambadur P, Rosenberg D, Mann G (2023) BloombergGPT: A Large Language Model for Finance. arXiv

BibTeX

@article{wu2023bloomberggpt,
  title = {BloombergGPT: A Large Language Model for Finance},
  author = {Wu, Shijie and Irsoy, Ozan and Lu, Steven and Dabravolski, Vadim and Dredze, Mark and Gehrmann, Sebastian and Kambadur, Prabhanjan and Rosenberg, David and Mann, Gideon},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.17564v3},
  eprint = {2303.17564}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors