Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length

Xuezhe MaXiaomeng YangWenhan XiongBeidi ChenLili YuHao ZhangJonathan MayLuke ZettlemoyerOmer LevyChunting Zhou

article2024NeurIPS63 citations

Introduces MEGALODON, a linear-complexity sequence architecture combining complex exponential moving averages with normalized gated attention that outperforms standard Transformer models like LLAMA2 at the 7-billion parameter scale across 2 trillion training tokens while scaling efficiently to virtually unlimited context lengths.

Listen

Modern artificial intelligence applications, such as analyzing extensive documents, maintaining long conversational context, and processing video, require models that can handle very long sequences of data. Standard Transformer architectures suffer from high computational costs that grow quadratically with sequence length, making them slow and expensive for long-context tasks. While alternative architectures like linear attention and state-space models offer lower computational complexity, they have historically lagged behind standard Transformers in training efficiency and overall task accuracy.

The article introduces and evaluates Megalodon, a neural network architecture designed for sequence modeling with unlimited context length. The primary objective is to demonstrate that Megalodon achieves linear computational and memory scaling while outperforming standard Transformer architectures in both pretraining efficiency and downstream accuracy.

To establish credibility under rigorous, controlled conditions, the researchers evaluated Megalodon through a head-to-head comparison against the standard Llama 2 model architecture. Both models were scaled to 7 billion parameters and trained on an identical dataset of 2 trillion tokens across 256 graphics processing units. Megalodon builds on the Moving Average Equipped Gated Attention framework by introducing complex exponential moving averages, a specialized timestep normalization technique for step-by-step sequence data, normalized attention, and an updated residual connection scheme to ensure stability during large-scale training. The authors also tested Megalodon on context lengths extending up to 2 million tokens, as well as on various standard benchmarks spanning text, speech, and image processing.

The evaluation yielded several key findings. First, Megalodon demonstrated superior training and data efficiency, reaching an overall training loss of 1.70, which significantly outperforms the 1.75 achieved by the 7-billion parameter Llama 2 and approaches the 1.67 achieved by the larger 13-billion parameter Llama 2 model. Second, Megalodon delivered major computational speedups at long context lengths: while roughly 6% slower than Llama 2 on short sequences of 4,000 tokens, it was approximately 32% faster when trained on 32,000-token sequences. Third, Megalodon consistently outperformed the 7-billion parameter Llama 2 on standard academic benchmarks and competitive long-context question-answering datasets. Fourth, tests extending sequence lengths up to 2 million tokens showed monotonic improvements in prediction performance, confirming effective long-context modeling. Finally, smaller-scale tests in raw audio and image classification confirmed that the architecture generalizes effectively across different data types.

These findings suggest that organizations can achieve higher model quality at reduced computational expense for long-context workloads. By lowering the processing overhead from quadratic to linear complexity, Megalodon reduces hardware costs and training timelines for processing extensive data sequences. Moreover, its ability to match or exceed the performance of a standard 13-billion parameter model while using only 7 billion parameters enables substantial operational efficiencies in deployment.

Based on these results, decision-makers considering long-context language systems should consider evaluating chunked linear attention architectures as an alternative to standard Transformer configurations. For teams managing production deployments, the main trade-off lies in sequence length: standard architectures remain slightly faster for short sequences under 4,000 tokens, but Megalodon provides clear computational and quality advantages for workloads that require long context windows. Future work should focus on validating the architecture at larger parameter scales and applying it to comprehensive multi-modal pretraining.

While confidence in these findings is high due to the strictly controlled 7-billion parameter pretraining setup, some limitations remain. The large-scale pretraining evaluation was limited to a single 7-billion parameter model size, and comparisons against certain external models were constrained by differing pretraining data volumes. Stakeholders should conduct pilot evaluations within their specific task pipelines before making full-scale infrastructure transitions.

arXiv: 2404.08801
Cover for Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length

Abstract

The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and state space models exist, they empirically underperform Transformers in pretraining efficiency and downstream task accuracy. We introduce MEGALODON, an neural architecture for efficient sequence modeling with unlimited context length. MEGALODON inherits the architecture of MEGA (exponential moving average with gated attention), and further introduces multiple technical components to improve its capability and stability, including complex exponential moving average (CEMA), timestep normalization layer, normalized attention mechanism and pre-norm with two-hop residual configuration. In a controlled head-to-head comparison with LLAMA2, MEGALODON achieves better efficiency than Transformer in the scale of 7 billion parameters and 2 trillion training tokens. MEGALODON reaches a training loss of 1.70, landing mid-way between LLAMA2-7B (1.75) and 13B (1.67). The improvements of MEGALODON over Transformers are robust throughout a range of benchmarks across different tasks and modalities. Code: https://github.com/XuezheMax/megalodon

Knowls

  1. Knowl 1 — Complex Exponential Moving Average (CEMA)

    model/method

    The Complex Exponential Moving Average (CEMA) extends the multi-dimensional damped exponential moving average (EMA) into the complex plane C\mathbb{C} to enhance the expressive capacity of diagonal state space layers in sequence modeling while maintaining linear computational complexity.

    Given an input sequence X=(x1,x2,…,xn)∈Rn×dX = (x_1, x_2, \dots, x_n) \in \mathbb{R}^{n \times d}, each feature dimension j∈{1,…,d}j \in \{1, \dots, d\} at timestep tt is expanded to an hh-dimensional hidden representation via an expansion vector βj∈Rh\beta_j \in \mathbb{R}^h: ut(j)=βjxt,j∈Rhu_t^{(j)} = \beta_j x_{t,j} \in \mathbb{R}^h

    The complex hidden state ht(j)∈Chh_t^{(j)} \in \mathbb{C}^h and the output scalar yt,j∈Ry_{t,j} \in \mathbb{R} are updated recursively according to: ht(j)=αj(cos⁡θj+isin⁡θj)⊙ut(j)+(1−αj⊙δj)(cos⁡θj+isin⁡θj)⊙ht−1(j)h_t^{(j)} = \alpha_j (\cos \theta_j + i\sin \theta_j) \odot u_t^{(j)} + (1 - \alpha_j \odot \delta_j)(\cos \theta_j + i\sin \theta_j) \odot h_{t-1}^{(j)} yt,j=Re⁡(ηj⊤ht(j))y_{t,j} = \operatorname{Re}(\eta_j^\top h_t^{(j)}) where:

    • α∈(0,1)d×h\alpha \in (0, 1)^{d \times h} is a learnable decay factor tensor.
    • δ∈(0,1)d×h\delta \in (0, 1)^{d \times h} is a learnable damping factor tensor.
    • η∈Cd×h\eta \in \mathbb{C}^{d \times h} is a complex projection matrix mapping the hh-dimensional hidden state back to a real scalar.
    • ⊙\odot denotes the element-wise Hadamard product.
    • Re⁡(⋅)\operatorname{Re}(\cdot) extracts the real part of a complex vector.
    • θj∈Rh\theta_j \in \mathbb{R}^h is an argument vector whose kk-th component (k∈{1,…,h}k \in \{1, \dots, h\}) is parameterized using a learnable base angle ωj∈R\omega_j \in \mathbb{R} as: θj,k=2πkhωj\theta_{j,k} = \frac{2\pi k}{h} \omega_j

    Decaying the magnitude of ht(j)h_t^{(j)} via the factor (1−αj⊙δj)(1 - \alpha_j \odot \delta_j) while rotating its phase via (cos⁡θj+isin⁡θj)(\cos \theta_j + i\sin \theta_j) preserves a decaying convolutional kernel structure over long sequences.

  2. Knowl 2 — Timestep Normalization for Autoregressive Sequence Modeling

    model/method

    Timestep Normalization extends group normalization to autoregressive sequence models by calculating cumulative statistics exclusively over historical and current timesteps, thereby preventing the leakage of future token information during training and inference.

    Given an input sequence X=(x1,x2,…,xn)∈Rn×dX = (x_1, x_2, \dots, x_n) \in \mathbb{R}^{n \times d} partitioned along the feature dimension into kk groups of size dg=d/kd_g = d / k, the cumulative mean μt\mu_t and cumulative variance σt2\sigma_t^2 for a group at timestep t∈{1,…,n}t \in \{1, \dots, n\} are computed across all timesteps from i=1i = 1 to tt and all within-group feature indices j=1j = 1 to dgd_g: μt=1t⋅dg∑i=1t∑j=1dgxi,j\mu_t = \frac{1}{t \cdot d_g} \sum_{i=1}^t \sum_{j=1}^{d_g} x_{i,j} σt2=1t⋅dg∑i=1t∑j=1dg(xi,j−μt)2\sigma_t^2 = \frac{1}{t \cdot d_g} \sum_{i=1}^t \sum_{j=1}^{d_g} (x_{i,j} - \mu_t)^2

    In GPU implementation, cumulative means and variances are accumulated numerically using Welford's algorithm combined with Kahan summation across threads partitioned along sequence and feature dimensions. In the MEGALODON architecture, Timestep Normalization is applied immediately prior to the attention layer to mitigate internal covariate shift along the temporal dimension.

  3. Knowl 3 — Normalized Attention Mechanism

    model/method

    The normalized attention mechanism stabilizes attention matrix computation by applying L2L_2 normalization directly to the shared contextual representation prior to generating query and key vectors.

    Given an input sequence X∈Rn×dX \in \mathbb{R}^{n \times d}, the sequence is first transformed by Complex Exponential Moving Average (CEMA) to yield contextual representation X′=CEMA(X)∈Rn×dX' = \text{CEMA}(X) \in \mathbb{R}^{n \times d}. A shared representation ZZ and its normalized counterpart Z′Z' are defined as: Z=X′Wz+bz∈Rn×zZ = X' W_z + b_z \in \mathbb{R}^{n \times z} Z′=Z∥Z∥∈Rn×zZ' = \frac{Z}{\|Z\|} \in \mathbb{R}^{n \times z} where Wz∈Rd×zW_z \in \mathbb{R}^{d \times z} and bz∈Rzb_z \in \mathbb{R}^z are learnable projection weights and biases, and ∥⋅∥\|\cdot\| is the L2L_2 norm computed along the representation dimension zz.

    Queries QQ, keys KK, and values VV are computed via: Q=κq⊙Z′+μq∈Rn×zQ = \kappa_q \odot Z' + \mu_q \in \mathbb{R}^{n \times z} K=κk⊙Z′+μk∈Rn×zK = \kappa_k \odot Z' + \mu_k \in \mathbb{R}^{n \times z} V=ϕsilu(XWv+bv)∈Rn×vV = \phi_{\text{silu}}(X W_v + b_v) \in \mathbb{R}^{n \times v} where κq,μq,κk,μk∈Rz\kappa_q, \mu_q, \kappa_k, \mu_k \in \mathbb{R}^z are learnable per-dimension scale and bias parameters, Wv∈Rd×vW_v \in \mathbb{R}^{d \times v} and bv∈Rvb_v \in \mathbb{R}^v are linear projection parameters, and ϕsilu\phi_{\text{silu}} is the SiLU activation function.

    The attention output O∈Rn×vO \in \mathbb{R}^{n \times v} is computed with standard softmax attention without requiring an explicit temperature scaling function τ(X)\tau(X): O=softmax⁡(QK⊤)VO = \operatorname{softmax}(Q K^\top) V

  4. Knowl 4 — Pre-Norm with Two-Hop Residual Connection

    model/method

    Pre-norm with two-hop residual connection reorganizes residual skip connections within each transformer block to prevent the unbounded growth of output variance across deep layers without requiring auxiliary gating networks.

    Standard pre-normalization accumulates the outputs of successive sub-layers into the residual stream via: Y^=Attention⁡(Norm⁡(X))+X\hat{Y} = \operatorname{Attention}(\operatorname{Norm}(X)) + X Y=FFN⁡(Norm⁡(Y^))+Y^=FFN⁡(Norm⁡(Y^))+Attention⁡(Norm⁡(X))+XY = \operatorname{FFN}(\operatorname{Norm}(\hat{Y})) + \hat{Y} = \operatorname{FFN}(\operatorname{Norm}(\hat{Y})) + \operatorname{Attention}(\operatorname{Norm}(X)) + X which causes the variance of YY to increase progressively with network depth.

    The pre-norm with two-hop residual configuration replaces the intermediate skip connection Y^\hat{Y} with the original block input XX across both sub-layers: Y^=Attention⁡(TimestepNorm⁡(X))+X\hat{Y} = \operatorname{Attention}(\operatorname{TimestepNorm}(X)) + X Y=FFN⁡(LayerNorm⁡(Y^))+XY = \operatorname{FFN}(\operatorname{LayerNorm}(\hat{Y})) + X

    Timestep Normalization is applied prior to the attention layer to normalize temporal representations, while standard Layer Normalization is applied prior to the feed-forward network (FFN) layer. Reusing the input XX as the direct residual for the FFN layer eliminates the parameter overhead of gated residual update connections while ensuring stable training at large scale.

  5. Knowl 5 — Plus-One Reparameterization for Normalization Layers

    model/method

    In neural network normalization layers (such as Layer Normalization or Timestep Normalization), the normalized tensor x−μσ\frac{x - \mu}{\sigma} is scaled by a learnable parameter γ\gamma and shifted by β\beta: y=γx−μσ+βy = \gamma \frac{x - \mu}{\sigma} + \beta When γ\gamma is initialized to 1 and trained with L2L_2 weight decay regularization, the weight decay penalty drives γ\gamma toward 0 over time, pulling the scale away from 1 and destabilizing training dynamics.

    The plus-one reparameterization modifies the transformation to: y=(γ+1)x−μσ+βy = (\gamma + 1) \frac{x - \mu}{\sigma} + \beta where γ\gamma is initialized to 0 and β\beta is initialized to 0. Under L2L_2 weight decay, the regularization penalty keeps γ\gamma centered around 0, ensuring that the effective scale factor γ+1\gamma + 1 remains stable around 1 throughout training.

  6. Knowl 6 — 4-Dimensional Parallelism for Chunk-Wise Attention

    model/method

    MEGALODON incorporates chunk parallelism along the temporal/sequence dimension alongside standard 3D distributed training parallelism (data, tensor, and pipeline parallelism).

    Because MEGALODON computes self-attention independently within local chunks of length cc (e.g., c=4096c = 4096), cross-chunk temporal dependencies are mediated exclusively by two boundary state vectors per block:

    1. The final complex hidden state ht∈Cd×hh_t \in \mathbb{C}^{d \times h} of the Complex Exponential Moving Average (CEMA) layer.
    2. The cumulative running mean μt\mu_t and running variance σt2\sigma_t^2 of the Timestep Normalization layer.

    In distributed training across devices assigned to consecutive sequence chunks, only these boundary state vectors are communicated across devices within each chunk-parallel group. Using asynchronous communication, this inter-device transfer is hidden behind the local matrix multiplications of other components in the same or adjacent blocks.

  7. Knowl 7 — Pretraining Efficiency and Convergence of MEGALODON-7B

    empirical result

    A 7-billion parameter MEGALODON model (32 blocks, hidden dimension d=4096d = 4096, 4 attention heads, chunk size c=4096c = 4096, pretraining context length 32K tokens, SwiGLU activation, RoPE base θ=100,000\theta = 100{,}000) trained on 2 trillion tokens across 256 NVIDIA A100 GPUs demonstrated the following convergence and speed characteristics compared to LLAMA2:

    • Training Loss: MEGALODON-7B reached a final negative log-likelihood (NLL) loss of 1.70 on the 2T token dataset, outperforming LLAMA2-7B (final NLL = 1.75) and approaching LLAMA2-13B (final NLL = 1.67) while exhibiting fewer training loss spikes.
    • Computational Speed: When measuring words/tokens per second (WPS) per device on 256 A100 GPUs with a global batch size of 4M tokens:
      • At 4K context length, MEGALODON-7B was approximately 6% slower than LLAMA2-7B accelerated with FlashAttention-2 (relative speed 1.40x vs. 1.48x).
      • At 32K context length, MEGALODON-7B achieved a relative speed of 1.32x, outperforming LLAMA2-7B at 32K (1.00x) by approximately 32%.
      • MEGALODON-7B at 32K context retained 94% of the operational device throughput of MEGALODON-7B at 4K context via chunk parallelism.
  8. Knowl 8 — Academic Benchmark Downstream Evaluation of MEGALODON-7B

    data/table

    When evaluated on standard zero-shot and few-shot academic benchmarks with short context lengths (<4K< 4\text{K} tokens), MEGALODON-7B pretrained on 2 trillion tokens consistently outperformed LLAMA2-7B and other open-source models of comparable scale trained on similar token budgets.

    Model Size Tokens MMLU BoolQ HellaSw PIQA SIQA WinoG Arc-e Arc-c NQ TQA
    Mamba 3B 0.6T 26.2 71.0 71.0 78.1 – 65.9 68.2 41.7 – –
    RWKV 7B 1.1T – – 70.8 77.3 – 68.4 74.9 46.1 – –
    MPT 7B 1T 26.8 75.0 76.4 80.6 48.5 68.3 70.2 42.6 20.8 50.4
    Mistral 7B – 60.1 83.2 81.3 82.2 47.0 74.2 80.0 54.9 23.2 62.5
    Gemma 8B 6T 64.3 83.2 81.2 81.2 51.8 72.3 81.5 53.2 23.0 63.4
    LLAMA2 13B 2T 54.8 81.7 80.7 80.5 50.3 72.8 77.3 49.4 31.2 65.1
    LLAMA2 7B 2T 45.3 77.4 77.2 78.8 48.3 69.2 75.2 45.9 25.7 58.5
    MEGALODON 7B 2T 49.8 80.5 77.5 80.1 49.6 71.4 79.8 53.1 25.7 60.5

    MEGALODON-7B exceeded LLAMA2-7B across all benchmarks: MMLU (+4.5), BoolQ (+3.1), HellaSwag (+0.3), PIQA (+1.3), SIQA (+1.3), WinoGrande (+2.2), ARC-easy (+4.6), ARC-challenge (+7.2), and TriviaQA (+2.0), while matching NaturalQuestions (25.7). On ARC-easy, ARC-challenge, PIQA, and WinoGrande, MEGALODON-7B surpassed the 13-billion parameter LLAMA2-13B model.

  9. Knowl 9 — Long-Context Extrapolation up to 2 Million Tokens and SCROLLS Evaluation

    empirical result

    MEGALODON achieves context length scaling and long-range question answering without performance degradation on ultra-long sequences:

    • Monotonic Perplexity Reduction up to 2M Context: Evaluated on a validation dataset of 1,920 books containing at least 2 million tokens per sequence, the validation perplexity of MEGALODON-7B decreased monotonically across context lengths spanning 4K, 8K, 16K, 32K, 64K, 128K, 256K, 512K, 1M, and 2M tokens.
    • SCROLLS Benchmark Results: On the SCROLLS long-context benchmark, MEGALODON-7B was evaluated against long-context 7B models on NarrativeQA (0-shot F1), Qasper (2-shot F1), and QMSum (1-shot geometric-ROUGE):
    Model NarrativeQA (F1) Qasper (F1) QMSum (ROUGE)
    Xgen-7B-8K 17.4 20.5 6.8
    MPT-7B-8K 18.8 24.7 8.8
    YaRN-7B-128k 20.9 26.2 11.4
    LLAMA2-7B-4K 18.8 19.8 10.1
    LLAMA2-7B-32K (LLAMA2-L) 23.5 28.3 14.5
    MEGALODON-7B 23.9 28.0 13.1

    MEGALODON-7B outperformed LLAMA2-7B-4K across all tasks and scored higher on NarrativeQA (23.9 vs. 23.5 F1) than LLAMA2-7B-32K (LLAMA2-L), which was continually trained on 500 billion tokens of long-context data.

  10. Knowl 10 — Multimodal and Medium-Scale Benchmark Performance

    data/table

    When evaluated across diverse modalities using a single architecture (softmax attention, rotary position embeddings, two-hop residual pre-norm, and Timestep/Group Normalization), MEGALODON outperformed baseline Transformers and state space models across text, image, and raw speech benchmarks:

    • Long Range Arena (LRA): On sequences of length 1K to 16K, MEGALODON achieved 88.63% average accuracy with full attention and 87.62% with chunked attention (c=128c = 128 or c=4096c = 4096), outperforming MEGA (88.21% full, 85.66% chunked), S4 (85.86%), and standard Transformer (59.24%), narrowing the gap between chunked and full attention to 1.01%.
    • PG-19 Long Document Modeling (1.3B Parameters): MEGALODON-1.3B attained a validation perplexity of 29.5 and test perplexity of 25.4, outperforming Block-Recurrent Transformer (26.5 test PPL), Perceiver AR (28.9 test PPL), and Compressive Transformer (33.6 test PPL).
    • ImageNet-1K Classification (90M Parameters): MEGALODON reached 83.1% Top-1 accuracy, outperforming MEGA (82.3%), DeiT-B (81.8%), and ViT-B (77.9%).
    • Raw Speech Classification (Speech Commands SC10, 300K Parameters): MEGALODON achieved 98.14% accuracy on 16,000-timestep raw audio inputs with chunk size c=1000c = 1000, exceeding S4 (97.50%) and MEGA (96.92%).
    • WikiText-103 Language Modeling (252M Parameters): MEGALODON achieved 17.23 word-level test perplexity with chunk size c=2048c = 2048, outperforming MEGA (18.07), Transformer-XL (18.30), and standard Transformer (18.66).

Coverage note — MT-Bench instruction finetuning results (Table 3) and low-level CUDA implementation details (such as fused DropKey before softmax and shared-memory Cooley-Tukey FFTConv) were omitted as secondary evaluation and implementation details.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  2. 2.Alexei Baevski and Michael Auli. Adaptive input representations for neural language modeling. In International Conference on Learning Representations, 2018.
  3. 3.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, pages 7432–7439, 2020.
  4. 4.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/.
  5. 5.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.
  6. 6.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  7. 7.James W Cooley and John W Tukey. An algorithm for the machine calculation of complex fourier series. Mathematics of computation, 19(90):297–301, 1965.
  8. 8.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2978–2988, 2019.
  9. 9.Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. In International Conference on Learning Representations (ICLR-2024), 2024.
  10. 10.Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. A dataset of information-seeking questions and answers anchored in research papers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-2021), pages 4599–4610, Online, June 2021. Association for Computational Linguistics.
  11. 11.Jared Q Davis, Albert Gu, Krzysztof Choromanski, Tri Dao, Christopher Re, Chelsea Finn, and Percy Liang. Catformer: Designing stable transformers via sensitivity analysis. In International Conference on Machine Learning, pages 2489–2499. PMLR, 2021.
  12. 12.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  13. 13.Daniel Y Fu, Tri Dao, Khaled Kamal Saab, Armin W Thomas, Atri Rudra, and Christopher Re. Hungry hungry hippos: Towards language modeling with state space models. In The Eleventh International Conference on Learning Representations (ICLR-2023), 2023.
  14. 14.Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces, 2023.
  15. 15.Albert Gu, Karan Goel, and Christopher Re. Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations (ICLR-2022), 2022a.
  16. 16.Albert Gu, Ankit Gupta, Karan Goel, and Christopher Re. On the parameterization and initialization of diagonal state space models. arXiv preprint arXiv:2206.11893, 2022b.
  17. 17.Stephen Hanson and Lorien Pratt. Comparing biases for minimal network construction with backpropagation. Advances in neural information processing systems, 1, 1988.
  18. 18.Curtis Hawthorne, Andrew Jaegle, Cătălina Cangea, Sebastian Borgeaud, Charlie Nash, Mateusz Malinowski, Sander Dieleman, Oriol Vinyals, Matthew Botvinick, Ian Simon, et al. General-purpose, long-context autoregressive modeling with perceiver ar. In International Conference on Machine Learning, pages 8535–8558. PMLR, 2022.
  19. 19.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  20. 20.Alex Henry, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen. Query-key normalization for transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4246–4253, 2020.
  21. 21.Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc Le. Transformer quality in linear time. In International Conference on Machine Learning (ICML-2022), pages 9099–9117. PMLR, 2022.
  22. 22.Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and zhifeng Chen. Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019.
  23. 23.J Stuart Hunter. The exponentially weighted moving average. Journal of quality technology, 18(4): 203–210, 1986.
  24. 24.DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, and Behnam Neyshabur. Block-recurrent transformers. Advances in neural information processing systems, 35:33248–33261, 2022.
  25. 25.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning (ICML-2015), pages 448–456. pmlr, 2015.
  26. 26.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  27. 27.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2017.
  28. 28.William Kahan. Pracniques: further remarks on reducing truncation errors. Communications of the ACM, 8(1):40, 1965.
  29. 29.Tomaš Koćisky, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gabor Melis, and Edward Grefenstette. The NarrativeQA reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328, 2018.
  30. 30.Alex Krizhevsky et al. Learning multiple layers of features from tiny images. Technical Report. University of Toronto, 2009.
  31. 31.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466, 2019.
  32. 32.Bonan Li, Yinhan Hu, Xuecheng Nie, Congying Han, Xiangjian Jiang, Tiande Guo, and Luoqi Liu. Dropkey for vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22700–22709, June 2023a.
  33. 33.Dacheng Li, Rulin Shao, Anze Xie, Eric P Xing, Joseph E Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang. Lightseq:: Sequence level parallelism for distributed training of long context transformers. In Workshop on Advancing Neural Network Training: Computational Efficiency, Scalability, and Resource Optimization (WANT@ NeurIPS 2023), 2023b.
  34. 34.Yuhong Li, Tianle Cai, Yi Zhang, Deming Chen, and Debadeepta Dey. What makes convolutional models great on long sequence modeling? In International Conference on Learning Representations (ICLR-2023), 2023c.
  35. 35.Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, and Thomas Serre. Learning long-range spatial dependencies with horizontal gated recurrent units. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018.
  36. 36.Hao Liu, Matei Zaharia, and Pieter Abbeel. Ring attention with blockwise transformers for near-infinite context. In International Conference on Learning Representations (ICLR-2024), 2024.
  37. 37.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12009–12019, 2022.
  38. 38.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019.
  39. 39.Chunjie Luo, Jianfeng Zhan, Xiaohe Xue, Lei Wang, Rui Ren, and Qiang Yang. Cosine normalization: Using cosine similarity instead of dot product in neural networks. In 27th International Conference on Artificial Neural Networks (ICANN-2018), pages 382–391. Springer, 2018.
  40. 40.Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer. Luna: Linear unified nested attention. Advances in Neural Information Processing Systems, 34:2441–2453, 2021.
  41. 41.Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. Mega: Moving average equipped gated attention. In The Eleventh International Conference on Learning Representations, 2023.
  42. 42.Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, pages 142–150, 2011.
  43. 43.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In International Conference on Learning Representations (ICLR-2017), 2017.
  44. 44.Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Riviere, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Leonard Hussenot, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amelie Heliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clement Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Pier Giuseppe Sessa, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clement Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy. Gemma: Open models based on gemini research and technology, 2024.
  45. 45.MosaicML. Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023.
  46. 46.Nikita Nangia and Samuel Bowman. Listops: A diagnostic dataset for latent tree learning. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 92–99, 2018.
  47. 47.Erik Nijkamp, Tian Xie, Hiroaki Hayashi, Bo Pang, Congying Xia, Chen Xing, Jesse Vig, Semih Yavuz, Philippe Laban, Ben Krause, Senthil Purushwalkam, Tong Niu, Wojciech Kryscinski, Lidiya Murakhovs’ka, Prafulla Kumar Choubey, Alex Fabbri, Ye Liu, Rui Meng, Lifu Tu, Meghana Bhat, Chien-Sheng Wu, Silvio Savarese, Yingbo Zhou, Shafiq Joty, and Caiming Xiong. Xgen-7b technical report, 2023.
  48. 48.Emilio Parisotto, Francis Song, Jack Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, et al. Stabilizing transformers for reinforcement learning. In International conference on machine learning, pages 7487–7498. PMLR, 2020.
  49. 49.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  50. 50.Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al. Rwkv: Reinventing rnns for the transformer era. arXiv preprint arXiv:2305.13048, 2023.
  51. 51.Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. YaRN: Efficient context window extension of large language models. In International Conference on Learning Representations (ICLR-2024), 2024.
  52. 52.Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Re. Hyena hierarchy: Towards larger convolutional language models. In International conference on machine learning (ICML-2023). PMLR, 2023.
  53. 53.Dragomir R Radev, Pradeep Muthukrishnan, Vahed Qazvinian, and Amjad Abu-Jbara. The acl anthology network corpus. Language Resources and Evaluation, 47(4):919–944, 2013.
  54. 54.Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. arXiv preprint, 2019. URL https://arxiv.org/abs/1911.05507.
  55. 55.Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap. Compressive transformers for long-range sequence modeling. In International Conference on Learning Representations (ICLR-2020), 2020.
  56. 56.Prajit Ramachandran, Barret Zoph, and Quoc V Le. Swish: a self-gated activation function. arXiv preprint arXiv:1710.05941, 7(1):5, 2017.
  57. 57.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  58. 58.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728, 2019.
  59. 59.Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, and Omer Levy. SCROLLS: Standardized CompaRison over long language sequences. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP-2022), pages 12007–12021, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics.
  60. 60.Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  61. 61.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019.
  62. 62.Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864, 2021.
  63. 63.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey. arXiv preprint arXiv:2009.06732, 2020.
  64. 64.Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. Long range arena : A benchmark for efficient transformers. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=qVyeW-grC2k.
  65. 65.Yi Tay, Mostafa Dehghani, Samira Abnar, Hyung Won Chung, William Fedus, Jinfeng Rao, Sharan Narang, Vinh Q. Tran, Dani Yogatama, and Donald Metzler. Scaling laws vs model architectures: How does inductive bias influence scaling?, 2022.
  66. 66.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021.
  67. 67.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  68. 68.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017.
  69. 69.Xindi Wang, Mahsa Salmani, Parsa Omidi, Xiangyu Ren, Mehdi Rezagholizadeh, and Armaghan Eshaghi. Beyond the limits: A survey of techniques to extend the context length in large language models, 2024.
  70. 70.Pete Warden. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209, 2018.
  71. 71.B. P. Welford. Note on a method for calculating corrected sums of squares and products. Technometrics, 4(3):419–420, 1962.
  72. 72.Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European conference on computer vision (ECCV-2018), pages 3–19, 2018.
  73. 73.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pages 10524–10533. PMLR, 2020.
  74. 74.Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al. Effective long-context scaling of foundation models. arXiv preprint arXiv:2309.16039, 2023.
  75. 75.Hongfei Xu, Qiuhui Liu, Deyi Xiong, and Josef van Genabith. Transformer with depth-wise lstm. arXiv preprint arXiv:2007.06257, 2020.
  76. 76.Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis. Megabyte: Predicting million-byte sequences with multiscale transformers. Advances in Neural Information Processing Systems, 36, 2024.
  77. 77.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL-2019). Association for Computational Linguistics, 2019.
  78. 78.Biao Zhang and Rico Sennrich. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  79. 79.Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev. QMSum: A new benchmark for query-based multi-domain meeting summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-2021), pages 5905–5921, Online, June 2021. Association for Computational Linguistics.
  80. 80.Yongchao Zhou, Uri Alon, Xinyun Chen, Xuezhi Wang, Rishabh Agarwal, and Denny Zhou. Transformers can achieve length generalization but not robustly, 2024.

Citation

MLA
Ma, X., et al. “Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 71831–54, https://proceedings.neurips.cc/paper_files/paper/2024/file/840abfadd04c967feaa2a49aba94a32d-Paper-Conference.pdf.
APA
Ma, X., Yang, X., Xiong, W., Chen, B., Yu, L., Zhang, H., May, J., Zettlemoyer, L., Levy, O., & Zhou, C. (2024). Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length. Advances in Neural Information Processing Systems, 37, 71831–71854. https://proceedings.neurips.cc/paper_files/paper/2024/file/840abfadd04c967feaa2a49aba94a32d-Paper-Conference.pdf
Chicago
Ma, X., X. Yang, W. Xiong, et al. 2024. “Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length”. Advances in Neural Information Processing Systems 37: 71831–54. https://proceedings.neurips.cc/paper_files/paper/2024/file/840abfadd04c967feaa2a49aba94a32d-Paper-Conference.pdf.
Harvard
Ma, X. et al. (2024) “Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 71831–71854. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/840abfadd04c967feaa2a49aba94a32d-Paper-Conference.pdf.
Vancouver
1. Ma X, Yang X, Xiong W, Chen B, Yu L, Zhang H, May J, Zettlemoyer L, Levy O, Zhou C (2024) Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 71831–71854

BibTeX

@inproceedings{ma2024megalodon,
  title = {Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length},
  author = {Ma, Xuezhe and Yang, Xiaomeng and Xiong, Wenhan and Chen, Beidi and Yu, Lili and Zhang, Hao and May, Jonathan and Zettlemoyer, Luke and Levy, Omer and Zhou, Chunting},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {71831-71854},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/840abfadd04c967feaa2a49aba94a32d-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors