Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
Xuezhe MaXiaomeng YangWenhan XiongBeidi ChenLili YuHao ZhangJonathan MayLuke ZettlemoyerOmer LevyChunting Zhou
Introduces MEGALODON, a linear-complexity sequence architecture combining complex exponential moving averages with normalized gated attention that outperforms standard Transformer models like LLAMA2 at the 7-billion parameter scale across 2 trillion training tokens while scaling efficiently to virtually unlimited context lengths.
Modern artificial intelligence applications, such as analyzing extensive documents, maintaining long conversational context, and processing video, require models that can handle very long sequences of data. Standard Transformer architectures suffer from high computational costs that grow quadratically with sequence length, making them slow and expensive for long-context tasks. While alternative architectures like linear attention and state-space models offer lower computational complexity, they have historically lagged behind standard Transformers in training efficiency and overall task accuracy.
The article introduces and evaluates Megalodon, a neural network architecture designed for sequence modeling with unlimited context length. The primary objective is to demonstrate that Megalodon achieves linear computational and memory scaling while outperforming standard Transformer architectures in both pretraining efficiency and downstream accuracy.
To establish credibility under rigorous, controlled conditions, the researchers evaluated Megalodon through a head-to-head comparison against the standard Llama 2 model architecture. Both models were scaled to 7 billion parameters and trained on an identical dataset of 2 trillion tokens across 256 graphics processing units. Megalodon builds on the Moving Average Equipped Gated Attention framework by introducing complex exponential moving averages, a specialized timestep normalization technique for step-by-step sequence data, normalized attention, and an updated residual connection scheme to ensure stability during large-scale training. The authors also tested Megalodon on context lengths extending up to 2 million tokens, as well as on various standard benchmarks spanning text, speech, and image processing.
The evaluation yielded several key findings. First, Megalodon demonstrated superior training and data efficiency, reaching an overall training loss of 1.70, which significantly outperforms the 1.75 achieved by the 7-billion parameter Llama 2 and approaches the 1.67 achieved by the larger 13-billion parameter Llama 2 model. Second, Megalodon delivered major computational speedups at long context lengths: while roughly 6% slower than Llama 2 on short sequences of 4,000 tokens, it was approximately 32% faster when trained on 32,000-token sequences. Third, Megalodon consistently outperformed the 7-billion parameter Llama 2 on standard academic benchmarks and competitive long-context question-answering datasets. Fourth, tests extending sequence lengths up to 2 million tokens showed monotonic improvements in prediction performance, confirming effective long-context modeling. Finally, smaller-scale tests in raw audio and image classification confirmed that the architecture generalizes effectively across different data types.
These findings suggest that organizations can achieve higher model quality at reduced computational expense for long-context workloads. By lowering the processing overhead from quadratic to linear complexity, Megalodon reduces hardware costs and training timelines for processing extensive data sequences. Moreover, its ability to match or exceed the performance of a standard 13-billion parameter model while using only 7 billion parameters enables substantial operational efficiencies in deployment.
Based on these results, decision-makers considering long-context language systems should consider evaluating chunked linear attention architectures as an alternative to standard Transformer configurations. For teams managing production deployments, the main trade-off lies in sequence length: standard architectures remain slightly faster for short sequences under 4,000 tokens, but Megalodon provides clear computational and quality advantages for workloads that require long context windows. Future work should focus on validating the architecture at larger parameter scales and applying it to comprehensive multi-modal pretraining.
While confidence in these findings is high due to the strictly controlled 7-billion parameter pretraining setup, some limitations remain. The large-scale pretraining evaluation was limited to a single 7-billion parameter model size, and comparisons against certain external models were constrained by differing pretraining data volumes. Stakeholders should conduct pilot evaluations within their specific task pipelines before making full-scale infrastructure transitions.
- Paper: Mamba: Linear-Time Sequence Modeling with Selective State Spaces, Albert Gu et al. (2023). Introduces selective state space modeling with linear complexity, serving as the foundational sub-quadratic baseline that Megalodon seeks to outperform in pretraining stability and task capability.
- Paper: LLaMA: Open and Efficient Foundation Language Models, Hugo Touvron et al. (2023). Establishes the standard open foundation model architecture and pretraining methodology that Megalodon matches and benchmarks against at scale.
- Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). Formulates linear attention by recasting self-attention into linear recurrences, establishing the theoretical mechanism underlying sub-quadratic sequence models.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). Pioneers positional bias mechanisms that enable sequence length extrapolation, directly motivating Megalodon's focus on unbounded context modeling.
- Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). Introduces segment-level recurrence and relative position encodings to overcome fixed context boundaries in language modeling.
- Paper: Hungry Hungry Hippos: Towards Language Modeling with State Space Models, Daniel Y. Fu et al. (2023). Analyzes the expressive gap between attention and state-space models in language modeling, informing the design tradeoffs that Megalodon addresses.
- Paper: Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling, Liliang Ren et al. (2025). Combines selective state space layers with sliding-window attention to achieve unlimited context length extrapolation with linear time complexity.
- Paper: Titans: Learning to Memorize at Test Time, Ali Behrouz et al. (2024). Extends long-context sub-quadratic sequence modeling by incorporating test-time adaptive memory mechanisms alongside localized attention.
- Paper: Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Tri Dao et al. (2024). Connects structured state-space models and attention mechanisms through structured matrix duality, offering a complementary path to scaling sub-quadratic architectures.
- Paper: Learning to (Learn at Test Time): RNNs with Expressive Hidden States, Yu Sun 0020 et al. (2025). Further investigates overcoming linear-complexity expressive limits by turning recurrent hidden states into expressive models trained at test time.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). Provides a comprehensive taxonomy and evaluation of emerging efficient LLM architectures, situating linear sequence models and hybrid attention designs within the broader field.
