FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Jay ShahGanesh BikshandiYing ZhangVijay ThakkarPradeep RamaniTri Dao
Introduces FlashAttention-3, a method that speeds up Transformer attention by up to 2x on Hopper GPUs by overlapping computation with data movement via hardware asynchrony and executing FP8 low-precision operations with 2.6x lower numerical error.
The attention mechanism is the core computational foundation of modern artificial intelligence models, including large language models. However, it represents a primary processing bottleneck because its computational demands scale quadratically with context length. While earlier solutions like FlashAttention-2 reduced data transfer between memory layers, they achieved only 35% utilization on modern GPU architectures such as NVIDIA Hopper H100 because they followed a synchronous processing model and did not fully exploit specialized asynchronous hardware features or low-precision numerical formats.
The article demonstrates an optimized attention algorithm, FlashAttention-3, designed to maximize throughput and numerical accuracy on modern GPUs. It evaluates how asynchronous hardware execution, software pipelining, and 8-bit floating-point (FP8) computations can accelerate both the training and execution of transformer models.
To achieve these gains, the researchers developed three main techniques: separating data movement and computation into specialized worker groups (warp-specialization), interleaving lower-throughput softmax calculations with asynchronous matrix multiplications across a two-stage software pipeline, and incorporating hardware-accelerated FP8 operations with block quantization and incoherent processing to prevent quantization errors caused by outlier values. The approach was empirically validated on NVIDIA H100 SXM5 GPUs across various sequence lengths (512 to 16,384 tokens), head dimensions, and masking configurations, averaging execution metrics across 100 benchmark iterations against standard PyTorch, FlashAttention-2, Triton, and proprietary cuDNN implementations.
The findings show that for 16-bit floating-point (FP16) inputs, FlashAttention-3 achieves a 1.5× to 2.0× speedup in the forward pass over FlashAttention-2, reaching up to 740 teraflops per second (75% GPU utilization) and a 1.5× to 1.75× speedup in the backward pass. For low-precision FP8 operations, performance approaches 1.2 petaflops per second, roughly doubling FP16 throughput. Furthermore, for long sequence lengths, FlashAttention-3 matches or outperforms closed-source, vendor-optimized libraries. In terms of numerical accuracy, FP8 FlashAttention-3 reduces error by 2.6× compared to standard baseline FP8 attention when processing outlier features.
These performance improvements have immediate practical implications: they substantially reduce computational training costs, increase inference speeds, and make expanding AI context windows to handle multi-document analysis, long video streams, and extensive codebases commercially feasible. The algorithm provides a drop-in replacement that delivers higher hardware utilization without requiring model architecture compromises.
Organizations developing or deploying large-scale language models should adopt FlashAttention-3 via open-source integrations (such as PyTorch and Hugging Face) to reduce infrastructure overhead. Teams should leverage FP8 low-precision pipelines where supported by hardware, utilizing the recommended block quantization to protect output quality against numerical degradation.
While confidence in the benchmarked hardware performance is high, several limitations remain. FlashAttention-3 was benchmarked specifically on Hopper architectures; deeper 3-stage pipelining showed diminished returns due to register pressure and compiler reordering; and the FP8 implementation currently lacks persistent kernel load-balancing optimizations, which slightly reduces efficiency on small sequence lengths with causal masking. Additional empirical validation is recommended when extending FP8 attention to full-scale foundation model pretraining.
- Paper: FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning, Tri Dao (2023). FlashAttention-3 directly builds on FlashAttention-2's work partitioning and tiling techniques to push GPU throughput further using Hopper-specific asynchronous features.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). This foundational paper introduces the exact, IO-aware tiling and online softmax recomputation algorithm that forms the core baseline for all subsequent FlashAttention architectures.
No sufficiently relevant recommendations were found.
