FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Ying ShengLianmin ZhengBinhang YuanZhuohan LiMax RyabininBeidi ChenPercy LiangChristopher RéIon StoicaCe Zhang
Presents FlexGen, an offloading engine that combines linear-programming-based tensor scheduling across GPU, CPU, and disk with 4-bit compression to achieve up to 100-fold higher throughput for 175B-parameter model inference on a single commodity GPU.
Large language models deliver powerful capabilities across many applications, but running generative inference on models with over 100 billion parameters typically requires multiple expensive, high-end accelerators due to massive memory demands. While interactive tools like chatbots require immediate responses, many critical enterprise workloads—including benchmarking, information extraction, and data wrangling—are latency-insensitive and process large batches of text. Standard offloading systems, which move model data between graphics processing units (GPUs), central processing units (CPUs), and solid-state disks (SSDs), inherit inefficient designs from training systems that restrict batch sizes and cause severe input/output bottlenecks on budget-friendly hardware.
The article introduces and evaluates FlexGen, a generation engine engineered for high-throughput language model inference on constrained resources, such as a single commodity GPU. The objective is to demonstrate that coordinating memory hierarchies, computation scheduling, and targeted 4-bit compression can dramatically lower the hardware barrier for large model inference without compromising generation accuracy.
The authors designed a systematic offloading framework that aggregates memory across GPU, CPU, and disk storage. They implemented a zig-zag block computation schedule to reuse layer weights across batches, delegated selected attention operations to the CPU to reduce memory movement, and applied a linear programming search to optimize tensor placement across storage tiers based on hardware profiles. To further reduce input/output volume and storage footprint, the authors introduced fine-grained 4-bit group-wise quantization for both model weights and the key-value attention cache. The approach was evaluated on a single 16-gigabyte NVIDIA T4 GPU across OPT models ranging from 6.7 billion to 175 billion parameters, with additional comparisons against decentralized collaborative inference and multi-GPU pipeline parallelism.
FlexGen achieves significantly higher throughput than existing state-of-the-art offloading engines across all tested benchmarks. For a 175-billion-parameter model on a single 16-gigabyte GPU, FlexGen achieved up to 69 times higher throughput than baseline offloading systems by scaling effective batch sizes to 256, whereas baseline systems ran out of memory beyond a batch size of 2. When combined with 4-bit compression, FlexGen maintained the model weights and attention cache entirely within CPU memory—avoiding slow disk swapping entirely—to achieve a generation throughput of 1.12 tokens per second, representing a 100-fold to 112-fold speedup over baselines. In addition, 4-bit group-wise compression and 10% sparse attention showed negligible loss in accuracy on standard language benchmarks, and FlexGen delivered superior per-GPU throughput compared to decentralized collaborative network setups under varied bandwidth and latency conditions.
These findings indicate that organizations can run high-throughput, back-of-house language model workloads on accessible, low-cost commodity hardware rather than investing in expensive, multi-GPU clusters. By trading single-request latency for massive batch throughput, enterprises can dramatically cut operational infrastructure expenses for large-scale data processing and model evaluation pipelines.
Organizations should consider deploying offloading-oriented batch engines like FlexGen for non-interactive generative tasks such as batch summarization, enterprise document processing, and model auditing. When configuring hardware, teams should prioritize generous CPU memory capacities, as main system memory is critical for avoiding disk bottlenecks during offloading. Future efforts should focus on integrating dynamic batching for variable-length inputs to reduce padding overheads and implementing non-contiguous memory management to support fully optimal schedule algorithms.
Confidence in these findings is high for throughput-oriented batch workloads across standardized transformer architectures. However, decision-makers should recognize boundary conditions: offloading strategies inherently carry higher single-batch latency, making them ill-suited for real-time interactive user interfaces. Furthermore, overall throughput remains sensitive to underlying CPU memory capacity and storage read/write speeds, and variable-length inputs may experience efficiency losses when using basic uniform padding.
- Paper: ZeRO-Offload: Democratizing Billion-Scale Model Training, Jie Ren et al. (2021). ZeRO-Offload establishes the foundational CPU/GPU memory offloading framework that FlexGen analyzes, optimizes, and redesigns for high-throughput batch inference.
- Paper: DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale, Reza Yazdani Aminabadi et al. (2022). DeepSpeed Inference demonstrates multi-tier GPU-CPU-NVMe offloading systems whose throughput limitations on commodity single-GPU hardware FlexGen directly targets and overcomes.
- Paper: LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, Tim Dettmers et al. (2022). Understanding LLM.int8()'s findings on outlier activations and quantization provides the prerequisite foundation for FlexGen's 4-bit weight and KV cache compression design.
- Paper: Orca: A Distributed Serving System for Transformer-Based Generative Models, Gyeong-In Yu et al. (2022). Orca introduces iteration-level scheduling and non-uniform batching for transformer serving that motivate the scheduling and memory efficiency challenges addressed in FlexGen.
- Paper: Fast Transformer Decoding: One Write-Head is All You Need, Noam Shazeer (2019). This paper analyzes the fundamental key-value cache memory bandwidth bottlenecks during autoregressive generation that FlexGen compresses and optimizes.
- Paper: Generating Long Sequences with Sparse Transformers, Rewon Child et al. (2019). Sparse Transformers establish key techniques for attention sparsity that inform FlexGen's evaluation of sparse attention alongside low-bit quantization.
- Paper: LLM in a flash: Efficient Large Language Model Inference with Limited Memory, Keivan Alizadeh et al. (2024). LLM in a flash advances FlexGen's offloading paradigm to personal flash storage by leveraging activation sparsity and sliding-window caching for edge inference.
- Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention, Woosuk Kwon et al. (2023). PagedAttention extends the memory management problem identified in FlexGen by introducing non-contiguous paged allocation to eliminate fragmentation in the KV cache.
- Paper: SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification, Xupeng Miao et al. (2024). SpecInfer applies tree-based speculative inference specifically to offloading scenarios, speeding up single-GPU memory-offloaded LLM serving.
- Paper: Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference, Piotr Nawrot et al. (2024). Dynamic Memory Compression continues FlexGen's focus on KV cache footprint reduction by learning dynamic token state accumulation during generation.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). AWQ builds upon 4-bit compression concepts explored in FlexGen by developing an activation-aware weight quantization method tailored for hardware deployment.
- Paper: PolarQuant: Quantizing KV Caches with Polar Transformation, Insu Han et al. (2025). PolarQuant extends the low-bit KV cache compression line of research from FlexGen to long-context transformer serving using polar transformations.
- Paper: ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition, Lu Ye et al. (2024). ChunkAttention builds on KV cache reuse and partitioning schemes for shared prefixes to further enhance multi-tenant serving throughput.
- Paper: DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving, Yinmin Zhong et al. (2024). DistServe takes the scheduling and pipeline insights of generative inference further by disaggregating prefill and decoding phases onto specialized hardware.
