FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Ying ShengLianmin ZhengBinhang YuanZhuohan LiMax RyabininBeidi ChenPercy LiangChristopher RéIon StoicaCe Zhang

article2023ICML943 citations

Presents FlexGen, an offloading engine that combines linear-programming-based tensor scheduling across GPU, CPU, and disk with 4-bit compression to achieve up to 100-fold higher throughput for 175B-parameter model inference on a single commodity GPU.

Listen

Large language models deliver powerful capabilities across many applications, but running generative inference on models with over 100 billion parameters typically requires multiple expensive, high-end accelerators due to massive memory demands. While interactive tools like chatbots require immediate responses, many critical enterprise workloads—including benchmarking, information extraction, and data wrangling—are latency-insensitive and process large batches of text. Standard offloading systems, which move model data between graphics processing units (GPUs), central processing units (CPUs), and solid-state disks (SSDs), inherit inefficient designs from training systems that restrict batch sizes and cause severe input/output bottlenecks on budget-friendly hardware.

The article introduces and evaluates FlexGen, a generation engine engineered for high-throughput language model inference on constrained resources, such as a single commodity GPU. The objective is to demonstrate that coordinating memory hierarchies, computation scheduling, and targeted 4-bit compression can dramatically lower the hardware barrier for large model inference without compromising generation accuracy.

The authors designed a systematic offloading framework that aggregates memory across GPU, CPU, and disk storage. They implemented a zig-zag block computation schedule to reuse layer weights across batches, delegated selected attention operations to the CPU to reduce memory movement, and applied a linear programming search to optimize tensor placement across storage tiers based on hardware profiles. To further reduce input/output volume and storage footprint, the authors introduced fine-grained 4-bit group-wise quantization for both model weights and the key-value attention cache. The approach was evaluated on a single 16-gigabyte NVIDIA T4 GPU across OPT models ranging from 6.7 billion to 175 billion parameters, with additional comparisons against decentralized collaborative inference and multi-GPU pipeline parallelism.

FlexGen achieves significantly higher throughput than existing state-of-the-art offloading engines across all tested benchmarks. For a 175-billion-parameter model on a single 16-gigabyte GPU, FlexGen achieved up to 69 times higher throughput than baseline offloading systems by scaling effective batch sizes to 256, whereas baseline systems ran out of memory beyond a batch size of 2. When combined with 4-bit compression, FlexGen maintained the model weights and attention cache entirely within CPU memory—avoiding slow disk swapping entirely—to achieve a generation throughput of 1.12 tokens per second, representing a 100-fold to 112-fold speedup over baselines. In addition, 4-bit group-wise compression and 10% sparse attention showed negligible loss in accuracy on standard language benchmarks, and FlexGen delivered superior per-GPU throughput compared to decentralized collaborative network setups under varied bandwidth and latency conditions.

These findings indicate that organizations can run high-throughput, back-of-house language model workloads on accessible, low-cost commodity hardware rather than investing in expensive, multi-GPU clusters. By trading single-request latency for massive batch throughput, enterprises can dramatically cut operational infrastructure expenses for large-scale data processing and model evaluation pipelines.

Organizations should consider deploying offloading-oriented batch engines like FlexGen for non-interactive generative tasks such as batch summarization, enterprise document processing, and model auditing. When configuring hardware, teams should prioritize generous CPU memory capacities, as main system memory is critical for avoiding disk bottlenecks during offloading. Future efforts should focus on integrating dynamic batching for variable-length inputs to reduce padding overheads and implementing non-contiguous memory management to support fully optimal schedule algorithms.

Confidence in these findings is high for throughput-oriented batch workloads across standardized transformer architectures. However, decision-makers should recognize boundary conditions: offloading strategies inherently carry higher single-batch latency, making them ill-suited for real-time interactive user interfaces. Furthermore, overall throughput remains sensitive to underlying CPU memory capacity and storage read/write speeds, and variable-length inputs may experience efficiency losses when using basic uniform padding.

arXiv: 2303.06865
Cover for FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Abstract

The high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators. Motivated by the emerging demand for latency-insensitive tasks with batched processing, this paper initiates the study of high-throughput LLM inference using limited resources, such as a single commodity GPU. We present FlexGen, a high-throughput generation engine for running LLMs with limited GPU memory. FlexGen can be flexibly configured under various hardware resource constraints by aggregating memory and computation from the GPU, CPU, and disk. By solving a linear programming problem, it searches for efficient patterns to store and access tensors. FlexGen further compresses the weights and the attention cache to 4 bits with negligible accuracy loss. These techniques enable FlexGen to have a larger space of batch size choices and thus significantly increase maximum throughput. As a result, when running OPT-175B on a single 16GB GPU, FlexGen achieves significantly higher throughput compared to state-of-the-art offloading systems, reaching a generation throughput of 1 token/s for the first time with an effective batch size of 144. On the HELM benchmark, FlexGen can benchmark a 30B model with a 16GB GPU on 7 representative sub-scenarios in 21 hours. The code is available at https://github.com/FMInference/FlexGen.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background: LLM Inference
  • 4. Offloading Strategy
  • 4.1. Problem Formulation
  • 4.2. Search Space
  • 4.3. Cost Model and Policy Search
  • 4.4. Extension to Multiple GPUs
  • 5. Approximate Methods
  • 6. Evaluation
  • 6.1. Offloading
  • 6.2. Approximations
  • 6.3. Offloading vs. Collaborative Inference
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. Appendix
  • A.1. Notations
  • A.2. Compute Schedule Optimality
  • A.2.1. ZIG-ZAG BLOCK SCHEDULE AND DIAGONAL BLOCK SCHEDULE
  • A.2.2. PROOF OF THEOREM 4.1
  • A.3. Cost Model
  • A.4. Tables and Additional Experimental Results

Knowls

  1. Knowl 1 — FlexGen’s unified offloading formulation

    model/method

    FlexGen treats high-throughput generative inference as a graph-traversal problem over a three-level memory hierarchy consisting of GPU memory, CPU DRAM, and disk. The GPU and CPU can compute, whereas the disk is storage-only. The system assumes an effectively infinite stream of prompts so that latency can be traded for a large effective batch size.

    Each graph node represents computing one GPU batch for one transformer layer and one generation step. A valid execution path must satisfy four conditions: (1) a node can be computed only after all nodes to its left in the same prompt row have been computed; (2) all of a node’s inputs—weights, activations, and KV cache—must be colocated on the device performing the computation; (3) activations remain available until the next layer consumes them, while KV-cache entries remain available until the final generated token for that prompt; and (4) the tensors resident on every device must fit within that device’s capacity.

    The optimization objective is to find a valid path minimizing total execution time, including both computation and tensor transfers. Unlike row-by-row offloading, FlexGen jointly chooses the schedule and the locations of weights, activations, and KV cache.

  2. Knowl 2 — Overlapped zig-zag block scheduling

    algorithm

    FlexGen’s implemented schedule groups several GPU batches into a block and traverses the computation graph in a zig-zag order. Let ll be the number of transformer layers, nn the number of generated tokens, gbsgbs the GPU batch size, and KK the number of GPU batches in one block. The effective batch size is bls=gbs×Kbls=gbs\times K.

    For each generation step, FlexGen processes layers sequentially. Within a layer, it processes the KK GPU batches while overlapping six independent operations: loading the next layer’s weights, storing the previous batch’s activations, storing the previous batch’s KV cache, loading the next batch’s activations, loading the next batch’s KV cache, and computing the current batch. A synchronization follows these operations before the next batch proceeds. Boundary operations that have no predecessor or successor are skipped.

    Input: prompt stream, layer count ll, generation length nn, GPU batch size gbsgbs, GPU-batch count KK
    Output: generated tokens for blocks of effective batch size gbs×Kgbs × K
    for generation step from 1 to $n do
        for layer from 1 to $l do
            for GPU batch from 1 to $K do
                launch in parallel loading the next layer's weights
                launch in parallel storing the previous batch's activations
                launch in parallel storing the previous batch's KV cache
                launch in parallel loading the next batch's activations
                launch in parallel loading the next batch's KV cache
                launch in parallel computing the current GPU batch
                synchronize all devices and streams
            end for
        end for
    end for

    FlexGen uses multiple CUDA streams and CPU threads to realize the overlap. The zig-zag schedule reuses weights across the GPU batches in a block, while activations and KV cache are transferred as needed.

  3. Knowl 3 — Cost-model-driven policy search

    model/method

    A FlexGen policy has eleven variables: block size blsbls, GPU batch size gbsgbs, and the placement fractions of weights, activations, and KV cache on GPU, CPU, and disk. For weights these fractions are (wg,wc,wd)(w_g,w_c,w_d); for activations they are (hg,hc,hd)(h_g,h_c,h_d); and for KV cache they are (cg,cc,cd)(c_g,c_c,c_d). Each triple sums to one.

    The analytical cost model estimates prefill latency TpreT_{pre} and average per-layer decoding latency TgenT_{gen}. Assuming transfers and computation can overlap, each is the maximum of the relevant GPU-to-CPU, CPU-to-GPU, disk-to-CPU, CPU-to-disk, and computation times. For a model with ll layers and output length nn, the latency of one block is modeled as

    T=lTpre+l(n−1)Tgen.T=lT_{pre}+l(n-1)T_{gen}.

    The objective is to minimize latency per effective batch element, T/blsT/bls, subject to GPU, CPU, and disk peak-memory constraints and the three placement-sum constraints:

    \min_{p}\quad & T/bls \\ \text{subject to}\quad & M_{GPU}(p)<C_{GPU},\quad M_{CPU}(p)<C_{CPU},\quad M_{disk}(p)<C_{disk},\\ & w_g+w_c+w_d=1,\\ & h_g+h_c+h_d=1,\\ & c_g+c_c+c_d=1, \end{aligned}$$ where $p$ is the nine-dimensional placement vector, $M_{GPU}$, $M_{CPU}$, and $M_{disk}$ are modeled peak usages, and $C_{GPU}$, $C_{CPU}$, and $C_{disk}$ are device capacities. FlexGen enumerates a small set of $(bls,gbs)$ choices—typically with $gbs$ a multiple of four and $bls<20$—and solves the remaining placement problem as a nine-variable linear program. Hardware bandwidth and compute parameters are fitted from profiling measurements. The relaxed policy may require small manual adjustments because real memory fragmentation and discrete tensor partitioning are not modeled exactly.
  4. Knowl 4 — I/O optimality guarantees for the schedules

    theoretical result

    For the offloading graph in which the model does not fit on one GPU, and under the theoretical analysis that excludes CPU computation, the zig-zag block schedule has I/O complexity at most twice that of an optimal valid schedule. The comparison concerns the tensor-movement cost, with the same generation workload and finite secondary-memory capacity.

    The paper also defines a diagonal block schedule for analysis. After a warm-up, it traverses diagonals of the computation graph so that each weight load uses memory more evenly. This diagonal schedule is asymptotically I/O-optimal and can support a larger block than the zig-zag schedule. If ss is prompt length and nn is output length, the ratio of the diagonal schedule’s admissible block size to the zig-zag schedule’s block size is

    blsdiagonalblszigzag=2s+2n2s+n+O ⁣(1n).\frac{bls_{diagonal}}{bls_{zigzag}}=\frac{2s+2n}{2s+n}+O\!\left(\frac{1}{n}\right).

    The ratio approaches 22 when nn is much larger than ss and approaches 11 when ss is much larger than nn. FlexGen implements the simpler zig-zag schedule because the diagonal schedule requires dynamically updating non-contiguous KV-cache buffers; the implementation preallocates contiguous cache buffers instead.

  5. Knowl 5 — Delegating attention computation to the CPU

    model/method

    FlexGen can delegate decoding attention-score computation to the CPU when the associated KV cache is stored outside GPU memory. Let bb be the GPU batch size, ss the current sequence length, and h1h_1 the transformer hidden size. Moving the KV cache to the GPU for attention requires transferring approximately bsh1×4bsh_1\times4 bytes, whereas sending only the current activation to the CPU requires approximately bh1×4bh_1\times4 bytes. Thus CPU-side attention reduces the relevant transfer volume by a factor of approximately ss.

    This delegation is beneficial when attention is I/O-bound and the KV cache is not already on the GPU; the paper identifies long sequences such as s≥512s\geq512 as a representative regime. The CPU still performs the slower arithmetic, but avoiding movement of the much larger cache can reduce total latency.

  6. Knowl 6 — Four-bit group-wise compression of weights and KV cache

    model/method

    FlexGen compresses both transformer weights and KV-cache tensors to 4-bit integers without retraining or calibration. For each group of gg contiguous real-valued tensor elements, it computes the group minimum xminx_{min} and maximum xmaxx_{max}, then maps an element xx to a qq-bit integer using asymmetric min–max quantization:

    xquant=round⁡ ⁣(x−xminxmax−xmin(2q−1)).x_{quant}=\operatorname{round}\!\left(\frac{x-x_{min}}{x_{max}-x_{min}}(2^q-1)\right).

    Here q=4q=4 in the main configuration, xx is an original real-valued element, and xminx_{min} and xmaxx_{max} are computed independently for each group. The quantized tensors are stored in compressed form and dequantized to FP16 before computation. FlexGen uses group size g=64g=64, groups weights along the output-channel dimension, and groups KV cache along the hidden dimension.

    The purpose is primarily to reduce memory use and I/O volume rather than to perform integer matrix multiplication. Because fine-grained compression and decompression add CPU overhead, FlexGen disables CPU computation delegation when 4-bit compression is enabled.

  7. Knowl 7 — Top-10-percent sparse attention approximation

    model/method

    FlexGen implements a Top-K sparse-attention approximation that retains only the top 10% of value-cache tokens for each query. After computing attention scores against the full key cache, it identifies the indices of the highest-scoring tokens, discards the remaining tokens, and loads only the corresponding subset of the value cache.

    This approximation is intended to reduce KV-cache I/O while preserving generation quality. The paper reports that applying the 10% value-cache selection to OPT-175B maintains model quality in its evaluated tasks, although the result is presented as a preliminary demonstration that additional approximation methods can be plugged into the FlexGen framework.

  8. Knowl 8 — Pipeline parallelism across multiple GPUs

    model/method

    For multiple GPUs, FlexGen uses pipeline rather than tensor parallelism because the target workload is throughput-oriented and pipeline parallelism has lower communication cost. An ll-layer transformer is partitioned approximately evenly across mm GPUs, each GPU runs the same offloading policy for its assigned layer range, and micro-batch pipelining combines the inter-GPU pipeline schedule with the single-device overlapped block runtime.

    The policy search can be reused independently for the reduced layer count on each pipeline stage. Partitioning reduces the memory pressure on each GPU, which can move the system from disk offloading to CPU-only offloading or permit a much larger effective batch size. Consequently, decoding throughput can scale super-linearly with the number of GPUs, even though end-to-end generation throughput may scale less than linearly because prefill contains pipeline bubbles.

  9. Knowl 9 — Single-GPU and multi-GPU throughput results

    data/table

    FlexGen was evaluated on NVIDIA T4 GPUs with 16 GB GPU memory, 208 GB CPU DRAM, and a 1.5 TB SSD. The workloads used OPT models, synthetic padded prompts, 32 generated tokens per prompt, and prompt lengths of 512 or 1024 tokens. Throughput is generated tokens divided by prefill plus decoding time. Petals results are normalized per GPU; its model-specific GPU counts were 1 for OPT-6.7B, 4 for OPT-30B, and 24 for OPT-175B.

    The following results compare maximum generation throughput in tokens per second:

    System Prompt length 512 Prompt length 1024
    6.7B 30B 175B 6.7B 30B 175B
    Accelerate 25.12 0.62 0.01 13.01 0.31 0.01
    DeepSpeed 9.28 0.60 0.01 4.59 0.29 OOM
    Petals 8.25 2.84 0.08 6.56 1.51 0.06
    FlexGen 25.26 7.32 0.69 13.72 3.50 0.35
    FlexGen (4-bit) 29.12 8.70 1.12 13.18 3.98 0.42

    For OPT-175B with prompt length 512, uncompressed FlexGen reaches 0.690.69 token/s versus 0.010.01 token/s for the offloading baselines; 4-bit compression raises this to 1.121.12 token/s by fitting the weights and KV cache in CPU memory and avoiding disk swapping. The largest compressed effective batch size in the main trade-off experiments is 144, yielding roughly 100 times the baseline maximum throughput under the reported latency configuration.

    With four T4 GPUs, pipeline-parallel FlexGen achieves the following generation and decoding throughputs for prompt length 512:

    System Generation throughput Decoding throughput
    6.7B 30B 175B 6.7B 30B 175B
    FlexGen (1 GPU) 25.26 7.32 0.69 38.28 11.52 0.83
    FlexGen (4 GPUs) 201.12 23.61 2.33 764.65 48.94 3.86
    DeepSpeed (4 GPUs) 50.00 6.40 0.05 50.20 6.40 0.05

    The stronger scaling of decoding than of full generation is attributed to prefill pipeline bubbles and the relatively short output length used in the benchmark.

  10. Knowl 10 — Ablation and accuracy validation of FlexGen techniques

    empirical result

    On one T4 GPU with prompt length 512 and output length 32, the individual FlexGen components materially affect throughput, especially for OPT-175B. The following ablation results are generation throughput in tokens per second:

    Configuration OPT-30B OPT-175B
    All optimizations 7.32 0.69
    No policy search 7.26 0.27
    No overlapping 5.86 0.59
    No CPU compute 4.03 0.62
    No disk 7.32 OOM
    DeepSpeed policy 1.57 0.01

    The results show that policy search is crucial for the disk-bound 175B model, overlapping improves both models, CPU computation helps more for the CPU-bound 30B model, and the 175B model cannot run under the no-disk configuration. Replacing FlexGen’s policy with the row-by-row DeepSpeed policy sharply reduces throughput.

    Accuracy was measured using Lambada next-word prediction accuracy and WikiText perplexity, where higher accuracy and lower perplexity are better:

    Dataset Metric OPT-30B OPT-175B
    FP16 4-bit FP16 4-bit
    Lambada accuracy 0.725 0.724 0.758 0.756
    WikiText perplexity 12.72 12.90 10.82 10.94

    Combining 4-bit quantization with 10% value-cache sparse attention produced Lambada accuracies of 0.718 for OPT-30B and 0.756 for OPT-175B, and WikiText perplexities of 12.90 and 10.94, respectively. The paper therefore reports negligible quality loss for these approximations, while 3-bit compression did not preserve accuracy. FlexGen also completed seven representative HELM sub-scenarios for OPT-IML-30B in 21 hours, including dataset download, initialization, generation, and metric computation.

Coverage note — The detailed full cost-model memory formulas, extended data-wrangling tables, additional hardware experiments, and complete latency-throughput Pareto tables were omitted because the ten knowls already capture the paper’s load-bearing method, theory, approximation mechanisms, and primary empirical evidence.

References

  1. 1.Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., et al. Deepspeed-inference: Enabling efficient inference of transformer models at unprecedented scale. In 2022 SC22: International Conference for High Performance Computing, Networking, Storage and Analysis (SC), pp. 646–660. IEEE Computer Society, 2022.
  2. 2.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  3. 3.Borzunov, A., Baranchuk, D., Dettmers, T., Ryabinin, M., Belkada, Y., Chumachenko, A., Samygin, P., and Raffel, C. Petals: Collaborative inference and fine-tuning of large models. arXiv preprint arXiv:2209.01188, 2022.
  4. 4.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  5. 5.Chen, X., Maniatis, P., Singh, R., Sutton, C., Dai, H., Lin, M., and Zhou, D. Spreadsheetcoder: Formula prediction from semi-structured context. In International Conference on Machine Learning, pp. 1661–1672. PMLR, 2021.
  6. 6.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  7. 7.Demmel, J. Communication-avoiding algorithms for linear algebra and beyond. In 2013 IEEE 27th International Symposium on Parallel and Distributed Processing, pp. 585–585. IEEE, 2013.
  8. 8.Dettmers, T. and Zettlemoyer, L. The case for 4-bit precision: k-bit inference scaling laws. arXiv preprint arXiv:2212.09720, 2022.
  9. 9.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. Llm.int8(): 8-bit matrix multiplication for transformers at scale. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022.
  10. 10.Fang, J., Yu, Y., Zhao, C., and Zhou, J. Turbotransformers: an efficient gpu serving system for transformer models. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, pp. 389–402, 2021.
  11. 11.Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022.
  12. 12.Frantar, E. and Alistarh, D. Massive language models can be accurately pruned in one-shot. arXiv preprint arXiv:2301.00774, 2023.
  13. 13.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. Gptq: Accurate post-training quantization for generative pretrained transformers. arXiv preprint arXiv:2210.17323, 2022.
  14. 14.Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. J. Mach. Learn. Res., 22(241):1–124, 2021.
  15. 15.Huang, C.-C., Jin, G., and Li, J. Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems, pp. 1341–1355, 2020.
  16. 16.Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al. Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in neural information processing systems, 32, 2019.
  17. 17.HuggingFace. Hugging face accelerate. https://huggingface.co/docs/accelerate/index, 2022.
  18. 18.Iyer, S., Lin, X. V., Pasunuru, R., Mihaylov, T., Simig, D., Yu, P., Shuster, K., Wang, T., Liu, Q., Koura, P. S., et al. Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017, 2022.
  19. 19.Jia-Wei, H. and Kung, H.-T. I/o complexity: The red-blue pebble game. In Proceedings of the thirteenth annual ACM symposium on Theory of computing, pp. 326–333, 1981.
  20. 20.Kwon, S. J., Kim, J., Bae, J., Yoo, K. M., Kim, J.-H., Park, B., Kim, B., Ha, J.-W., Sung, N., and Lee, D. Alphatuning: Quantization-aware parameter-efficient adaptation of large-scale pre-trained language models. arXiv preprint arXiv:2210.03858, 2022.
  21. 21.Li, Y., Phanishayee, A., Murray, D., Tarnawski, J., and Kim, N. S. Harmony: Overcoming the hurdles of gpu memory capacity to train massive dnn models on commodity servers. arXiv preprint arXiv:2202.01306, 2022.
  22. 22.Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
  23. 23.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016.
  24. 24.Morton, A. Pagecachemangagement. https://code.google.com/archive/p/pagecache-mangagement/source/default/source, 2008.
  25. 25.Narayan, A., Chami, I., Orr, L., and Re, C. Can foundation models wrangle your data? arXiv preprint arXiv:2205.09911, 2022.
  26. 26.Narayan, S., Cohen, S. B., and Lapata, M. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. arXiv preprint arXiv:1808.08745, 2018.
  27. 27.Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al. Efficient large-scale language model training on gpu clusters using megatron-lm. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–15, 2021.
  28. 28.NVIDIA. Fastertransformer. https://github.com/NVIDIA/FasterTransformer, 2022.
  29. 29.Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N.-Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernandez, R. The lambada dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1525–1534, 2016.
  30. 30.Park, G., Park, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D. nuqmm: Quantized matmul for efficient inference of large-scale generative language models. arXiv preprint arXiv:2206.09557, 2022.
  31. 31.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  32. 32.Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., and Dean, J. Efficiently scaling transformer inference. arXiv preprint arXiv:2211.05102, 2022.
  33. 33.Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y. Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–14, 2021.
  34. 34.Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y. Zero-offload: Democratizing billion-scale model training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21), pp. 551–564, 2021.
  35. 35.Ryabinin, M., Dettmers, T., Diskin, M., and Borzunov, A. Swarm parallelism: Training large models can be surprisingly communication-efficient. arXiv preprint arXiv:2301.11913, 2023.
  36. 36.Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagne, R., Luccioni, A. S., Yvon, F., Gallé, M., et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  37. 37.Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K. Q-bert: Hessian based ultra low precision quantization of bert. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp. 8815–8821, 2020.
  38. 38.Steiner, B., Elhoushi, M., Kahn, J., and Hegarty, J. Olla: Optimizing the lifetime and location of arrays to reduce the memory usage of neural networks. 2022. doi: 10.48550/arXiv.2210.12924.
  39. 39.Wang, L., Ye, J., Zhao, Y., Wu, W., Li, A., Song, S. L., Xu, Z., and Kraska, T. Superneurons: Dynamic gpu memory management for training deep neural networks. In Proceedings of the 23rd ACM SIGPLAN symposium on principles and practice of parallel programming, pp. 41–53, 2018.
  40. 40.Wang, X., Xiong, Y., Wei, Y., Wang, M., and Li, L. Lightseq: A high performance inference library for transformers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Papers, pp. 113–120, 2021.
  41. 41.Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S. Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv preprint arXiv:2211.10438, 2022.
  42. 42.Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022.
  43. 43.Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G. Orca: A distributed serving system for {Transformer-Based} generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), pp. 521–538, 2022.
  44. 44.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  45. 45.Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Gonzalez, J. E., et al. Alpa: Automating inter-and intra-operator parallelism for distributed deep learning. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), 2022.

Citation

MLA
Sheng, Y., et al. “FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”. International Conference on Machine Learning, vol. 202, 2023, pp. 31094–116, https://proceedings.mlr.press/v202/sheng23a.html.
APA
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Re, C., Stoica, I., & Zhang, C. (2023). FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU. International Conference on Machine Learning, 202, 31094–31116. https://proceedings.mlr.press/v202/sheng23a.html
Chicago
Sheng, Y., L. Zheng, B. Yuan, et al. 2023. “FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”. International Conference on Machine Learning 202: 31094–116. https://proceedings.mlr.press/v202/sheng23a.html.
Harvard
Sheng, Y. et al. (2023) “FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”, International Conference on Machine Learning. PMLR, pp. 31094–31116. Available at: https://proceedings.mlr.press/v202/sheng23a.html.
Vancouver
1. Sheng Y, Zheng L, Yuan B, Li Z, Ryabinin M, Chen B, Liang P, Re C, Stoica I, Zhang C (2023) FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU. In: International Conference on Machine Learning. PMLR, pp 31094–31116

BibTeX

@InProceedings{pmlr-v202-sheng23a,
  title = 	 {{F}lex{G}en: High-Throughput Generative Inference of Large Language Models with a Single {GPU}},
  author =       {Sheng, Ying and Zheng, Lianmin and Yuan, Binhang and Li, Zhuohan and Ryabinin, Max and Chen, Beidi and Liang, Percy and Re, Christopher and Stoica, Ion and Zhang, Ce},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {31094--31116},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/sheng23a/sheng23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/sheng23a.html},
  abstract = 	 {The high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators. Motivated by the emerging demand for latency-insensitive tasks with batched processing, this paper initiates the study of high-throughput LLM inference using limited resources, such as a single commodity GPU. We present FlexGen, a high-throughput generation engine for running LLMs with limited GPU memory. FlexGen can be flexibly configured under various hardware resource constraints by aggregating memory and computation from the GPU, CPU, and disk. By solving a linear programming problem, it searches for efficient patterns to store and access tensors. FlexGen further compresses the weights and the attention cache to 4 bits with negligible accuracy loss. These techniques enable FlexGen to have a larger space of batch size choices and thus significantly increase maximum throughput. As a result, when running OPT-175B on a single 16GB GPU, FlexGen achieves significantly higher throughput compared to state-of-the-art offloading systems, reaching a generation throughput of 1 token/s for the first time with an effective batch size of 144. On the HELM benchmark, FlexGen can benchmark a 30B model with a 16GB GPU on 7 representative sub-scenarios in 21 hours. The code is available at https://github.com/FMInference/FlexGen.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/