LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Keivan AlizadehIman MirzadehDmitry BelenkoS. KhatamifardMinsik ChoCarlo C. del MundoMohammad RastegariMehrdad Farajtabar

article2024ACL308 citations

Presents hardware-informed techniques that exploit activation sparsity and chunked data access to run large language models exceeding device DRAM directly from flash memory with up to a twentyfold speedup over standard offloading methods.

Listen

Deploying modern large language models on personal and edge devices is currently constrained by limited high-speed system memory (DRAM). Standard inference requires loading the full model into DRAM, which prevents personal devices such as smartphones and laptops from running models that exceed their memory capacity. The article evaluates a framework designed to run large language models that are up to twice the size of available DRAM by storing parameters in larger flash storage and loading them on demand during inference.

To overcome the significant latency and throughput penalties of flash storage, the article demonstrates a hardware-informed approach evaluated across diverse hardware backends, including Apple Silicon central processing units (CPUs) and graphics processing units (GPUs) as well as NVIDIA GPUs. The evaluation focuses on popular model families, such as OPT, Falcon, Persimmon, Phi-2, and Llama 2. Rather than loading full layers, the framework exploits natural neuron activation sparsity within feed-forward networks by using a lightweight low-rank predictor to forecast which weights are required for upcoming tokens. It also applies a sliding-window caching technique to retain recently activated neurons in memory and bundles corresponding weight rows and columns to double flash read chunk sizes for increased input/output throughput.

The findings show substantial performance gains over conventional on-demand loading approaches. Combining sparsity prediction, windowing, and bundling reduces data transfer by over 90% in tested configurations, enabling models twice the size of available DRAM to execute efficiently. The proposed method accelerates single-token inference speed by 4 to 5 times on CPUs, about 7 times on Apple Metal GPUs, and 20 to 25 times on NVIDIA GPUs compared to naive baseline loading. When speculative decoding is integrated on GPUs, the inference pipeline achieves an additional 1.4 times speedup without degrading baseline task accuracy or output perplexity.

These results demonstrate that high-capability language models can operate directly on consumer hardware without incurring the high manufacturing costs of expanding device DRAM. This capability enhances user privacy, reduces dependence on centralized cloud infrastructure, and lowers response latency. However, while instantaneous operational power is lower due to sparse execution, the total energy consumed across token generation is slightly higher due to the extended duration of flash memory transfers compared to fully DRAM-resident models.

Organizations developing on-device artificial intelligence should implement sparsity-aware parameter streaming and hardware-aligned memory caching to deploy larger models on constrained devices. Product teams should further explore combining this approach with 4-bit model quantization to reduce the memory footprint beneath 2 gigabytes on mobile platforms. Prior to commercial deployment, technical teams should conduct broader evaluations covering multi-batch inference, prompt processing stages, and device-level thermal and battery dissipation over sustained workloads. The reported findings provide high confidence for single-sequence, on-device text generation under tightly constrained memory conditions.

arXiv: 2312.11514
Cover for LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Abstract

Large language models (LLMs) are central to modern natural language processing, delivering exceptional performance in various tasks. However, their substantial computational and memory requirements present challenges, especially for devices with limited DRAM capacity. This paper tackles the challenge of efficiently running LLMs that exceed the available DRAM capacity by storing the model parameters in flash memory, but bringing them on demand to DRAM. Our method involves constructing an inference cost model that takes into account the characteristics of flash memory, guiding us to optimize in two critical areas: reducing the volume of data transferred from flash and reading data in larger, more contiguous chunks. Within this hardware-informed framework, we introduce two principal techniques. First, “windowing” strategically reduces data transfer by reusing previously activated neurons, and second, “row-column bundling”, tailored to the sequential data access strengths of flash memory, increases the size of data chunks read from flash memory. These methods collectively enable running models up to twice the size of the available DRAM, with up to 4x and 20x increase in inference speed compared to naive loading approaches in CPU and GPU, respectively. Our integration of sparsity awareness, context-adaptive loading, and a hardware-oriented design paves the way for effective inference of LLMs on devices with limited memory.

Table of Contents

  • 1 Introduction
  • 2 Flash Memory & LLM Inference
  • 2.1 Bandwidth and Energy Constraints
  • 2.2 Read Throughput
  • 3 Load From Flash
  • 3.1 Reducing Data Transfer
  • 3.2 Increasing Transfer Throughput
  • 3.3 Optimized Data Management in DRAM
  • 4 Experiments and Results
  • 4.1 Experimental Setup
  • 4.2 Faster Load From Flash
  • 4.3 The Memory-Latency Tradeoff
  • 5 Ablation analysis
  • 5.1 The Impact of Longer Generation
  • 5.2 Speculative Decoding
  • 5.3 A Note on Power Consumption
  • 6 Related Works
  • 7 Discussion
  • 8 Limitations
  • Acknowledgements
  • References
  • A Appendix Overview
  • B Low-Rank Activation Predictor: Additional Results
  • B.1 Sparsity patterns of predictors
  • B.2 Accuracy of models using predictors
  • B.3 Overhead of predictors
  • C Extended Results
  • C.1 Results for OPT 6.7B Model
  • C.2 Results for Falcon 7B Model
  • C.3 Persimmon 8B
  • C.4 Phi-2
  • C.5 Llama 2
  • D Bundling Based on Co-activation
  • E Extended Related Works
  • F Small Device Implications
  • G Qualitative Evaluations

Knowls

  1. Knowl 1 — Flash-resident inference cost model

    model/method

    The paper introduces an inference design for transformer language models whose full parameter set does not fit in device DRAM or GPU memory. All model parameters are stored in flash, while only parameters needed for the current generation step are transferred into DRAM or accelerator-accessible memory. The latency objective is decomposed into three components:

    L=Lflash+Lmemory+Lcompute,L = L_{\mathrm{flash}} + L_{\mathrm{memory}} + L_{\mathrm{compute}},

    where LL is total per-token inference latency, LflashL_{\mathrm{flash}} is the time to transfer parameters from flash to DRAM, LmemoryL_{\mathrm{memory}} is the cost of inserting and removing loaded parameters in DRAM, and LcomputeL_{\mathrm{compute}} is the neural-network computation time. The proposed techniques target the first two terms rather than optimizing computation itself. By combining selective weight loading, larger flash reads, and efficient DRAM management, the system is intended to run models up to approximately twice the available DRAM capacity.

  2. Knowl 2 — Low-rank predictor for selective FFN weight loading

    model/method

    The system exploits activation sparsity in transformer feed-forward networks (FFNs). Embedding parameters and attention-module matrices remain resident in DRAM, while FFN weights associated with inactive intermediate neurons remain in flash. For each FFN layer, a low-rank predictor receives the output of that layer's attention module and predicts which intermediate neurons will produce positive or otherwise active outputs. If N=dmodelN=d_{\mathrm{model}} is the attention-output dimension, M=dffnM=d_{\mathrm{ffn}} is the number of FFN intermediate neurons, and rr is the predictor rank, the predictor maps RN→Rr→RM\mathbb{R}^{N}\rightarrow\mathbb{R}^{r}\rightarrow\mathbb{R}^{M} and applies a sigmoid threshold to obtain an active-neuron mask. Only the predicted active neurons' up-projection columns and down-projection rows are loaded from flash.

    The predictors use balanced positive/negative training loss, 10,000 samples from the C4 training set, and two training epochs; training each predictor took approximately four hours on an NVIDIA A100 GPU. Unlike an earlier predictor design that uses the preceding FFN output, this predictor uses only the current layer's attention output, which is sufficient for the hardware-aware loading procedure. The approach relies on sparse or sparsified models: the paper reports approximately 97% FFN sparsity for OPT 6.7B, 95% for a sparsified Falcon 7B, 90% for a sparsified Llama 2, and more than 90% for the relevant sparse configurations.

  3. Knowl 3 — Sliding-window reuse of active neurons

    model/method

    To avoid repeatedly loading the same FFN weights, the system maintains in DRAM the weights of neurons predicted active for a recent window of input tokens. Let kk be the number of recent tokens represented by the window, and let sagg(k)s_{\mathrm{agg}}(k) be the number or fraction of distinct neuron weights used across those kk tokens. When a new token arrives, only the incremental set

    sagg(k+1)−sagg(k)s_{\mathrm{agg}}(k+1)-s_{\mathrm{agg}}(k)

    is fetched from flash, while weights associated with tokens that have left the window are evicted. The window size is chosen as large as the DRAM budget permits. Because the cumulative set of active neurons grows sublinearly with the number of recent tokens, the incremental transfer per new token decreases for larger windows, at the cost of retaining more FFN weights in DRAM. In the OPT 6.7B implementation, a window configuration using the recent four-token setting reduced the FFN data required per token from roughly 10% of DRAM capacity for predictor-only loading to about 2.4%.

  4. Knowl 4 — Row-column bundling for larger flash reads

    model/method

    For an FFN intermediate neuron ii, the iith column of the up-projection matrix and the iith row of the down-projection matrix are consumed together. The system stores these two pieces contiguously in flash, so a predicted active neuron can be loaded with one larger read rather than two separate small reads. If each weight element occupies bb bytes and the model dimension is dmodeld_{\mathrm{model}}, a separately stored row or column has size dmodelbd_{\mathrm{model}}b, whereas the bundled pair has size 2dmodelb2d_{\mathrm{model}}b. This doubles the nominal read chunk and amortizes flash latency-to-first-byte.

    The method may intentionally read some unneeded data and discard it, because a larger contiguous read can have higher effective throughput than several smaller sparse reads. The paper also tested bundling neurons according to co-activation patterns. Although closest-neighbor neurons co-activated frequently, that strategy repeatedly included highly active neurons and therefore increased duplicate loading rather than improving transfer efficiency.

  5. Knowl 5 — Preallocated DRAM structure for dynamic FFN weights

    model/method

    Dynamic insertion and deletion of FFN neurons can be expensive if each update reallocates and copies a matrix. The system therefore preallocates, for layer ii, a matrix with capacity Ri×2dmodelR_i\times 2d_{\mathrm{model}}, where RiR_i is the maximum number of neurons expected for the selected window and dmodeld_{\mathrm{model}} is the transformer model dimension. Each active row stores the concatenated up-projection column and down-projection row for one intermediate neuron. Associated metadata contains the original-neuron pointer, the up-projection bias or scalar, the number of occupied rows, and the set of neurons active in the last kk tokens.

    When neurons leave the window, their rows and metadata are replaced by the most recently occupied rows so that active entries remain contiguous. If cc neurons are deleted, identifying them takes linear time in the relevant neuron metadata and rewriting the corresponding matrix data costs O(c dmodel)O(c\,d_{\mathrm{model}}). Newly required rows are then appended at the end without reallocating or copying the existing matrix. During inference, the first dmodeld_{\mathrm{model}} columns of the active matrix provide the up projection, while the transposed remaining columns provide the down projection. Reordering intermediate neurons does not change the final FFN output, enabling this compact representation.

  6. Knowl 6 — Flash hardware observations guiding the design

    experimental setup

    The hardware study establishes that flash capacity is much larger than DRAM but has substantially lower transfer bandwidth: the unified-memory illustration uses approximately 100 GB of flash versus approximately 10 GB of DRAM, with roughly 1 GB/s flash-to-DRAM bandwidth versus roughly 100 GB/s for DRAM-to-CPU/GPU access. On an Apple M1 Max, a 1 GiB uncached sequential read exceeded 6 GiB/s, but small random reads were much slower because operating-system, driver, interrupt, controller, and transfer-start latencies are paid for each read.

    Random-read throughput increased with both read-chunk size and the number of concurrent threads. The paper therefore uses parallel reads over 32 threads and targets chunks of at least 32 KiB. For sparse reads, reading larger chunks—even when some contents are discarded—can be preferable to reading only the exact required weights in many small requests. This hardware behavior motivates both row-column bundling and multithreaded loading.

  7. Knowl 7 — Evaluation protocol for limited-memory inference

    experimental setup

    The experiments process one sequence at a time and reserve part of memory for the key-value cache, focusing the remaining capacity on model parameters. The primary evaluations use OPT 6.7B and a sparsified Falcon 7B, with additional evaluations on Persimmon 8B, Phi-2 2.7B, and sparsified Llama 2 7B. Each C4 validation example contributes a 128-token prompt followed by 256 generated tokens. The main condition allocates approximately half of the model's total memory requirement to model computation; Phi-2 uses a 65% allocation because its sparsity is lower and its model is smaller.

    The hardware consists of an Apple M1 Max with 1 TB flash, an Apple M2 Ultra with 2 TB flash, and a Linux system with a 24 GB NVIDIA RTX 4090. CPU execution uses float32 on Apple systems, Metal GPU execution uses float16, and the NVIDIA GPU uses bfloat16. The naive baseline loads the model on demand for each forward pass. A hybrid baseline persistently stores half of the model and loads the other half at every generation step without exploiting sparsity. The reported I/O comparisons use theoretical best-case transfer costs for these baselines, and flash benchmarks disable operating-system file caching to measure a conservative uncached throughput.

  8. Knowl 8 — Ablation of predictor, windowing, and bundling

    data/table

    The following ablation measures OPT 6.7B 16-bit inference on an Apple M1 Max when half of the model memory is available. The columns distinguish the hybrid baseline, the activation predictor, sliding-window reuse, and row-column bundling. Predictor and windowing sharply reduce the transferred volume, while bundling restores part of the throughput lost by scattered sparse reads.

    Could not parse LaTeX table

    The ablation shows that reducing data volume alone is not sufficient: sparse scattered reads reduce throughput from 6.10 GB/s to 1.25 GB/s. Bundling increases the measured throughput to 2.25 GB/s in the complete configuration and reduces I/O latency to 87 ms, compared with 2,196 ms for naive loading.

  9. Knowl 9 — End-to-end latency improvements across models and hardware

    data/table

    The complete method, denoted by “All,” uses the activation predictor, sliding-window reuse, row-column bundling, and dynamic DRAM management. The following per-inference latency measurements are in milliseconds and decompose total latency into flash I/O, DRAM management, neural computation, and their sum. The measurements use the limited-memory protocol described for the evaluation models.

    Could not parse LaTeX table

    The complete method reduces OPT 6.7B total latency from 3,182 to 669 ms on CPU, from 2,389 to 565 ms on Metal M1, from 2,270 to 305 ms on Metal M2, and from 2,218 to 84 ms on the NVIDIA GPU. It also outperforms both naive and hybrid loading for Falcon 7B, Persimmon 8B, Phi-2, and Llama 2 on CPU.

  10. Knowl 10 — Accuracy and decoding robustness of predictor-based loading

    empirical result

    For OPT 6.7B, the activation predictor preserved the reported zero-shot task scores while enabling selective loading. The scores below compare the original model with the model using predictors; Arc Easy, Arc Challenge, and HellaSwag are zero-shot accuracy percentages.

    Could not parse LaTeX table

    For the OPT configuration with rank-128 predictors in the first 28 layers and rank-1024 predictors in the final four layers, the paper reports approximately 5% false negatives and 7% false positives. False negatives were usually near-zero preactivations, so omitting them had little observed effect on the listed zero-shot scores. Predictor quality is not uniformly lossless across every sparsified model and threshold: the paper reports some metric degradation for certain Falcon, Persimmon, and Phi-2 settings, making predictor rank and threshold model-dependent.

    The method also remains effective for longer generation: generating 1,000 tokens with OPT 6.7B did not produce the expected flash-thermal-throttling degradation, and average flash latency decreased after the initially empty cache was populated. With speculative decoding on OPT 6.7B, the system used a draft length of λ=4\lambda=4 and achieved a reported 1.4× speedup over its non-speculative implementation. The paper conjectures that, if the acceptance ratio is α\alpha, retaining a window ending near token α(λ+1)\alpha(\lambda+1) is preferable to blindly retaining the most recent tokens, but presents this as a heuristic rather than a proven rule.

Coverage note — Detailed qualitative generations, the full cross-model predictor-accuracy tables, quantitative power measurements, and the paper's broader single-batch/multi-batch and prompt-processing limitations were omitted as secondary validation or scope material rather than load-bearing components of the proposed flash-loading method.

References

  1. 1.Arash Ahmadian, Saurabh Dash, Hongyu Chen, Bharat Venkitesh, Stephen Gou, Phil Blunsom, A. Ustun, and Sara Hooker. 2023. Intriguing properties of quantization at scale. ArXiv, abs/2305.19268.
  2. 2.Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Maitha Alhammadi, Mazzotta Daniele, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. The falcon series of language models: Towards open frontier models.
  3. 3.Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, et al. 2022. Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale. In SC22: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–15. IEEE.
  4. 4.Sangmin Bae, Jongwoo Ko, Hwanjun Song, and Se-Young Yun. 2023. Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding. ArXiv, abs/2310.05424.
  5. 5.Cenk Baykal, Dylan Cutler, Nishanth Dikkala, Nikhil Ghosh, Rina Panigrahy, and Xin Wang. 2023. Alternating updates for efficient transformers. ArXiv, abs/2301.13310.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  7. 7.Yew Ken Chia, Pengfei Hong, Lidong Bing, and Soujanya Poria. 2023. Instructeval: Towards holistic evaluation of instruction-tuned large language models.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  9. 9.Han Dai, Yi Zhang, Ziyu Gong, Nanqing Yang, Wei Dai, Eric Song, and Qiankun Xie. 2021. Spatten: Efficient sparse attention architecture with cascade token and head pruning. In Advances in Neural Information Processing Systems, volume 34.
  10. 10.Erich Elsen, Augustus Odena, Maxwell Nye, Sagnak Tasırlar, Tri Dao, Curtis Hawthorne, Deepak Moparthi, and Arushi Somani. 2023. Releasing Persimmon-8B.
  11. 11.Mingyu Gao, Jie Yu, Wentai Li, Michael C Dai, Nam Sung Kim, and Krste Asanovic. 2022. computedram: In-memory compute using off-the-shelf dram. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, pages 1065–1079.
  12. 12.Google Gemini Team. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  13. 13.Alex Graves. 2016. Adaptive computation time for recurrent neural networks. In International Conference on Machine Learning, pages 3500–3509. PMLR.
  14. 14.Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023. Textbooks are all you need. CoRR, abs/2306.11644.
  15. 15.Jongmin Ham, Jinha Kim, Jinwoong Choi, Cheolwoo Cho, Seulki Hong, Kyeongsu Han, and Taejoo Chung. 2016. Graphssd: a high performance flash-based storage system for large-scale graph processing. In 2016 USENIX Annual Technical Conference (USENIXATC 16), pages 243–256.
  16. 16.Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A Horowitz, and William J Dally. 2016a. Eie: efficient inference engine on compressed deep neural network. arXiv preprint arXiv:1602.01528.
  17. 17.Song Han, Huizi Mao, and William J Dally. 2016b. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. In International Conference on Learning Representations (ICLR).
  18. 18.Awni Hannun, Jagrit Digani, Angelos Katharopoulos, and Ronan Collobert. 2023. MLX: Efficient and flexible machine learning on apple silicon.
  19. 19.Zhenyu He, Zexuan Zhong, Tianle Cai, Jason D Lee, and Di He. 2023. Rest: Retrieval-based speculative decoding. ArXiv, abs/2311.08252.
  20. 20.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  21. 21.Duc Nien Hoang, Minsik Cho, Thomas Merth, Mohammad Rastegari, and Zhangyang Wang. 2023. (dynamic) prompting might be all you need to repair compressed llms. ArXiv, abs/2310.00867.
  22. 22.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  23. 23.Ajay Jaiswal, Zhe Gan, Xianzhi Du, Bowen Zhang, Zhangyang Wang, and Yinfei Yang. 2023. Compressing llms: The truth is rarely pure and never simple. ArXiv, abs/2310.01382.
  24. 24.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b. CoRR, abs/2310.06825.
  25. 25.Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2022. Fast inference from transformers via speculative decoding.
  26. 26.Liang Li, Qingyuan Li, Bo Zhang, and Xiangxiang Chu. 2023. Norm tweaking: High-performance low-bit quantization of large language models. ArXiv, abs/2309.02784.
  27. 27.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023. Awq: Activation-aware weight quantization for llm compression and acceleration. ArXiv, abs/2306.00978.
  28. 28.Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. 2023a. Llm-qat: Data-free quantization aware training for large language models. CoRR.
  29. 29.Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al. 2023b. Deja vu: Contextual sparsity for efficient llms at inference time. In International Conference on Machine Learning, pages 22137–22176. PMLR.
  30. 30.Moinuddin K Meswani, Sergey Blagodurov, David Roberts, John Slice, Mike Ignatowski, and Gabriel Loh. 2015. Neural cache: Bit-serial in-cache acceleration of deep neural networks. In 2015 48th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pages 383–394. IEEE.
  31. 31.Iman Mirzadeh, Keivan Alizadeh, Sachin Mehta, Carlo C Del Mundo, Oncel Tuzel, Golnoosh Samei, Mohammad Rastegari, and Mehrdad Farajtabar. 2023. Relu strikes back: Exploiting activation sparsity in large language models.
  32. 32.Sharan Narang, Logan Feistel, Erich Elsen Undersander, Cindy Song, and Gregory Diamos. 2022. Firefly: A lightweight system for running multi-billion parameter models on commodity hardware. In 2022 ACM/IEEE 49th Annual International Symposium on Computer Architecture (ISCA), pages 757–771. IEEE.
  33. 33.Sharan Narang, Erich Elsen Undersander, and Gregory Diamos. 2021. Sparse gpu kernels for deep learning. In International Conference on Learning Representations.
  34. 34.Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W Keckler, and William J Dally. 2017. Timeloop: A systematic approach to dnn accelerator evaluation. In 2017 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), pages 241–251. IEEE.
  35. 35.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 8024–8035.
  36. 36.Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021. Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. In SC21: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–14.
  37. 37.Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W Keckler. 2013. vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design. In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), page Article 13. IEEE Computer Society.
  38. 38.Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqiang Li, Kaipeng Zhang, Peng Gao, Yu Jiao Qiao, and Ping Luo. 2023. Omniquant: Omnidirectionally calibrated quantization for large language models. ArXiv, abs/2308.13137.
  39. 39.Yifan Shao, Mengjiao Li, Wenhao Cai, Qi Wang, Dhananjay Narayanan, and Parthasarathy Ranganathan. 2022. Hotpot: Warmed-up gigascale inference with tightly-coupled compute and reuse in flash. In Proceedings of the 55th Annual IEEE/ACM International Symposium on Microarchitecture, pages 335–349.
  40. 40.Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023. Flexgen: High-throughput generative inference of large language models with a single GPU. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 31094–31116. PMLR.
  41. 41.Chenyang Song, Xu Han, Zhengyan Zhang, Shengding Hu, Xiyu Shi, Kuai Li, Chen Chen, Zhiyuan Liu, Guangli Li, Tao Yang, and Maosong Sun. 2024. Prosparse: Introducing and enhancing intrinsic activation sparsity within large language models.
  42. 42.Vedant Subramani, Marios Savvides, Li Ping, and Sharan Narang. 2022. Adapt: Parameter adaptive token-wise inference for vision transformers. In Proceedings of the 55th Annual IEEE/ACM International Symposium on Microarchitecture.
  43. 43.Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter. 2023. A simple and effective pruning approach for large language models. ArXiv, abs/2306.11695.
  44. 44.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  45. 45.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  46. 46.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019. Huggingface’s transformers: State-of-the-art natural language processing. CoRR, abs/1910.03771.
  47. 47.Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song. 2023. Flash-llm: Enabling low-cost and highly-efficient large generative model inference with unstructured sparsity. Proc. VLDB Endow., 17:211–224.
  48. 48.Zirui Xu, Zirui Liu, Beidi Chen, Yuxin Tang, Jue Wang, Kaixiong Zhou, Xia Hu, and Anshumali Shrivastava. 2023. Compress, then prompt: Improving accuracy-efficiency trade-off of llm inference with transferable prompt. ArXiv, abs/2305.11186.
  49. 49.Rongjie Yi, Liwei Guo, Shiyun Wei, Ao Zhou, Shangguang Wang, and Mengwei Xu. 2023. Edgemoe: Fast on-device inference of moe-based large language models. ArXiv, abs/2308.14352.
  50. 50.Jinchao Zhang, Jue Wang, Huan Li, Lidan Shou, Ke Chen, Gang Chen, and Sharad Mehrotra. 2023. Draft & verify: Lossless large language model acceleration via self-speculative decoding. ArXiv, abs/2309.08168.
  51. 51.Shizhao Zhang, Han Dai, Tian Sheng, Jiawei Zhang, Xiaoyong Li, Qun Xu, Mengjia Dai, Yunsong Xiao, Chao Ma, Rui Tang, et al. 2022a. Llm quantization: Quantization-aware training for large language models. In Advances in Neural Information Processing Systems, volume 35.
  52. 52.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022b. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068.
  53. 53.Zhengyan Zhang, Yixin Song, Guanghui Yu, Xu Han, Yankai Lin, Chaojun Xiao, Chenyang Song, Zhiyuan Liu, Zeyu Mi, and Maosong Sun. 2024. Relu2 wins: Discovering efficient activation functions for sparse llms.
  54. 54.Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2023. Atom: Low-bit quantization for efficient and accurate llm serving. ArXiv, abs/2310.19102.

Citation

MLA
Alizadeh, K., et al. “LLM in a Flash: Efficient Large Language Model Inference with Limited Memory”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 12562–84, https://doi.org/10.18653/v1/2024.acl-long.678.
APA
Alizadeh, K., Mirzadeh, S. I., Belenko, D., Khatamifard, S., Cho, M., Mundo, C. C. D., Rastegari, M., & Farajtabar, M. (2024). LLM in a flash: Efficient Large Language Model Inference with Limited Memory. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12562–12584. https://doi.org/10.18653/v1/2024.acl-long.678
Chicago
Alizadeh, K., S. I. Mirzadeh, D. Belenko, et al. 2024. “LLM in a Flash: Efficient Large Language Model Inference with Limited Memory”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12562–84. https://doi.org/10.18653/v1/2024.acl-long.678.
Harvard
Alizadeh, K. et al. (2024) “LLM in a flash: Efficient Large Language Model Inference with Limited Memory”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12562–12584. Available at: https://doi.org/10.18653/v1/2024.acl-long.678.
Vancouver
1. Alizadeh K, Mirzadeh SI, Belenko D, Khatamifard S, Cho M, Mundo CCD, Rastegari M, Farajtabar M (2024) LLM in a flash: Efficient Large Language Model Inference with Limited Memory. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12562–12584

BibTeX

@inproceedings{alizadeh-etal-2024-llm,
    title = "{LLM} in a flash: Efficient Large Language Model Inference with Limited Memory",
    author = "Alizadeh, Keivan  and
      Mirzadeh, Seyed Iman  and
      Belenko, Dmitry  and
      Khatamifard, S.  and
      Cho, Minsik  and
      Del Mundo, Carlo C  and
      Rastegari, Mohammad  and
      Farajtabar, Mehrdad",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.678/",
    doi = "10.18653/v1/2024.acl-long.678",
    pages = "12562--12584"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/