LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Keivan AlizadehIman MirzadehDmitry BelenkoS. KhatamifardMinsik ChoCarlo C. del MundoMohammad RastegariMehrdad Farajtabar
Presents hardware-informed techniques that exploit activation sparsity and chunked data access to run large language models exceeding device DRAM directly from flash memory with up to a twentyfold speedup over standard offloading methods.
Deploying modern large language models on personal and edge devices is currently constrained by limited high-speed system memory (DRAM). Standard inference requires loading the full model into DRAM, which prevents personal devices such as smartphones and laptops from running models that exceed their memory capacity. The article evaluates a framework designed to run large language models that are up to twice the size of available DRAM by storing parameters in larger flash storage and loading them on demand during inference.
To overcome the significant latency and throughput penalties of flash storage, the article demonstrates a hardware-informed approach evaluated across diverse hardware backends, including Apple Silicon central processing units (CPUs) and graphics processing units (GPUs) as well as NVIDIA GPUs. The evaluation focuses on popular model families, such as OPT, Falcon, Persimmon, Phi-2, and Llama 2. Rather than loading full layers, the framework exploits natural neuron activation sparsity within feed-forward networks by using a lightweight low-rank predictor to forecast which weights are required for upcoming tokens. It also applies a sliding-window caching technique to retain recently activated neurons in memory and bundles corresponding weight rows and columns to double flash read chunk sizes for increased input/output throughput.
The findings show substantial performance gains over conventional on-demand loading approaches. Combining sparsity prediction, windowing, and bundling reduces data transfer by over 90% in tested configurations, enabling models twice the size of available DRAM to execute efficiently. The proposed method accelerates single-token inference speed by 4 to 5 times on CPUs, about 7 times on Apple Metal GPUs, and 20 to 25 times on NVIDIA GPUs compared to naive baseline loading. When speculative decoding is integrated on GPUs, the inference pipeline achieves an additional 1.4 times speedup without degrading baseline task accuracy or output perplexity.
These results demonstrate that high-capability language models can operate directly on consumer hardware without incurring the high manufacturing costs of expanding device DRAM. This capability enhances user privacy, reduces dependence on centralized cloud infrastructure, and lowers response latency. However, while instantaneous operational power is lower due to sparse execution, the total energy consumed across token generation is slightly higher due to the extended duration of flash memory transfers compared to fully DRAM-resident models.
Organizations developing on-device artificial intelligence should implement sparsity-aware parameter streaming and hardware-aligned memory caching to deploy larger models on constrained devices. Product teams should further explore combining this approach with 4-bit model quantization to reduce the memory footprint beneath 2 gigabytes on mobile platforms. Prior to commercial deployment, technical teams should conduct broader evaluations covering multi-batch inference, prompt processing stages, and device-level thermal and battery dissipation over sustained workloads. The reported findings provide high confidence for single-sequence, on-device text generation under tightly constrained memory conditions.
- Paper: Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time, Zichang Liu et al. (2023). Deja Vu establishes the foundational concept of dynamic contextual sparsity and neuron activation prediction during LLM inference upon which on-demand parameter loading is built.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). FlashAttention introduces the core principles of IO-aware memory hierarchies and hardware-oriented data movement optimization that guide flash-to-DRAM transfer design.
- Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention, Woosuk Kwon et al. (2023). PagedAttention demonstrates OS-inspired memory management and non-contiguous memory allocation techniques for LLMs that motivate structured memory bundling and caching strategies.
- Paper: LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, Tim Dettmers et al. (2022). LLM.int8() provides essential background on memory footprint reduction and outlier-aware matrix operations for running large models on constrained hardware.
- Paper: Scaling Embedding Layers in Language Models, Da Yu et al. (2025). SCONE extends the paradigm of offloading model components to secondary storage by caching contextualized n-gram embeddings in off-accelerator memory during inference.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). This comprehensive survey contextualizes hardware-aware IO kernels, sparse activation strategies, and efficient serving architectures across the broader landscape of modern LLM design.
- Paper: Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference, Piotr Nawrot et al. (2024). Dynamic Memory Compression complements parameter offloading by dynamically compressing runtime key-value memory during decoding to further alleviate DRAM constraints.
- Paper: ThinK: Thinner Key Cache by Query-Driven Pruning, Yuhui Xu et al. (2025). ThinK builds upon activation-aware memory reduction by applying query-driven pruning to key cache channels in constrained deployment setups.
- Paper: Efficient Reasoning on the Edge, Yelysei Bondarenko et al. (2026). Efficient Reasoning on the Edge applies hardware-constrained model optimizations and dynamic caching to execute complex multi-step reasoning models on consumer edge devices.
