Built independently by an author, for readers. Read the story and support ChapterPal

keyword

memory footprint

A memory footprint refers to the total amount of primary storage or active memory, such as RAM or GPU memory, that a software program, model, or dataset occupies while executing or resident in a system. It encompasses the memory required for storing executable code, static and dynamic data, runtime buffers, and auxiliary states generated during computation, such as parameters, activations, and gradients. Managing and minimizing the memory footprint is essential in computational systems to optimize performance, prevent out-of-memory errors, and enable the deployment of demanding applications on resource-constrained hardware such as edge devices, embedded systems, or individual graphic processing units.

7 items

Scaling Up Dataset Distillation to ImageNet-1K with Constant Memory

Scaling Up Dataset Distillation to ImageNet-1K with Constant Memory

Justin Cui, Ruochen Wang, Si Si, Cho-Jui Hsieh

Why you should read this

Presents a constant-memory trajectory matching algorithm paired with a teacher-guided soft label assignment that scales dataset distillation to ImageNet-1K on a single GPU with up to 50 images per class.

Dataset Distillation is a newly emerging area that aims to distill large datasets into much smaller and highly informative synthetic ones to accelerate training and reduce storage. Among various dataset distillation methods, trajectory-matching-based methods (MTT) have achieved SOTA performance in many tasks, e.g., on CIFAR-10/100. However, due to exorbitant memory consumption when unrolling optimization through SGD steps, MTT fails to scale to large-scale datasets such as ImageNet-1K. Can we scale this SOTA method to ImageNet-1K and does its effectiveness on CIFAR transfer to ImageNet-1K? To answer these questions, we first propose a procedure to exactly compute the unrolled gradient with constant memory complexity, which allows us to scale MTT to ImageNet-1K seamlessly with ∼ 6x reduction in memory footprint. We further discover that it is challenging for MTT to handle datasets with a large number of classes, and propose a novel soft label assignment that drastically improves its convergence. The resulting algorithm sets new SOTA on ImageNet-1K: we can scale up to 50 IPCs (Image Per Class) on ImageNet-1K on a single GPU (all previous methods can only scale to 2 IPCs on ImageNet-1K), leading to the best accuracy (only 5.9% accuracy drop against full dataset training) while utilizing only 4.2% of the number of data points - an 18.2% absolute gain over prior SOTA.

Added

2026-10-05

POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging

POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging

Shishir G. Patil, Paras Jain, Prabal Dutta, Ion Stoica, Joseph Gonzalez

OrganizationsUniversity of California Berkeley

Why you should read this

Presents an integer linear programming framework that jointly optimizes activation rematerialization and paging to enable energy-efficient fine-tuning of large neural networks like BERT and ResNet-18 directly on memory-constrained microcontroller-class edge devices.

Fine-tuning models on edge devices like mobile phones would enable privacy-preserving personalization over sensitive data. However, edge training has historically been limited to relatively small models with simple architectures because training is both memory and energy intensive. We present POET, an algorithm to enable training large neural networks on memory-scarce battery-operated edge devices. POET jointly optimizes the integrated search search spaces of rematerialization and paging, two algorithms to reduce the memory consumption of backpropagation. Given a memory budget and a run-time constraint, we formulate a mixed-integer linear program (MILP) for energy-optimal training. Our approach enables training significantly larger models on embedded devices while reducing energy consumption while not modifying mathematical correctness of backpropagation. We demonstrate that it is possible to fine-tune both ResNet-18 and BERT within the memory constraints of a Cortex-M class embedded device while outperforming current edge training methods in energy efficiency. POET is an open-source project available at https://github.com/ShishirPatil/poet

Added

2026-10-03

Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs

Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs

Yeonhong Park, Jake Hyun, SangLyul Cho, Bonggeun Sim, Jae W. Lee

OrganizationsSeoul National University

Why you should read this

Proposes a post-training quantization framework and a specialized GPU serving engine that pack multiple Large Language Models of varying bit-widths into the memory footprint of a single high-precision model without sacrificing inference speed or output quality.

Recently, considerable efforts have been directed towards compressing Large Language Models (LLMs), which showcase groundbreaking capabilities across diverse applications but entail significant deployment costs due to their large sizes. Meanwhile, much less attention has been given to mitigating the costs associated with deploying multiple LLMs of varying sizes despite its practical significance. Thus, this paper introduces any-precision LLM, extending the concept of any-precision DNN to LLMs. Addressing challenges in any-precision LLM, we propose a lightweight method for any-precision quantization of LLMs, leveraging a post-training quantization framework, and develop a specialized software engine for its efficient serving. As a result, our solution significantly reduces the high costs of deploying multiple, different-sized LLMs by overlaying LLMs quantized to varying bit-widths, such as 3, 4, ..., n bits, into a memory footprint comparable to a single n-bit LLM. All the supported LLMs with varying bit-widths demonstrate state-of-the-art model quality and inference throughput, proving itself to be a compelling option for deployment of multiple, different-sized LLMs. The code is available at https://github.com/SNU-ARC/any-precision-llm.

Added

2026-10-03

Random-Access Infinite Context Length for Transformers

Random-Access Infinite Context Length for Transformers

Amirkeivan Mohtashami, Martin Jaggi

OrganizationsÉcole Polytechnique Fédérale de Lausanne

Why you should read this

Introduces landmark attention, a mechanism that uses dedicated tokens to retrieve relevant context blocks directly within attention, reducing memory and computation to allow fine-tuned models like LLaMA 7B to scale inference past 32k tokens.

While Transformers have shown remarkable success in natural language processing, their attention mechanism's large memory requirements have limited their ability to handle longer contexts. Prior approaches, such as recurrent memory or retrieval-based augmentation, have either compromised the random-access flexibility of attention (i.e., the capability to select any token in the entire context) or relied on separate mechanisms for relevant context retrieval, which may not be compatible with the model's attention. In this paper, we present a novel approach that allows access to the complete context while retaining random-access flexibility, closely resembling running attention on the entire context. Our method uses a landmark token to represent each block of the input and trains the attention to use it for selecting relevant blocks, enabling retrieval of blocks directly through the attention mechanism instead of by relying on a separate mechanism. Our approach seamlessly integrates with specialized data structures and the system's memory hierarchy, enabling processing of arbitrarily long context lengths. We demonstrate that our method can obtain comparable performance with Transformer-XL while significantly reducing the number of retrieved tokens in each step. Finally, we show that fine-tuning LLaMA 7B with our method successfully extends its context length capacity to over 32k tokens, allowing for inference at the context lengths of GPT-4. We release the implementation of landmark attention and the code to reproduce our experiments at https://github.com/epfml/landmark-attention/.

Added

2026-09-26

Resource-Efficient Neural Networks for Embedded Systems

Resource-Efficient Neural Networks for Embedded Systems

Wolfgang Roth, Günther Schindler, Bernhard Klein, Robert Peharz, Sebastian Tschiatschek, Holger Fröning, Franz Pernkopf, Zoubin Ghahramani

OrganizationsFaculty of Computer ScienceGraz University of TechnologyHeidelberg UniversityInstitute for Theoretical Computer ScienceInstitute of Computer EngineeringLaboratory of Signal Processing and Speech CommunicationUniversity of CambridgeUniversity of Vienna

Why you should read this

Presents a systematic review and empirical analysis of deep neural network optimization techniques, including quantization, pruning, and structural efficiency, to guide the trade-offs between prediction accuracy, latency, and energy consumption across embedded CPUs, GPUs, and FPGAs.

While machine learning is traditionally a resource intensive task, embedded systems, autonomous navigation, and the vision of the Internet of Things fuel the interest in resource-efficient approaches. These approaches aim for a carefully chosen trade-off between performance and resource consumption in terms of computation and energy. The development of such approaches is among the major challenges in current machine learning research and key to ensure a smooth transition of machine learning technology from a scientific environment with virtually unlimited computing resources into everyday’s applications. In this article, we provide an overview of the current state of the art of machine learning techniques facilitating these real-world requirements. In particular, we focus on resource-efficient inference based on deep neural networks (DNNs), the predominant machine learning models of the past decade. We give a comprehensive overview of the vast literature that can be mainly split into three non-mutually exclusive categories: (i) quantized neural networks, (ii) network pruning, and (iii) structural efficiency. These techniques can be applied dur-

Added

2026-09-26

Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding

Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding

Song Han, Huizi Mao, William J. Dally

OrganizationsNVIDIAStanford UniversityTsinghua University

Why you should read this

Introduces a three-stage pipeline combining weight pruning, trained quantization, and Huffman coding that reduces deep neural network sizes by up to 49x without losing accuracy, enabling fast, energy-efficient inference on resource-constrained devices.

Neural networks are both computationally intensive and memory intensive, making them difficult to deploy on embedded systems with limited hardware resources. To address this limitation, we introduce "deep compression", a three stage pipeline: pruning, trained quantization and Huffman coding, that work together to reduce the storage requirement of neural networks by 35x to 49x without affecting their accuracy. Our method first prunes the network by learning only the important connections. Next, we quantize the weights to enforce weight sharing, finally, we apply Huffman coding. After the first two steps we retrain the network to fine tune the remaining connections and the quantized centroids. Pruning, reduces the number of connections by 9x to 13x; Quantization then reduces the number of bits that represent each connection from 32 to 5. On the ImageNet dataset, our method reduced the storage required by AlexNet by 35x, from 240MB to 6.9MB, without loss of accuracy. Our method reduced the size of VGG-16 by 49x from 552MB to 11.3MB, again with no loss of accuracy. This allows fitting the model into on-chip SRAM cache rather than off-chip DRAM memory. Our compression method also facilitates the use of complex neural networks in mobile applications where application size and download bandwidth are constrained. Benchmarked on CPU, GPU and mobile GPU, compressed network has 3x to 4x layerwise speedup and 3x to 7x better energy efficiency.

Added

2026-09-07