POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging
Shishir G. PatilParas JainPrabal DuttaIon StoicaJoseph Gonzalez
Presents an integer linear programming framework that jointly optimizes activation rematerialization and paging to enable energy-efficient fine-tuning of large neural networks like BERT and ResNet-18 directly on memory-constrained microcontroller-class edge devices.
Deploying machine learning models directly on edge devices such as smartphones and microcontrollers enables private personalization and offline functionality. However, while executing model inference on edge hardware is common, training or fine-tuning advanced neural networks remains constrained by severe memory and battery limitations. Existing approaches to lower training memory either discard and recompute intermediate calculations (rematerialization) or offload them to auxiliary storage (paging), but using either technique in isolation leads to excessive energy consumption and execution delays.
The article demonstrates an algorithmic framework called POET (Private Optimal Energy Training) that enables training modern, large neural networks on memory-constrained, battery-powered edge hardware. It evaluates how integrating and optimizing both rematerialization and paging within a unified schedule can minimize total energy consumption while strictly adhering to hardware memory budgets and runtime deadlines.
To accomplish this, the authors formulated the training schedule as a mixed-integer linear program. POET incorporates hardware-specific profiles of execution time, memory allocation, and power draw across neural network operations. It then determines mathematically optimal execution schedules that decide whether to retain, recompute, or offload intermediate tensors without altering the mathematical correctness of standard backpropagation. The framework was evaluated across four diverse edge hardware platforms—ranging from a 32 KB microcontroller to an 8 GB embedded module—using standard vision and language models including ResNet-18, VGG16, and BERT.
The evaluation revealed several key findings. First, combining rematerialization for cheap operations with paging for compute-heavy operations reduces energy overhead by up to 141% compared to heuristic paging methods like Capuchin. Second, POET successfully enabled the fine-tuning of large models like ResNet-18 and BERT on microcontroller-class devices with as little as 32 KB to 256 KB of internal memory. Third, compared to prior integrated systems like POFO, POET lowered peak memory consumption by 8.3% and improved training throughput by 13% while generalizing to complex, non-linear model architectures. Fourth, the optimizer generated schedules that consume up to 35% less energy than state-of-the-art rematerialization baselines while maintaining strict deadline constraints.
These findings indicate that organizations can achieve local, privacy-preserving model personalization on existing edge hardware without sacrificing model accuracy or investing in expensive device hardware upgrades. By ensuring mathematical equivalence to standard training, POET removes the risk of model degradation common in approximation methods. This unlocks viable offline learning workflows in connectivity-constrained or privacy-sensitive settings, including healthcare, speech recognition, and remote sensing.
Organizations planning edge deployments should profile target hardware and adopt integrated scheduling to optimize on-device learning during device idle times. While the results demonstrate high confidence across the tested architectures, the solver time was capped at 10 minutes, which led to timeouts on complex models like BERT on lower-tier hardware; extending optimization time or adopting heuristic solvers may be required for larger production networks. Future implementations should also evaluate broader combinations of activation compression and parameter paging.
- Paper: Training Deep Nets with Sublinear Memory Cost, Tianqi Chen et al. (2016). Read this foundational work on dropping and recomputing intermediate activations first; it establishes the rematerialization strategy that POET integrates with paging.
No sufficiently relevant recommendations were found.
