POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging

Shishir G. PatilParas JainPrabal DuttaIon StoicaJoseph Gonzalez

article2022ICML58 citations

Presents an integer linear programming framework that jointly optimizes activation rematerialization and paging to enable energy-efficient fine-tuning of large neural networks like BERT and ResNet-18 directly on memory-constrained microcontroller-class edge devices.

Listen

Deploying machine learning models directly on edge devices such as smartphones and microcontrollers enables private personalization and offline functionality. However, while executing model inference on edge hardware is common, training or fine-tuning advanced neural networks remains constrained by severe memory and battery limitations. Existing approaches to lower training memory either discard and recompute intermediate calculations (rematerialization) or offload them to auxiliary storage (paging), but using either technique in isolation leads to excessive energy consumption and execution delays.

The article demonstrates an algorithmic framework called POET (Private Optimal Energy Training) that enables training modern, large neural networks on memory-constrained, battery-powered edge hardware. It evaluates how integrating and optimizing both rematerialization and paging within a unified schedule can minimize total energy consumption while strictly adhering to hardware memory budgets and runtime deadlines.

To accomplish this, the authors formulated the training schedule as a mixed-integer linear program. POET incorporates hardware-specific profiles of execution time, memory allocation, and power draw across neural network operations. It then determines mathematically optimal execution schedules that decide whether to retain, recompute, or offload intermediate tensors without altering the mathematical correctness of standard backpropagation. The framework was evaluated across four diverse edge hardware platforms—ranging from a 32 KB microcontroller to an 8 GB embedded module—using standard vision and language models including ResNet-18, VGG16, and BERT.

The evaluation revealed several key findings. First, combining rematerialization for cheap operations with paging for compute-heavy operations reduces energy overhead by up to 141% compared to heuristic paging methods like Capuchin. Second, POET successfully enabled the fine-tuning of large models like ResNet-18 and BERT on microcontroller-class devices with as little as 32 KB to 256 KB of internal memory. Third, compared to prior integrated systems like POFO, POET lowered peak memory consumption by 8.3% and improved training throughput by 13% while generalizing to complex, non-linear model architectures. Fourth, the optimizer generated schedules that consume up to 35% less energy than state-of-the-art rematerialization baselines while maintaining strict deadline constraints.

These findings indicate that organizations can achieve local, privacy-preserving model personalization on existing edge hardware without sacrificing model accuracy or investing in expensive device hardware upgrades. By ensuring mathematical equivalence to standard training, POET removes the risk of model degradation common in approximation methods. This unlocks viable offline learning workflows in connectivity-constrained or privacy-sensitive settings, including healthcare, speech recognition, and remote sensing.

Organizations planning edge deployments should profile target hardware and adopt integrated scheduling to optimize on-device learning during device idle times. While the results demonstrate high confidence across the tested architectures, the solver time was capped at 10 minutes, which led to timeouts on complex models like BERT on lower-tier hardware; extending optimization time or adopting heuristic solvers may be required for larger production networks. Future implementations should also evaluate broader combinations of activation compression and parameter paging.

  • Paper: Training Deep Nets with Sublinear Memory Cost, Tianqi Chen et al. (2016). Read this foundational work on dropping and recomputing intermediate activations first; it establishes the rematerialization strategy that POET integrates with paging.

No sufficiently relevant recommendations were found.

Cover for POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging

Abstract

Fine-tuning models on edge devices like mobile phones would enable privacy-preserving personalization over sensitive data. However, edge training has historically been limited to relatively small models with simple architectures because training is both memory and energy intensive. We present POET, an algorithm to enable training large neural networks on memory-scarce battery-operated edge devices. POET jointly optimizes the integrated search search spaces of rematerialization and paging, two algorithms to reduce the memory consumption of backpropagation. Given a memory budget and a run-time constraint, we formulate a mixed-integer linear program (MILP) for energy-optimal training. Our approach enables training significantly larger models on embedded devices while reducing energy consumption while not modifying mathematical correctness of backpropagation. We demonstrate that it is possible to fine-tune both ResNet-18 and BERT within the memory constraints of a Cortex-M class embedded device while outperforming current edge training methods in energy efficiency. POET is an open-source project available at https://github.com/ShishirPatil/poet

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background
  • 4. Integrated paging and rematerialization
  • 5. POET: Private Optimal Energy Training
  • 5.1. Optimal Rematerialization
  • 5.2. Optimal integrated paging and rematerialization
  • 5.3. Expressing an energy consumption objective
  • 5.4. Ensuring minimum training throughput
  • 5.5. Paging latency hiding via transfer planner
  • 6. Evaluation
  • 6.1. Experimental setup
  • 6.2. How much energy consumption does POET reduce across models and platforms?
  • 6.3. How does POET benefit from integrated rematerialization and paging?
  • 6.4. How does POET adapt to varying runtimes?
  • 7. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — POET compiles a training graph into an energy-aware edge execution schedule

    model/method

    POET (Private Optimal Energy Training) is a graph-level compiler that schedules training for memory-constrained edge devices. It profiles the target device and model operators for memory use, runtime, and energy, then uses those costs to choose when to retain, recompute, or page activation tensors to secondary storage. A solver finds the schedule under user-specified memory and training-time limits; POET then turns that schedule into an execution graph for the device. Profiling is automated and performed for a given workload and hardware, while the solved schedule is small enough to ship to the edge device.

  2. Knowl 2 — A joint integer program searches rematerialization and paging schedules

    model/method

    POET represents the forward-and-backward training computation as a directed acyclic graph with a topological execution order. Its integer program jointly chooses whether to compute or recompute each operation, keep each result resident in RAM, page it out to auxiliary storage, or page it back in. Binary decision variables describe these choices at each schedule step. Constraints enforce operation dependencies, valid transitions between RAM and auxiliary-storage residency, and availability of a tensor before it is used or transferred. The program also enforces a peak RAM limit. Since paging and recomputation choices are solved together, a decision to page one tensor can account for the recomputation costs it creates elsewhere in the graph. The formulation permits repeated paging and rematerialization of an operation; it is not restricted to one such action per tensor. When the solver returns an optimum, the schedule has minimum energy among the schedules represented by this formulation and constraints.

  3. Knowl 3 — The schedule minimizes profiled energy subject to a training deadline

    equation

    For each operation ii in the training graph and schedule step tt, let Rt,iR_{t,i} indicate that the operation is computed or recomputed, It,iI_{t,i} indicate a page-in, and Ot,iO_{t,i} indicate a page-out. These are binary decision variables. Let EicompE_i^{\mathrm{comp}}, EiinE_i^{\mathrm{in}}, and EioutE_i^{\mathrm{out}} be the profiled energy costs, in joules, for those respective actions. POET minimizes their total energy:

    min⁡∑t∑i(Rt,iEicomp+It,iEiin+Ot,iEiout).\min \sum_t \sum_i \left(R_{t,i}E_i^{\mathrm{comp}} + I_{t,i}E_i^{\mathrm{in}} + O_{t,i}E_i^{\mathrm{out}}\right).

    Let τicomp\tau_i^{\mathrm{comp}} be the profiled compute runtime of operation ii, in seconds, and let DD be the maximum permitted training time, in seconds. The MILP enforces the compute-runtime deadline

    ∑t∑iRt,iτicomp≤D.\sum_t \sum_i R_{t,i}\tau_i^{\mathrm{comp}} \le D.

    The energy costs are profiled on the target hardware, including the energy of the relevant device components during computation and data transfer. POET also profiles paging latency and handles transfer timing in its execution planner.

  4. Knowl 4 — Rematerialization and paging suit different operator costs

    model/method

    POET combines two ways to reduce activation memory because their energy and runtime trade-offs differ. It favors rematerialization for memory-intensive but cheap-to-compute operations: the activation is discarded and later recomputed. It favors paging for compute-intensive operations whose recomputation would be costly: the activation is copied to secondary storage and later restored. In the paper's illustrative eight-layer schedule, inexpensive nonlinear operations are rematerialized, while more expensive convolution or matrix-multiplication results are paged. The joint optimizer can mix these choices within one training schedule rather than applying a paging-first or recomputation-first rule.

  5. Knowl 5 — A transfer planner schedules page-ins around latency and bus contention

    model/method

    After the integer program selects page-in, page-out, and recomputation actions, POET's transfer planner converts the schedule into a fine-grained operator execution plan. It uses profiled paging latency to arrange transfers so tensors arrive when needed. If planned page-ins contend for the memory bus, the planner can start one earlier and update the corresponding RAM-residency schedule. The execution plan also includes computation and deallocation actions, freeing tensors when they are no longer needed or are scheduled for rematerialization. This planning can exploit overlap between transfers and computation where the hardware permits it.

  6. Knowl 6 — POET preserves the training computation under stated device assumptions

    assumption

    POET changes when activation values are stored, recomputed, or transferred; it does not approximate the values or alter the training routine. Its schedules therefore preserve the mathematical computation of backpropagation. The formulation assumes sequential operation execution without inter-operator parallelism, and that model parameters and gradients occupy a contiguous memory region that is not paged. It assumes auxiliary storage such as flash or an SD card is available; if it is not, POET can fall back to rematerialization alone. The supported scheduling graph can be non-linear, rather than being limited to a chain of layers.

  7. Knowl 7 — Evaluation spans four edge platforms and three model workloads

    experimental setup

    POET was evaluated on four battery-powered devices: an Arduino MKR1000 with a Cortex-M0-class processor (48 MHz, 32 KB RAM, no floating-point unit); an nrf52840 with a Cortex-M4F-class processor (64 MHz, 256 KB RAM); a Raspberry Pi 4B+ with an A72-class processor (1.5 GHz, 2 GB RAM); and an Nvidia Jetson TX2 with an A57-class processor (2 GHz, 8 GB RAM). Each device had at least 32 GB of flash storage available for paging. Workloads were VGG16 and ResNet-18 trained on CIFAR-10, and BERT; experiments used batch size 1. Baselines included PyTorch's default scheduling and several rematerialization methods, with additional comparisons against Capuchin and POFO. MILP solves were capped at 10 minutes on commodity CPUs, and operator costs were profiled for the target hardware.

  8. Knowl 8 — POET reduces energy across varied models and devices

    empirical result

    Across the tested ResNet-18, VGG, and BERT workloads and four hardware platforms, POET produced the lowest-energy schedules in most evaluated configurations while operating under the imposed memory and timing constraints. The multi-panel evaluation chart compares energy relative to a full-memory training configuration as available activation RAM is reduced; POET's schedules generally retain lower relative energy than the rematerialization baselines. For ResNet-18 on the Jetson TX2, the paper reports up to 35% less energy than DTR at tighter memory budgets. BERT solves on the Cortex-M4 and TX2 timed out at the 10-minute solver limit, so those cases do not establish that POET had found an optimum within that budget.

  9. Knowl 9 — POET improves on paging-first scheduling baselines

    empirical result

    Two comparisons illustrate the benefit of POET's joint scheduling. For ResNet-18 on the Raspberry Pi 4B+, Capuchin's paging-first heuristic incurred 73%–141% more energy than the full-memory baseline, while POET's overhead was below 1%; the paper's chart caption gives Capuchin's upper endpoint as 140%. For ResNet-18 on the Jetson TX2, POET used 285,873 kB of memory and took 82.36 ms, compared with POFO's 311,808 kB and 94.79 ms. These correspond to 8.3% lower peak memory and 13% shorter runtime for POET in that comparison.

  10. Knowl 10 — The best paging–rematerialization mix depends on the deadline

    empirical result

    An ablation on VGG trained on CIFAR-10 compared POET using paging only, rematerialization only, and both strategies together under per-epoch runtime budgets of 0.5, 0.6, 0.8, and 0.9 ms. At the looser budgets of 0.6–0.9 ms, the integrated solution's energy closely followed the rematerialization-only solution, which was more energy-efficient than paging alone. At the tight 0.5 ms budget, paging was preferable because rematerialization adds serial computation, whereas paging latency can be hidden by scheduling transfers alongside computation. The integrated schedule consumed up to 40% less energy than either single-strategy solution.

Coverage note — The proposed future extensions—activation compression and paging model parameters—are omitted because the paper presents them as directions for future work, not evaluated contributions.

References

  1. 1.Beaumont, O., Eyraud-Dubois, L., and Shilova, A. Efficient combination of rematerialization and offloading for training dnns. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 23844–23857. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/c8461bf13fca8a2b9912ab2eb1668e4b-Paper.pdf.
  2. 2.Blalock, D., Gonzalez Ortiz, J. J., Frankle, J., and Guttag, J. What is the state of neural network pruning? In Dhillon, I., Papailiopoulos, D., and Sze, V. (eds.), Proceedings of Machine Learning and Systems, volume 2, pp. 129–146, 2020. URL https://proceedings.mlsys.org/paper/2020/file/d2ddea18f00665ce8623e36bd4e3c7c5-Paper.pdf.
  3. 3.Cai, H., Zhu, L., and Han, S. ProxylessNAS: Direct neural architecture search on target task and hardware. In International Conference on Learning Representations, 2019. URL https://arxiv.org/pdf/1812.00332.pdf.
  4. 4.Chen, J., Zheng, L., Yao, Z., Wang, D., Stoica, I., Mahoney, M. W., and Gonzalez, J. E. Actnn: Reducing training memory footprint via 2-bit activation compressed training. In International Conference on Machine Learning, 2021.
  5. 5.Chen, T., Xu, B., Zhang, C., and Guestrin, C. Training deep nets with sublinear memory cost. CoRR, abs/1604.06174, 2016. URL http://arxiv.org/abs/1604.06174.
  6. 6.Dennis, D. K., Gopinath, S., Gupta, C., Kumar, A., Kusupati, A., Patil, S. G., and Simhadri, H. V. EdgeML: Machine Learning for resource-constrained edge devices. 2019. URL https://github.com/Microsoft/EdgeML. http://github.com/Microsoft/EdgeML.
  7. 7.Devlin, J., Chang, M., Lee, K., and Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. CoRR, abs/1810.04805, 2018. URL http://arxiv.org/abs/1810.04805.
  8. 8.Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K. Hawq: Hessian aware quantization of neural networks with mixed-precision. In The IEEE International Conference on Computer Vision (ICCV), October 2019.
  9. 9.Frankle, J. and Carbin, M. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=rJl-b3RcF7.
  10. 10.Griewank, A. and Walther, A. Algorithm 799: Revolve: An implementation of checkpointing for the reverse or adjoint mode of computational differentiation. ACM Trans. Math. Softw., 26(1):19–45, March 2000. ISSN 0098-3500. doi: 10.1145/347837.347846. URL https://doi.org/10.1145/347837.347846.
  11. 11.Hard, A., Rao, K., Mathews, R., Ramaswamy, S., Beaufays, F., Augenstein, S., Eichner, H., Kiddon, C., and Ramage, D. Federated learning for mobile keyboard prediction. arXiv preprint arXiv:1811.03604, 2018.
  12. 12.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  13. 13.Huang, C.-C., Jin, G., and Li, J. SwapAdvisor: Pushing Deep Learning Beyond the GPU Memory Limit via Smart Swapping, pp. 1341–1355. Association for Computing Machinery, New York, NY, USA, 2020. ISBN 9781450371025. URL https://doi.org/10.1145/3373376.3378530.
  14. 14.Iandola, F. N., Han, S., Moskewicz, M. W., Ashraf, K., Dally, W. J., and Keutzer, K. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and ¡0.5mb model size, 2016.
  15. 15.Jain, P., Jain, A., Nrusimha, A., Gholami, A., Abbeel, P., Keutzer, K., Stoica, I., and Gonzalez, J. E. Checkmate: Breaking the memory wall with optimal tensor rematerialization. arXiv preprint arXiv:1910.02653, 2020.
  16. 16.Jang, J. and Adib, F. Underwater backscatter networking. In Proceedings of the ACM Special Interest Group on Data Communication, SIGCOMM ’19, pp. 187–199, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450359566. doi: 10.1145/3341302.3342091. URL https://doi.org/10.1145/3341302.3342091.
  17. 17.Kapoor, G. e. a. Coreml, apple. 2019. URL https://developer.apple.com/documentation/coreml.
  18. 18.Kirisame, M., Lyubomirsky, S., Haan, A., Brennan, J., He, M., Roesch, J., Chen, T., and Tatlock, Z. Dynamic tensor rematerialization. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Vfs_2RnOD0H.
  19. 19.Lane, N. D., Bhattacharya, S., Georgiev, P., Forlivesi, C., Jiao, L., Qendro, L., and Kawsar, F. DeepX: A software accelerator for low-power deep learning inference on mobile devices. In Proceedings of the 15th International Conference on Information Processing in Sensor Networks, IPSN ’16. IEEE Press, 2016. ISBN 9781509008025.
  20. 20.Lee, J., Chirkov, N., Ignasheva, E., Pisarchyk, Y., Shieh, M., Riccardi, F., Sarokin, R., Kulik, A., and Grundmann, M. On-device neural net inference with mobile GPUs. CoRR, abs/1907.01989, 2019. URL http://arxiv.org/abs/1907.01989.
  21. 21.Levis, P., Patel, N., Culler, D., and Shenker, S. Trickle: A self-regulating algorithm for code propagation and maintenance in wireless sensor networks. In Proceedings of the 1st Conference on Symposium on Networked Systems Design and Implementation - Volume 1, NSDI’04, pp. 2, USA, 2004. USENIX Association.
  22. 22.Li, T., Sahu, A. K., Talwalkar, A., and Smith, V. Federated learning: Challenges, methods, and future directions. IEEE Signal Processing Magazine, 37(3):50–60, 2020. doi: 10.1109/MSP.2020.2975749.
  23. 23.Park, E., Ahn, J., and Yoo, S. Weighted-entropy-based quantization for deep neural networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7197–7205, July 2017. doi: 10.1109/CVPR.2017.761.
  24. 24.Patil, S. G., Dennis, D. K., Pabbaraju, C., Shaheer, N., Simhadri, H. V., Seshadri, V., Varma, M., and Jain, P. Gesturepod: Enabling on-device gesture-based interaction for white cane users. In Proceedings of the 32Nd Annual ACM Symposium on User Interface Software and Technology, UIST ’19, pp. 403–415, New York, NY, USA, 2019. ACM. ISBN 978-1-4503-6816-2. doi: 10.1145/3332165.3347881. URL http://doi.acm.org/10.1145/3332165.3347881.
  25. 25.Patterson, D., Gonzalez, J., Hölzle, U., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., and Dean, J. The carbon footprint of machine learning training will plateau, then shrink. arXiv preprint arXiv:2204.05149, 2022.
  26. 26.Paulik, M., Seigel, M., Mason, H., Telaar, D., Kluivers, J., van Dalen, R. C., Lau, C. W., Carlson, L., Granqvist, F., Vandevelde, C., Agarwal, S., Freudiger, J., Byde, A., Bhowmick, A., Kapoor, G., Beaumont, S., Cahill, A., Hughes, D., Javidbakht, O., Dong, F., Rishi, R., and Hung, S. Federated evaluation and tuning for on-device personalization: System design & applications. CoRR, abs/2102.08503, 2021. URL https://arxiv.org/abs/2102.08503.
  27. 27.Peng, Q., Shi, X., Dai, H., Jin, H., Ma, W., Xiong, Q., Yang, F., and Qian, X. Capuchin: Tensor-based gpu memory management for deep learning. In ASPLOS, March 2020. URL https://www.microsoft.com/en-us/research/publication/capuchin-tensor-based-gpu-memory-\management-for-deep-learning/.
  28. 28.Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y. ZeRO-Offload: Democratizing Billion-Scale model training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21), pp. 551–564. USENIX Association, July 2021. ISBN 978-1-939133-23-6. URL https://www.usenix.org/conference/atc21/presentation/ren-jie.
  29. 29.Shah, A., Wu, C.-Y., Mohan, J., Chidambaram, V., and Kraehenbuehl, P. Memory optimization for deep networks. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=bnY0jm4l59.
  30. 30.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  31. 31.Tan, M. and Le, Q. V. Efficientnetv2: Smaller models and faster training. CoRR, abs/2104.00298, 2021. URL https://arxiv.org/abs/2104.00298.
  32. 32.Vasisht, D., Kapetanovic, Z., Won, J., Jin, X., Chandra, R., Sinha, S., Kapoor, A., Sudarshan, M., and Stratman, S. Farmbeats: An iot platform for data-driven agriculture. In 14th {USENIX} Symposium on Networked Systems Design and Implementation ({NSDI} 17), pp. 515–529, 2017.
  33. 33.Wang, Y., Jiang, Z., Chen, X., Xu, P., Zhao, Y., Lin, Y., and Wang, Z. E2-train: Training state-of-the-art cnns with over 80% energy savings. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32, pp. 5138–5150. Curran Associates, Inc., 2019. URL http://papers.nips.cc/paper/8757-e2-train-training-state-of-the-art-cnns-with-over-80-energy-savings.pdf.

Citation

MLA
Patil, S. G., et al. “POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging”. International Conference on Machine Learning, vol. 162, 2022, pp. 17573–83, https://proceedings.mlr.press/v162/patil22b.html.
APA
Patil, S. G., Jain, P., Dutta, P., Stoica, I., & Gonzalez, J. (2022). POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging. International Conference on Machine Learning, 162, 17573–17583. https://proceedings.mlr.press/v162/patil22b.html
Chicago
Patil, S. G., P. Jain, P. Dutta, I. Stoica, and J. Gonzalez. 2022. “POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging”. International Conference on Machine Learning 162: 17573–83. https://proceedings.mlr.press/v162/patil22b.html.
Harvard
Patil, S.G. et al. (2022) “POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging”, International Conference on Machine Learning. PMLR, pp. 17573–17583. Available at: https://proceedings.mlr.press/v162/patil22b.html.
Vancouver
1. Patil SG, Jain P, Dutta P, Stoica I, Gonzalez J (2022) POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging. In: International Conference on Machine Learning. PMLR, pp 17573–17583

BibTeX

@InProceedings{pmlr-v162-patil22b,
  title = 	 {{POET}: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging},
  author =       {Patil, Shishir G. and Jain, Paras and Dutta, Prabal and Stoica, Ion and Gonzalez, Joseph},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {17573--17583},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/patil22b/patil22b.pdf},
  url = 	 {https://proceedings.mlr.press/v162/patil22b.html},
  abstract = 	 {Fine-tuning models on edge devices like mobile phones would enable privacy-preserving personalization over sensitive data. However, edge training has historically been limited to relatively small models with simple architectures because training is both memory and energy intensive. We present POET, an algorithm to enable training large neural networks on memory-scarce battery-operated edge devices. POET jointly optimizes the integrated search search spaces of rematerialization and paging, two algorithms to reduce the memory consumption of backpropagation. Given a memory budget and a run-time constraint, we formulate a mixed-integer linear program (MILP) for energy-optimal training. Our approach enables training significantly larger models on embedded devices while reducing energy consumption while not modifying mathematical correctness of backpropagation. We demonstrate that it is possible to fine-tune both ResNet-18 and BERT within the memory constraints of a Cortex-M class embedded device while outperforming current edge training methods in energy efficiency. POET is an open-source project available at https://github.com/ShishirPatil/poet}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/