Green AI

Roy SchwartzJesse DodgeNoah A. SmithOren Etzioni

article2019arXiv1,911 citations

Advocates making computational efficiency a standard evaluation metric alongside accuracy to reduce deep learning's carbon footprint and lower the financial barrier to entry for AI researchers.

Listen

Recent advances in artificial intelligence have relied on increasingly massive deep learning models that require immense computing power. Between 2012 and 2018, the computational requirements to train top models surged roughly 300,000-fold, doubling every few months and outpacing historical hardware improvements. This trend has created significant environmental costs through large carbon footprints and generated steep financial barriers that exclude researchers from smaller academic institutions, emerging economies, or resource-constrained settings.

The article evaluates the escalating computational costs across mainstream artificial intelligence research—termed Red AI—and advocates for a cultural shift toward Green AI, which prioritizes computational efficiency alongside traditional accuracy measures.

To assess the scope of the problem, the authors conducted a literature review of 60 recent papers sampled from top artificial intelligence conferences (ACL, NeurIPS, and CVPR) and analyzed the mathematical scaling relationships between model performance, model size, training dataset volume, and tuning experiments. They evaluated several potential efficiency metrics—including carbon emissions, electricity consumption, elapsed execution time, parameter counts, and total floating-point operations—to determine the best standard for cross-hardware evaluation.

The analysis produced several key findings. First, existing research overwhelmingly prioritizes performance gains over efficiency: between 75% and 90% of sampled conference papers targeted accuracy, while only 10% to 20% targeted efficiency improvements. Second, the financial and computational cost of generating a machine learning result scales linearly as the product of single-example processing cost, training dataset size, and hyperparameter tuning iterations, each of which has expanded dramatically in recent years. Third, increasing computation yields sharply diminishing returns: achieving marginal linear gains in accuracy requires exponential increases in model parameters, dataset volume, and experimentation. For instance, increasing floating-point operations by nearly 35% between leading vision models yielded only a 0.5% gain in accuracy. Fourth, floating-point operations offer the most practical, hardware-independent metric to quantify computational work, energy consumption, and model efficiency.

These findings indicate that simply pouring more computation into larger models is economically and environmentally unsustainable. A single-minded focus on accuracy benchmarks obscures the hidden financial price tag of model development, distorts research incentives, and centralizes innovation within a small group of well-funded industrial laboratories.

The article recommends that researchers report the computational cost—specifically floating-point operations and budget-accuracy curves—alongside accuracy metrics when publishing results. It urges conference organizers and reviewers to recognize and reward efficiency improvements as first-class scientific contributions. Furthermore, the article encourages continued public release of pre-trained models to reduce redundant computation and advises developers to implement data-efficient pre-training methods and early-stopping rules during tuning. While total floating-point operations do not fully capture hardware memory constraints or differing implementation qualities, the evidence strongly supports adopting standard efficiency reporting to ensure that future artificial intelligence development is both ecologically sustainable and democratically accessible.

arXiv: 1907.10597
  • Paper: Energy and Policy Considerations for Deep Learning in NLP, Emma Strubell et al. (2019). Read this empirical account of deep-learning energy use and carbon emissions first; it establishes the concrete environmental and financial costs that Green AI turns into a field-wide call for efficiency reporting.
Cover for Green AI

Abstract

The computations required for deep learning research have been doubling every few months, resulting in an estimated 300,000x increase from 2012 to 2018 [2]. These computations have a surprisingly large carbon footprint [38]. Ironically, deep learning was inspired by the human brain, which is remarkably energy efficient. Moreover, the financial cost of the computations can make it difficult for academics, students, and researchers, in particular those from emerging economies, to engage in deep learning research.

This position paper advocates a practical solution by making efficiency an evaluation criterion for research alongside accuracy and related measures. In addition, we propose reporting the financial cost or "price tag" of developing, training, and running models to provide baselines for the investigation of increasingly efficient methods. Our goal is to make AI both greener and more inclusive---enabling any inspired undergraduate with a laptop to write high-quality research papers. Green AI is an emerging focus at the Allen Institute for AI.

Table of Contents

  • 1 Introduction and Motivation
  • 2 Red AI
  • 3 Green AI
  • 3.1 Measures of Efficiency
  • 3.2 FPO Cost of Existing Models
  • 3.3 Additional Ways to Promote Green AI
  • 4 Related Work
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Definitions of Red AI and Green AI

    definition

    Red AI refers to artificial intelligence research that seeks to achieve state-of-the-art results in accuracy or task performance primarily through the expenditure of massive computational resources, essentially buying performance gains via scale.

    Green AI refers to artificial intelligence research that yields novel scientific and practical results without increasing computational cost, and ideally reducing it. Green AI advocates treating computational efficiency as an essential evaluation metric alongside accuracy and task performance to foster environmental sustainability, energy efficiency, and inclusivity across the research community.

  2. Knowl 2 — Linear Cost Decomposition of AI Model Development

    equation

    The total computational cost required to produce a machine learning result, Cost(R)\text{Cost}(R), can be approximated as scaling linearly with three primary factors:

    Cost(R)∝E⋅D⋅H\text{Cost}(R) \propto E \cdot D \cdot H

    where:

    • EE is the computational cost of processing a single example through the model architecture during training or inference.
    • DD is the size of the training dataset (controlling the number of per-example model executions across training epochs).
    • HH is the number of hyperparameter tuning experiments, architectural variants, or ablation runs conducted during model development.
  3. Knowl 3 — Floating Point Operations (FPO) as an AI Efficiency Metric

    model/method

    Floating Point Operations (FPO) quantifies the total computational work required to train, tune, or execute a neural network model. FPO is calculated analytically by assigning standardized computational costs to two base elementary operations: addition (ADD\text{ADD}) and multiplication (MUL\text{MUL}). Any complex neural network operation (including convolutions, non-linear activations, attention layers, and matrix multiplications) is evaluated as a recursive function of these two elementary operations.

    FPO provides three main advantages as an efficiency benchmark:

    1. Hardware Agnosticism: FPO is invariant to the specific hardware architecture (GPU/TPU/CPU), clock speeds, and background workloads, enabling fair comparisons across different computing environments.
    2. Physical Correlation: Analytical FPO directly measures running machine work and correlates with both execution runtime and energy consumption.
    3. Granular Operation Accounting: Unlike asymptotic time complexity (OO-notation), FPO captures constant factors and exact computational work per step.
  4. Knowl 4 — Comparative Evaluation of Efficiency Metrics in AI Research

    model/method

    Several metrics exist for evaluating the computational and environmental efficiency of machine learning models, each presenting distinct trade-offs:

    • Carbon Emission: Measures total greenhouse gas emissions (e.g., CO2e\text{CO}_2\text{e}). Although it reflects direct environmental harm, it depends heavily on the local electricity infrastructure and grid mix where hardware is hosted, making comparisons across locations or timeframes unstandardized.
    • Electricity Usage: Measures total electrical energy consumed (e.g., in kilowatt-hours). While time- and location-agnostic, it depends on specific GPU/TPU hardware architectures and cooling overhead.
    • Elapsed Real Time: Measures physical runtime. It is easy to record but fluctuates with hardware generation, cluster scheduling, memory bus contention, and parallelization schemes.
    • Number of Parameters: Measures the total learnable or fixed weights. It is hardware-independent and reflects memory footprint, but fails to distinguish model depth from width, often yielding identical parameter counts for models performing vastly different amounts of work.
    • Floating Point Operations (FPO): Analytically computes total arithmetic work (ADD\text{ADD} and MUL\text{MUL}), providing a hardware-independent, location-independent, and reproducible standard.
  5. Knowl 5 — Dominance of Accuracy Targets Over Efficiency in Top AI Conferences

    empirical result

    An empirical survey of 60 papers sampled across premier AI and computer vision conferences demonstrated an overwhelming bias toward accuracy over computational efficiency as the primary reported contribution:

    • ACL 2018: 90% of sampled papers targeted accuracy improvements; only 10% targeted computational efficiency.
    • CVPR 2019: 75% of sampled papers targeted accuracy improvements; only 20% targeted computational efficiency.
    • NeurIPS 2018: 80% of sampled papers targeted empirical accuracy improvements; 20% targeted efficiency (though 55% incorporated theoretical convergence or regret bounds).

    Across empirical machine learning conferences, research predominantly optimizes task accuracy, with minimal attention given to computational expense or inference speed.

  6. Knowl 6 — Diminishing Accuracy Returns from Compute Scaling in Vision Architectures

    empirical result

    Empirical evaluation of standard convolutional architectures on the ImageNet object classification benchmark shows that scaling floating-point operations (FPO) produces diminishing returns in accuracy:

    • Model Scaling: An increase of approximately 35%35\% in FPO between ResNet-152 and ResNeXt yields only a 0.5%0.5\% absolute gain in top-1 accuracy.
    • Parameter Count Divergence from Compute Work: AlexNet possesses more parameters (61.1 million) than ResNet-50 (25.6 million) and ResNet-152 (60.2 million), but requires substantially fewer floating-point operations (0.7×1090.7 \times 10^9 FPO vs. 11.6×10911.6 \times 10^9 FPO for ResNet-152) while achieving significantly lower top-1 accuracy (56.4%56.4\% vs. 78.4%78.4\%).
    • Depth Scaling: Scaling the depth of the ResNet architecture from 18 to 152 layers increases computation from 1.8×1091.8 \times 10^9 FPO to 11.6×10911.6 \times 10^9 FPO, producing a logarithmic relationship where linear gains in top-1 accuracy (70.1%70.1\% to 78.4%78.4\%) demand exponential increases in computational work.
  7. Knowl 7 — Budget-Accuracy Curves for Computational Reporting

    model/method

    A budget-accuracy curve evaluates machine learning models by plotting expected validation performance as a function of the total computational budget (measured in floating point operations, number of hyperparameter evaluations, or hardware hours) spent across both model training and tuning.

    Because the observed performance of an architecture depends strongly on the hyperparameter search budget available during development, reporting budget-accuracy curves allows practitioners to assess model performance under varying compute constraints, identifies whether performance gains are due to architectural improvements or extensive search budgets, and highlights optimization stability.

  8. Knowl 8 — Limitations of Floating Point Operations (FPO)

    limitation

    Reporting Floating Point Operations (FPO) as an efficiency measure in machine learning has two primary limitations:

    1. Omission of Memory Footprint and Access Costs: FPO quantifies arithmetic operations but ignores peak memory consumption, memory bandwidth saturation, and memory access patterns, which are often the limiting factors for deploying models on resource-constrained devices and contribute substantially to energy dissipation.
    2. Software Implementation Sensitivity: The actual number of operations and execution efficiency in practice depend on the quality of the software implementation, library kernels, and compiler optimizations; two implementations of the identical abstract model can perform different amounts of work.

Coverage note — Omitted background summaries of external historical compute growth figures (Amodei & Hernandez, 2018) and external NLP carbon footprint estimates (Strubell et al., 2019) as they represent prior literature rather than this paper's own contributions.

References

  1. 1.Prabal Acharyya, Sean D Rosario, Roey Flor, Ritvik Joshi, Dian Li, Roberto Linares, and Hongbao Zhang. Autopilot of cement plants for reduction of fuel consumption and emissions, 2019. ICML Workshop on Climate Change.
  2. 2.Dario Amodei and Danny Hernandez. AI and compute, 2018. Blog post.
  3. 3.James S. Bergstra, Remi Bardenet, Yoshua Bengio, and Balazs Kegl. Algorithms for hyper-parameter optimization. In Proc. of NeurIPS, 2011.
  4. 4.Alfredo Canziani, Adam Paszke, and Eugenio Culurciello. An analysis of deep neural network models for practical applications. In Proc. of ISCAS, 2017.
  5. 5.Yunpeng Chen, Jianan Li, Huaxin Xiao, Xiaojie Jin, Shuicheng Yan, and Jiashi Feng. Dual path networks. In Proc. of NeurIPS, 2017.
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In Proc. of CVPR, 2009.
  7. 7.Tim Dettmers and Luke Zettlemoyer. Sparse networks from scratch: Faster training without losing performance, 2019. arXiv:1907.04840.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proc. of NAACL, 2019.
  9. 9.Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. Show your work: Improved reporting of experimental results. In Proc. of EMNLP, 2019.
  10. 10.Jesse Dodge, Kevin Jamieson, and Noah A. Smith. Open loop hyperparameter optimization and determinantal point processes. In Proc. of AutoML, 2017.
  11. 11.Clement Duhart, Gershon Dublon, Brian Mayton, Glorianna Davenport, and Joseph A. Paradiso. Deep learning for wildlife conservation and restoration efforts, 2019. ICML Workshop on Climate Change.
  12. 12.Ariel Gordon, Elad Eban, Ofir Nachum, Bo Chen, Hao Wu, Tien-Ju Yang, and Edward Choi. MorphNet: Fast & simple resource-constrained structure learning of deep networks. In Proc. of CVPR, 2018.
  13. 13.Alon Halevy, Peter Norvig, and Fernando Pereira. The unreasonable effectiveness of data. IEEE Intelligent Systems, 24:8–12, 2009.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. of CVPR, 2016.
  15. 15.Sepp Hochreiter and Jurgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  16. 16.Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. MobileNets: Efficient convolutional neural networks for mobile vision applications, 2017. arXiv:1704.04861.
  17. 17.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In Proc. of CVPR, 2018.
  18. 18.Jonathan Huang, Vivek Rathod, Chen Sun, Menglong Zhu, Anoop Korattikara, Alireza Fathi, Ian Fischer, Zbigniew Wojna, Yang Song, Sergio Guadarrama, and Kevin Murphy. Speed/accuracy trade-offs for modern convolutional object detectors. In Proc. of CVPR, 2017.
  19. 19.Sanket Kamthe and Marc Peter Deisenroth. Data-efficient reinforcement learning with probabilistic model predictive control. In Proc. of AISTATS, 2018.
  20. 20.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Proc. of NeurIPS, 2012.
  21. 21.Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar. Hyperband: Banditbased configuration evaluation for hyperparameter optimization. In Proc. of ICLR, 2017.
  22. 22.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg. Ssd: Single shot multibox detector. In Proc. of ECCV, 2016.
  23. 23.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized bert pretraining approach, 2019. arXiv:1907.11692.
  24. 24.Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. ShuffleNet V2: Practical guidelines for efficient cnn architecture design. In Proc. of ECCV, 2018.
  25. 25.Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten. Exploring the limits of weakly supervised pretraining. In Proc. ECCV, 2018.
  26. 26.Gabor Melis, Chris Dyer, and Phil Blunsom. On the state of the art of evaluation in neural language models. In Proc. of EMNLP, 2018.
  27. 27.Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. Pruning convolutional neural networks for resource efficient inference. In Proc. of ICLR, 2017.
  28. 28.Gordon E. Moore. Cramming more components onto integrated circuits, 1965.
  29. 29.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. In Proc. of NAACL, 2018.
  30. 30.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners, 2019. OpenAI Blog.
  31. 31.Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. Xnor-net: Imagenet classification using binary convolutional neural networks. In Proc. of ECCV, 2016.
  32. 32.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proc. of CVPR, 2016.
  33. 33.David Rolnick, Priya L. Donti, Lynn H. Kaack, Kelly Kochanski, Alexandre Lacoste, Kris Sankaran, Andrew Slavin Ross, Nikola Milojevic-Dupont, Natasha Jaques, Anna Waldman-Brown, Alexandra Luccioni, Tegan Maharaj, Evan D. Sherwin, S. Karthik Mukkavilli, Konrad P. Kording, Carla Gomes, Andrew Y. Ng, Demis Hassabis, John C. Platt, Felix Creutzig, Jennifer Chayes, and Yoshua Bengio. Tackling climate change with machine learning, 2019. arXiv:1905.12616.
  34. 34.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. MobileNetV2: Inverted residuals and linear bottlenecks. In Proc. of CVPR, 2018.
  35. 35.Roy Schwartz, Sam Thomson, and Noah A. Smith. SoPa: Bridging CNNs, RNNs, and weighted finite-state machines. In Proc. of ACL, 2018.
  36. 36.Yoav Shoham, Raymond Perrault, Erik Brynjolfsson, Jack Clark, James Manyika, Juan Carlos Niebles, Terah Lyons, John Etchemendy, and Z Bauer. The AI index 2018 annual report. AI Index Steering Committee, Human-Centered AI Initiative, Stanford University. Available at http://cdn.aiindex.org/2018/AI%20Index%202018%20Annual%20Report.pdf, 202018, 2018.
  37. 37.David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587):484, 2016.
  38. 38.David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis. Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017. arXiv:1712.01815.
  39. 39.David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of Go without human knowledge. Nature, 550(7676):354, 2017.
  40. 40.Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for deep learning in NLP. In Proc. of ACL, 2019.
  41. 41.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In Proc. of ICCV, 2017.
  42. 42.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proc. of NeurIPS, 2017.
  43. 43.Tom Veniat and Ludovic Denoyer. Learning time/memory-efficient deep architectures with budgeted super networks. In Proc. of CVPR, 2018.
  44. 44.Aaron Walsman, Yonatan Bisk, Saadia Gabriel, Dipendra Misra, Yoav Artzi, Yejin Choi, and Dieter Fox. Early fusion for goal directed robotic vision. In Proc. of IROS, 2019.
  45. 45.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. SuperGLUE: A stickier benchmark for general-purpose language understanding systems, 2019. arXiv:1905.00537.
  46. 46.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proc. of ICLR, 2019.
  47. 47.Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Proc. of CVPR, 2017.
  48. 48.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. XLNet: Generalized autoregressive pretraining for language understanding, 2019. arXiv:1906.08237.
  49. 49.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. Defending against neural fake news, 2019. arXiv:1905.12616.
  50. 50.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. ShuffleNet: An extremely efficient convolutional neural network for mobile devices. In Proc. of CVPR, 2018.
  51. 51.Barret Zoph and Quoc V. Le. Neural architecture search with reinforcement learning. In Proc. of ICLR, 2017.

Citation

MLA
Schwartz, R., et al. “Green AI”. arXiv, 2019, http://arxiv.org/abs/1907.10597v3.
APA
Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2019). Green AI. arXiv. http://arxiv.org/abs/1907.10597v3
Chicago
Schwartz, R., J. Dodge, N. A. Smith, and O. Etzioni. 2019. “Green AI”. arXiv. http://arxiv.org/abs/1907.10597v3.
Harvard
Schwartz, R. et al. (2019) “Green AI”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1907.10597v3.
Vancouver
1. Schwartz R, Dodge J, Smith NA, Etzioni O (2019) Green AI. arXiv

BibTeX

@article{schwartz2019green,
  title = {Green AI},
  author = {Schwartz, Roy and Dodge, Jesse and Smith, Noah A. and Etzioni, Oren},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1907.10597v3},
  eprint = {1907.10597}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors