Green AI
Roy SchwartzJesse DodgeNoah A. SmithOren Etzioni
Advocates making computational efficiency a standard evaluation metric alongside accuracy to reduce deep learning's carbon footprint and lower the financial barrier to entry for AI researchers.
Recent advances in artificial intelligence have relied on increasingly massive deep learning models that require immense computing power. Between 2012 and 2018, the computational requirements to train top models surged roughly 300,000-fold, doubling every few months and outpacing historical hardware improvements. This trend has created significant environmental costs through large carbon footprints and generated steep financial barriers that exclude researchers from smaller academic institutions, emerging economies, or resource-constrained settings.
The article evaluates the escalating computational costs across mainstream artificial intelligence research—termed Red AI—and advocates for a cultural shift toward Green AI, which prioritizes computational efficiency alongside traditional accuracy measures.
To assess the scope of the problem, the authors conducted a literature review of 60 recent papers sampled from top artificial intelligence conferences (ACL, NeurIPS, and CVPR) and analyzed the mathematical scaling relationships between model performance, model size, training dataset volume, and tuning experiments. They evaluated several potential efficiency metrics—including carbon emissions, electricity consumption, elapsed execution time, parameter counts, and total floating-point operations—to determine the best standard for cross-hardware evaluation.
The analysis produced several key findings. First, existing research overwhelmingly prioritizes performance gains over efficiency: between 75% and 90% of sampled conference papers targeted accuracy, while only 10% to 20% targeted efficiency improvements. Second, the financial and computational cost of generating a machine learning result scales linearly as the product of single-example processing cost, training dataset size, and hyperparameter tuning iterations, each of which has expanded dramatically in recent years. Third, increasing computation yields sharply diminishing returns: achieving marginal linear gains in accuracy requires exponential increases in model parameters, dataset volume, and experimentation. For instance, increasing floating-point operations by nearly 35% between leading vision models yielded only a 0.5% gain in accuracy. Fourth, floating-point operations offer the most practical, hardware-independent metric to quantify computational work, energy consumption, and model efficiency.
These findings indicate that simply pouring more computation into larger models is economically and environmentally unsustainable. A single-minded focus on accuracy benchmarks obscures the hidden financial price tag of model development, distorts research incentives, and centralizes innovation within a small group of well-funded industrial laboratories.
The article recommends that researchers report the computational cost—specifically floating-point operations and budget-accuracy curves—alongside accuracy metrics when publishing results. It urges conference organizers and reviewers to recognize and reward efficiency improvements as first-class scientific contributions. Furthermore, the article encourages continued public release of pre-trained models to reduce redundant computation and advises developers to implement data-efficient pre-training methods and early-stopping rules during tuning. While total floating-point operations do not fully capture hardware memory constraints or differing implementation qualities, the evidence strongly supports adopting standard efficiency reporting to ensure that future artificial intelligence development is both ecologically sustainable and democratically accessible.
- Paper: Energy and Policy Considerations for Deep Learning in NLP, Emma Strubell et al. (2019). Read this empirical account of deep-learning energy use and carbon emissions first; it establishes the concrete environmental and financial costs that Green AI turns into a field-wide call for efficiency reporting.
- Paper: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜, Emily M. Bender et al. (2021). This later critique carries Green AI’s concerns about compute costs into a broader examination of how model scaling affects the environment and equity, alongside data and social harms.
- Paper: Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities, Brian R. Bartoldson et al. (2023). This later survey advances Green AI’s efficiency agenda by systematizing algorithmic speedup methods and examining why FLOPs alone can fail to predict practical training costs.
