Dark silicon and the end of multicore scaling
H. EsmaeilzadehEmily R. BlemRenée St. AmantKarthikeyan SankaralingamD. Burger
Demonstrates through empirical modeling across CPU and GPU topologies that the breakdown of Dennard scaling forces over half of future chip area to remain unpowered dark silicon, severely limiting multicore speedups to well below historical Moore's Law projections.
For decades, semiconductor progress relied on simultaneously shrinking transistors and increasing energy efficiency. However, the breakdown of voltage scaling has created severe power constraints for modern microprocessors. The computing industry pivoted toward multicore designs to sustain historical performance growth, but growing power density threatens to stall this approach as well.
This article evaluates how much performance improvement multicore scaling can deliver across future technology generations and measures the extent of unpowered chip area, known as dark silicon. The goal was to determine whether simply adding more processor cores remains a viable long-term strategy for computing hardware.
To conduct this assessment, the article combined device-level scaling roadmaps with empirical data from over 150 real-world processors to establish optimal power, area, and performance frontiers. The authors then built an analytical modeling framework to evaluate multiple chip organizations and topologies across realistic parallel workloads from the standard benchmark suite, maintaining fixed chip area and power constraints through future technology nodes down to 8 nanometers.
The findings show that multicore scaling will fall dramatically short of historical growth expectations. Under optimistic industry projections, processor performance will improve by an average of only 7.9 times over five technology generations, leaving a nearly 24-fold shortfall compared to historical performance doubling targets. Under more conservative projections, average performance improves by only 3.7 times. Power constraints directly force severe chip underutilization: at the 22-nanometer node, at least 21% of a chip must remain unpowered dark silicon, rising to more than 50% at 8 nanometers. Furthermore, neither massively threaded graphic architectures nor topology variations such as asymmetric or dynamic designs can close this performance gap.
These results demonstrate that the multicore era cannot sustain the historical economics of semiconductor scaling. Without major shifts, the semiconductor industry faces a severe return-on-investment wall where manufacturing smaller transistors yields diminishing practical performance gains.
To prevent performance stagnation, computer architects and hardware developers must move beyond adding standard processor cores. The article recommends pursuing radical architectural innovations, such as highly specialized accelerators and alternative low-power microarchitectures, to fundamentally shift energy-efficiency frontiers.
These projections rely on optimistic assumptions, including idealized thread parallelism, zero synchronization overheads, and favorable memory scaling. In real deployments, memory bottlenecks and non-core component power will likely degrade performance further. The conclusions therefore provide a reliable, robust upper bound, confirming with high confidence that traditional multicore scaling cannot maintain historical computing growth.
- Paper: Real-time dynamic voltage scaling for low-power embedded operating systems, Padmanabhan Pillai et al. (2001). Its deadline-aware voltage-scaling algorithms show how reducing processor voltage and frequency trades performance for energy, grounding the power constraint that motivates dark-silicon analysis.
- Paper: Scheduling for reduced CPU energy, Mark Weiser et al. (1994). Its operating-system approach to dynamic CPU speed and voltage adjustment makes the energy benefits—and performance costs—of voltage scaling concrete before the source examines their hardware limits.
- Paper: Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks, Yu-hsin Chen et al. (2016). It answers the source’s call to move beyond general-purpose cores by designing a CNN accelerator that improves energy efficiency through dataflow and reduced data movement.
- Paper: In-datacenter performance analysis of a tensor processing unit, Norman P. Jouppi et al. (2017). It carries the source’s proposed shift toward specialized accelerators into production, measuring the performance and energy advantages of a datacenter tensor-processing unit.
