Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multicore scaling

Multicore scaling is a computer architecture approach that increases overall processing performance by integrating a greater number of independent processor cores onto a single chip rather than increasing the clock frequency or complexity of an individual core. This design strategy arose as physical power and thermal limits, particularly the end of Dennard scaling, restricted the ability to continuously speed up single-threaded processors. By dividing parallel workloads across multiple cores, multicore scaling seeks to sustain generational computing gains alongside increasing transistor density, though its practical limits are governed by power consumption constraints, heat dissipation, memory communication bottlenecks, and the degree of inherent parallelism in software.

1 item

Dark silicon and the end of multicore scaling

Dark silicon and the end of multicore scaling

H. Esmaeilzadeh, Emily R. Blem, Renée St. Amant, Karthikeyan Sankaralingam, D. Burger

OrganizationsMicrosoftUniversity of Texas at AustinUniversity of WashingtonUniversity of Wisconsin Madison

Why you should read this

Demonstrates through empirical modeling across CPU and GPU topologies that the breakdown of Dennard scaling forces over half of future chip area to remain unpowered dark silicon, severely limiting multicore speedups to well below historical Moore's Law projections.

Since 2005, processor designers have increased core counts to exploit Moore’s Law scaling, rather than focusing on single-core performance. The failure of Dennard scaling, to which the shift to multicore parts is partially a response, may soon limit multicore scaling just as single-core scaling has been curtailed. This paper models multicore scaling limits by combining device scaling, single-core scaling, and multicore scaling to measure the speedup potential for a set of parallel workloads for the next five technology generations. For device scaling, we use both the ITRS projections and a set of more conservative device scaling parameters. To model single-core scaling, we combine measurements from over 150 processors to derive Pareto-optimal frontiers for area/performance and power/performance. Finally, to model multicore scaling, we build a detailed performance model of upper-bound performance and lower-bound core power. The multicore designs we study include single-threaded CPU-like and massively threaded GPU-like multicore chip organizations with symmetric, asymmetric, dynamic, and composed topologies. The study shows that regardless of chip organization and topology, multicore scaling is power limited to a degree not widely appreciated by the computing community. Even at 22 nm (just one year from now), 21% of a fixed-size chip must be powered off, and at 8 nm, this number grows to more than 50%. Through 2024, only 7.9× average speedup is possible across commonly used parallel workloads, leaving a nearly 24-fold gap from a target of doubled performance per generation.

Added

2026-09-16