topic
processing core
A processing core is an independent execution unit within a processor that reads, decodes, and executes computer program instructions. Each core operates as a distinct computation engine equipped with its own arithmetic logic unit, registers, control circuitry, and dedicated cache memory. While early central processing units featured only a single core, modern microprocessors commonly integrate multiple processing cores onto a single integrated circuit to support parallel computing and simultaneous task execution. This architecture enables computing systems, ranging from embedded devices to high-performance multiprocessor servers, to process concurrent workloads efficiently while sharing system buses, higher-level caches, and main memory.
2 items

Scheduling for reduced CPU energy
Mark Weiser, Brent Welch, Alan Demers, Scott Shenker
Why you should read this
Demonstrates through trace-driven operating system simulations that dynamically scaling CPU clock speed and voltage during low-demand periods substantially reduces processor energy consumption with minimal impact on performance.
The energy usage of computer systems is becoming more important, especially for battery operated systems. Displays, disks, and cpus, in that order, use the most energy. Reducing the energy used by displays and disks has been studied elsewhere; this paper considers a new method for reducing the energy used by the cpu. We introduce a new metric for cpu energy performance, millions-of-instructions-per-joule (MIPJ). We examine a class of methods to reduce MIPJ that are characterized by dynamic control of system clock speed by the operating system scheduler. Reducing clock speed alone does not reduce MIPJ, since to do the same work the system must run longer. However, a number of methods are available for reducing energy with reduced clock-speed, such as reducing the voltage [Chandrakasan et al 1992][Horowitz 1993] or using reversible [Younis and Knight 1993] or adiabatic logic [Athas et al 1994]. What are the right scheduling algorithms for taking advantage of reduced clock-speed, especially in the presence of applications demanding ever more instructions-per-second? We consider several methods for varying the clock speed dynamically under control of the operating system, and examine the performance of these methods against workstation traces. The primary result is that by adjusting the clock speed at a fine grain, substantial CPU energy can be saved with a limited impact on performance.
Added
2026-09-25

Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks
Jierun Chen, Shiu-hong Kao, Hao He, Weipeng Zhuo, Song Wen, Chul-Ho Lee, Shueng-Han Gary Chan
Why you should read this
Proposes Partial Convolution and the FasterNet architecture family, which minimize memory access overhead to achieve significantly higher inference speeds across GPUs, CPUs, and mobile processors without sacrificing vision accuracy.
To design fast neural networks, many works have been focusing on reducing the number of floating-point operations (FLOPs). We observe that such reduction in FLOPs, however, does not necessarily lead to a similar level of reduction in latency. This mainly stems from inefficiently low floating-point operations per second (FLOPS). To achieve faster networks, we revisit popular operators and demonstrate that such low FLOPS is mainly due to frequent memory access of the operators, especially the depthwise convolution. We hence propose a novel partial convolution (PConv) that extracts spatial features more efficiently, by cutting down redundant computation and memory access simultaneously. Building upon our PConv, we further propose FasterNet, a new family of neural networks, which attains substantially higher running speed than others on a wide range of devices, without compromising on accuracy for various vision tasks. For example, on ImageNet-1k, our tiny FasterNet-T0 is , , and faster than MobileViT-XXS on GPU, CPU, and ARM processors, respectively, while being more accurate. Our large FasterNet-L achieves impressive top-1 accuracy, on par with the emerging Swin-B, while having higher inference throughput on GPU, as well as saving compute time on CPU. Code is available at \url{this https URL}.
Added
2026-09-18
