A single GPU refers to a computing hardware setup or execution environment that relies on only one graphics processing unit to perform computational workloads. In parallel computing and machine learning, this architecture contrasts with multi-GPU configurations and distributed computing clusters, constraining all specialized parallel processing and high-bandwidth video memory to the physical limits of that solitary accelerator. Because high-demand tasks such as deep learning model training or inference often exceed the onboard memory capacity of an individual processor, single-GPU execution frequently requires specialized software techniques, including memory offloading, compression, and precision reduction, to efficiently process large datasets and models on a single device.