PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation
Jason AnselEdward YangHorace HeN. GimelsheinAnimesh JainMichael VoznesenskyBin BaoPeter BellD. BerardEvgeni Burovski
Presents TorchDynamo and TorchInductor, the core compiler systems powering PyTorch 2's torch.compile, which achieve significant training and inference speedups across real-world models by dynamically capturing computation graphs from Python bytecode without sacrificing language flexibility.
Modern machine learning workflows predominantly use eager-mode frameworks like PyTorch because they are flexible and straightforward to debug. However, eager execution evaluates operations one by one, preventing the automated compiler optimizations across operators—such as kernel fusion—that graph-mode systems use to achieve peak performance. Prior attempts to capture computation graphs in PyTorch frequently compromised usability by producing incorrect behaviors, requiring extensive code rewrites, or imposing significant runtime delays.
The article evaluates TorchDynamo and TorchInductor, two open-source extensions integrated into PyTorch 2 through the "torch.compile" interface. The main objective is to demonstrate that these tools can reliably capture computation graphs without sacrificing Python’s flexibility while delivering substantial training and inference speedups on both graphic processing units (GPUs) and central processing units (CPUs).
The authors tested this approach using extensive benchmarking across more than 180 real-world machine learning models drawn from three prominent suites: TorchBench, HuggingFace, and TIMM. They compared TorchDynamo’s graph capture reliability and overhead against earlier technologies, and evaluated TorchInductor’s execution speeds against six existing deep learning compiler backends across various precision levels and hardware configurations.
The findings show that TorchDynamo captured graphs across 93% of diverse TorchBench models and 100% of HuggingFace and TIMM models, vastly outperforming legacy tools like TorchScript, which failed entirely on HuggingFace and worked on only 45% of TorchBench. Steady-state capture overhead for TorchDynamo remained minimal at 1% to 5%, avoiding the large delays seen in systems like Lazy Tensors. TorchInductor delivered a geometric mean speedup of 2.27x for inference and 1.41x for training on NVIDIA A100 GPUs, consistently outperforming the six alternative compilers across GPU and CPU benchmarks. Ablation analysis revealed that combining operations via kernel inlining and fusion was the single largest contributor to these performance gains.
These results demonstrate that organizations can achieve compilation-level efficiency and reduce hardware computation costs without refactoring existing model code or retraining engineers in rigid graph-mode frameworks. By gracefully handling partial graphs when unsupported Python features or third-party libraries are encountered, the system provides practical acceleration while preserving standard development workflows.
For technical teams seeking to optimize machine learning performance, adopting the latest release of PyTorch 2 and enabling compilation via the standard compiler interface is recommended. Teams should test workloads using provided tuning options, such as autotuning and CUDA Graphs, to identify optimal performance configurations. While the system demonstrates high reliability across diverse architectures, users should note that dynamic rank inputs are unsupported, unbacked symbolic integers from data-dependent operations can cause graph breaks, and speedup magnitudes will vary based on model architecture and hardware setup.
- Paper: PyTorch: An Imperative Style, High-Performance Deep Learning Library, Adam Paszke et al. (2019). It introduces the foundational imperative, eager-execution architecture of PyTorch that PyTorch 2 directly compiles and accelerates through dynamic graph capture.
- Paper: Triton: an intermediate language and compiler for tiled neural network computations, Philippe Tillet et al. (2019). It presents Triton, the underlying intermediate language and compiler backend that TorchInductor heavily targets to generate fused, high-performance GPU kernels.
- Paper: TVM: an automated end-to-end optimizing compiler for deep learning, Tianqi Chen et al. (2018). It establishes foundational deep learning compiler techniques for automated operator fusion and code generation that motivate and contrast with PyTorch 2's compilation approach.
- Paper: Automatic differentiation in PyTorch, Adam Paszke et al. (2017). It outlines the core tape-based automatic differentiation design of PyTorch that TorchDynamo and AOTAutograd trace and compile into static subgraphs.
No sufficiently relevant recommendations were found.
