PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation

Jason AnselEdward YangHorace HeN. GimelsheinAnimesh JainMichael VoznesenskyBin BaoPeter BellD. BerardEvgeni Burovski

article2024ASPLOS1,382 citations

Presents TorchDynamo and TorchInductor, the core compiler systems powering PyTorch 2's torch.compile, which achieve significant training and inference speedups across real-world models by dynamically capturing computation graphs from Python bytecode without sacrificing language flexibility.

Listen

Modern machine learning workflows predominantly use eager-mode frameworks like PyTorch because they are flexible and straightforward to debug. However, eager execution evaluates operations one by one, preventing the automated compiler optimizations across operators—such as kernel fusion—that graph-mode systems use to achieve peak performance. Prior attempts to capture computation graphs in PyTorch frequently compromised usability by producing incorrect behaviors, requiring extensive code rewrites, or imposing significant runtime delays.

The article evaluates TorchDynamo and TorchInductor, two open-source extensions integrated into PyTorch 2 through the "torch.compile" interface. The main objective is to demonstrate that these tools can reliably capture computation graphs without sacrificing Python’s flexibility while delivering substantial training and inference speedups on both graphic processing units (GPUs) and central processing units (CPUs).

The authors tested this approach using extensive benchmarking across more than 180 real-world machine learning models drawn from three prominent suites: TorchBench, HuggingFace, and TIMM. They compared TorchDynamo’s graph capture reliability and overhead against earlier technologies, and evaluated TorchInductor’s execution speeds against six existing deep learning compiler backends across various precision levels and hardware configurations.

The findings show that TorchDynamo captured graphs across 93% of diverse TorchBench models and 100% of HuggingFace and TIMM models, vastly outperforming legacy tools like TorchScript, which failed entirely on HuggingFace and worked on only 45% of TorchBench. Steady-state capture overhead for TorchDynamo remained minimal at 1% to 5%, avoiding the large delays seen in systems like Lazy Tensors. TorchInductor delivered a geometric mean speedup of 2.27x for inference and 1.41x for training on NVIDIA A100 GPUs, consistently outperforming the six alternative compilers across GPU and CPU benchmarks. Ablation analysis revealed that combining operations via kernel inlining and fusion was the single largest contributor to these performance gains.

These results demonstrate that organizations can achieve compilation-level efficiency and reduce hardware computation costs without refactoring existing model code or retraining engineers in rigid graph-mode frameworks. By gracefully handling partial graphs when unsupported Python features or third-party libraries are encountered, the system provides practical acceleration while preserving standard development workflows.

For technical teams seeking to optimize machine learning performance, adopting the latest release of PyTorch 2 and enabling compilation via the standard compiler interface is recommended. Teams should test workloads using provided tuning options, such as autotuning and CUDA Graphs, to identify optimal performance configurations. While the system demonstrates high reliability across diverse architectures, users should note that dynamic rank inputs are unsupported, unbacked symbolic integers from data-dependent operations can cause graph breaks, and speedup magnitudes will vary based on model architecture and hardware setup.

No sufficiently relevant recommendations were found.

Cover for PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation

Abstract

This paper introduces two extensions to the popular PyTorch machine learning framework, TorchDynamo and TorchInductor, which implement the torch.compile feature released in PyTorch 2. TorchDynamo is a Python-level just-in-time (JIT) compiler that enables graph compilation in PyTorch programs without sacrificing the flexibility of Python. It achieves this by dynamically modifying Python bytecode

Table of Contents

  • 1 Introduction
  • 2 Prior Attempts at PyTorch Graph Capture
  • 2.1 torch.jit.trace
  • 2.2 torch.jit.script
  • 2.3 Lazy Tensors
  • 2.4 torch.fx.symbolic_trace
  • 2.5 torch.onnx.export
  • 2.6 Comparison To Graph Capture In JAX
  • 3 TorchDynamo Design and Implementation
  • 3.1 Usage API
  • 3.2 CPython Frame Evaluation Hook
  • 3.3 Guards
  • 3.4 Symbolic Evaluation
  • 3.5 Modeling Python Data Structures
  • 3.6 Inlining, Control Flow, and Closures
  • 3.7 Mutation and Side Effects
  • 3.8 Graph Breaks and Continuation Functions
  • 3.9 AOTAutograd
  • 4 TorchInductor Design and Implementation
  • 4.1 Design Principles and Key Technologies
  • 4.2 Decompositions
  • 4.3 Lowerings and Define-By-Run Loop-Level IR
  • 4.4 Scheduling
  • 4.5 Triton Code Generation
  • 4.6 C++ Code Generation
  • 4.7 Wrapper Codegen
  • 4.8 Related Deep Learning Compilers
  • 5 Dynamic Shapes
  • 5.1 Symbolic Shape Guards
  • 5.2 Optimizing Dynamic Shapes Reasoning
  • 5.3 Hint-Free (Unbacked) Symbolic Integers
  • 6 Experimental Results
  • 6.1 TorchDynamo's Ability to Capture Graphs
  • 6.2 Overheads of Graph Capture
  • 6.3 TorchInductor Speedups
  • 6.4 Sources of TorchInductor Speedups
  • 7 Conclusions
  • Acknowledgements
  • A Artifact Appendix
  • A.1 Abstract
  • A.2 Artifact check-list (meta-information)
  • A.3 Description
  • A.3.1 How to access.
  • A.3.2 Hardware dependencies.
  • A.3.3 Software dependencies.
  • A.4 Installation
  • A.5 Experiment workflow
  • A.6 Evaluation and expected results
  • A.7 Experiment customization
  • A.8 Notes
  • References

Knowls

  1. Knowl 1 — TorchDynamo Just-In-Time Bytecode Transformation via CPython Frame Evaluation

    model/method

    TorchDynamo is a Python-level just-in-time (JIT) compiler that intercepts Python bytecode execution to extract PyTorch operations into intermediate FX graphs while leaving unsupported Python code intact. It hooks into the CPython frame evaluation API (PEP 523), which provides an eval_frame function pointer in PyInterpreterState that overrides default evaluation (_PyEval_EvalFrameDefault).

    When a function frame is evaluated under torch.compile, TorchDynamo executes the following workflow:

    1. Filter and Cache Check: Checks whether the frame belongs to an excluded module (such as the Python standard library or NumPy) or exceeds cache limits. If previously compiled, it evaluates generated guard functions associated with the frame; if the guard returns true, it executes the cached transformed bytecode via _PyEval_EvalFrameDefault.
    2. Symbolic Bytecode Interpretation: Steps through the frame's bytecode instruction by instruction, maintaining symbolic representations of the stack, local variables, closure cells, side effects, and accumulated PyTorch operations as a torch.fx.Graph.
    3. Graph Compilation: Passes the extracted FX graph to an optimizing backend compiler (such as TorchInductor).
    4. Guard Generation: Emits a combined Python guard function that rechecks all dynamic assumptions (such as tensor properties, object types, and PyTorch global flags) on subsequent invocations.
    5. Bytecode Rewriting: Emits replacement bytecode that calls the compiled artifact, executes pending side effects, and returns or branches.
    6. Continuation Handling: If analysis encounters unsupported operations or non-inlined control flow, it generates continuation functions (resume_at_XX) to execute remaining bytecode in CPython before optionally resuming JIT analysis.
  2. Knowl 2 — TorchInductor Define-By-Run Loop-Level Intermediate Representation

    model/method

    TorchInductor translates PyTorch FX graphs into high-performance GPU (Triton) and CPU (C++ with OpenMP) code using a define-by-run loop-level intermediate representation (IR). Instead of static AST nodes for loop nests, loop bodies in the IR are defined by executable Python functions that make calls to a virtualized primitive namespace ops.*.

    The IR models tensors and computation through specific abstractions:

    • TensorBox and StorageBox: Mirror PyTorch torch.Tensor and torch.Storage abstractions to track views, layout strides, aliasing, and in-place mutations.
    • ComputedBuffer: Represents a tensor computed via generated code, configured with concrete or symbolic sizes and strides (via FixedLayout).
    • Pointwise: Represents data-parallel elementwise operations whose computation is defined by an inner Python closure that accepts symbolic index coordinates (expressed with SymPy symbols, e.g., i0,i1i_0, i_1) and emits ops.* calls.
    • Reduction and Scatter: Represent reduction operations along specific dimensions and scatter-write operations.

    The virtualized ops.* namespace consists of 54 primitive operators, including:

    • ops.load(buffer_name, sympy_index) and ops.store(buffer_name, sympy_index, value): Symbolic tensor memory accesses.
    • ops.reduction(buffer_name, reduction_type, sympy_index, value): In-place reductions supporting types such as sum, prod, min, max, argmin, argmax, any, and welford_combine.
    • ops.index_expr(expr) and ops.indirect_indexing(expr): Conversion between computed tensor values and symbolic SymPy indexing coordinates.
    • ops.masked(condition, inner_fn): Conditional execution mapping to mask operations in Triton or branch conditionals in C++.

    Because the IR is executable Python, transformations, analysis passes (such as memory tracking and strength reduction), and backend code generation are implemented by overriding the ops implementation.

  3. Knowl 3 — TorchInductor Greedy Kernel Fusion and Scheduling Algorithm

    algorithm

    TorchInductor determines kernel scheduling, buffer reuse, and fusion boundaries across intermediate representation (IR) buffers using a greedy score-based fusion algorithm.

    Every buffer in the define-by-run IR is converted into a scheduler node: standard compute kernels (SchedulerNode), external library calls or handwritten kernels (ExternKernelSchedulerNode), synchronization dependencies (NopKernelSchedulerNode), or fused kernel aggregates (FusedSchedulerNode). Dependencies are constructed by analyzing read and write sets annotated with symbolic memory addresses.

    Input: List of scheduler nodes VV, dependency graph G=(V,E)G = (V, E), fusion heuristic configurations
    Output: Ordered list of scheduled and fused kernels KK
    Initialize active fusion candidates F←∅F \leftarrow \emptyset
    for each pair of nodes (u,v)(u, v) with an edge in GG or shared inputs do
        if Scheduler.can_fuse(u, v) is True then
            score ←\leftarrow Scheduler.score_fusion(u, v)
            F←F∪{(u,v,score)}F \leftarrow F \cup \{(u, v, \text{score})\}
        end if
    end for
    while FF is not empty do
        Select (u∗,v∗,max_score)←arg⁡max⁡(u,v,s)∈Fs(u^*, v^*, \text{max\_score}) \leftarrow \arg\max_{(u, v, s) \in F} s
        F←F∖{(u∗,v∗,max_score)}F \leftarrow F \setminus \{(u^*, v^*, \text{max\_score})\}
        
        if Scheduler.can_fuse(u∗,v∗u^*, v^*) remains legal under current graph state then
            w←FuseNodes(u∗,v∗)w \leftarrow \text{FuseNodes}(u^*, v^*)
            Update graph GG: replace u∗u^* and v∗v^* with fused node ww
            Update pending fusion candidates in FF referencing u∗u^* or v∗v^* to reference ww
            
            for each neighbor nn of ww in GG do
                if Scheduler.can_fuse(w, n) is True then
                    snew←Scheduler.score_fusion(w,n)s_{new} \leftarrow \text{Scheduler.score\_fusion}(w, n)
                    F←F∪{(w,n,snew)}F \leftarrow F \cup \{(w, n, s_{new})\}
                end if
            end for
        end if
    end while
    K←TopologicalSort(G)K \leftarrow \text{TopologicalSort}(G)
    return KK

    The legality check Scheduler.can_fuse(node1, node2) verifies memory access patterns (preventing fusion when read/write address alignments or iteration spaces conflict, such as forward versus reverse iteration) and backend constraints (e.g., reduction-broadcast-reduction fusion is permitted for Triton but disallowed for C++).

    The priority metric Scheduler.score_fusion(node1, node2) ranks candidates by:

    1. Fusion category compatibility (pointwise-pointwise, pointwise-reduction, template).
    2. Estimated bytes of GPU/CPU memory traffic eliminated by keeping intermediate values in registers/cache.
    3. Proximity (shortest topological distance) between nodes in the original graph.
  4. Knowl 4 — TorchDynamo Graph Breaks and Continuation Function Generation

    model/method

    When TorchDynamo encounters a Python construct or bytecode instruction that cannot be safely captured into an FX graph—such as unsupported C-extensions, calls to external non-PyTorch libraries, or data-dependent branching on tensor values—it executes a graph break rather than failing entirely.

    To handle a graph break:

    1. The symbolic evaluator terminates the current trace and compiles the accumulated operations up to that instruction into a partial FX graph.
    2. TorchDynamo generates new bytecode that invokes the compiled partial graph and updates program state (including applying deferred mutations).
    3. It synthesizes one or more continuation functions (named resume_at_XX, where XX is the target bytecode offset) that encapsulate the remainder of the function:
    def resume_at_XX(*livevars):
        # Restore stack, exception, and local variable state
        JUMP_ABSOLUTE XX
        # Original function bytecode
    
    1. For unconditional operations or function calls, one continuation function is generated. For data-dependent Python control flow that cannot be eliminated via guards, two continuation functions are generated corresponding to both branch targets.
    2. When execution transitions into a continuation function, the CPython frame evaluation hook intercepts the new frame, allowing TorchDynamo to recursively resume symbolic tracing and compile subsequent graph fragments.
  5. Knowl 5 — AOTAutograd Ahead-of-Time Forward-Backward Graph Partitioning and Decomposition

    model/method

    AOTAutograd captures training graphs ahead of time by tracing both forward and backward computation into explicit FX graphs, eliminating the need for runtime dynamic tape-based autograd dispatch.

    The AOTAutograd workflow operates as follows:

    1. Joint Tracing with Fake Tensors: Executes PyTorch's eager autograd engine using lightweight metadata tensors (FakeTensor) carrying shapes, strides, and data types without backing memory storage. It captures a joint forward and backward computation graph.
    2. Decomposition: Recursively decomposes complex or compound PyTorch operators into a standardized functional core of primitive operators using a decomposition registry (decomposing over 191 operators / 387 including overloads), avoiding cycles until a fixed point is reached.
    3. Functionalization: Eliminates in-place mutations and view-aliasing side effects by substituting functional equivalents.
    4. Min-Cut Graph Partitioning: Applies a min-cut graph partitioning algorithm to split the joint graph into distinct forward and backward graphs. The partitioning optimizes peak memory consumption by deciding which intermediate activations should be saved for the backward pass versus which cheap activations should be rematerialized (recomputed) during the backward pass.
  6. Knowl 6 — Symbolic Shape Guards and Dynamic Shapes Reasoning

    model/method

    To compile programs with dynamic tensor dimensions (such as varying batch sizes, sequence lengths, or ragged tensors) without compiling separate kernels for every concrete shape configuration, TorchDynamo and TorchInductor employ a symbolic shape reasoning engine under the assumption that tensor ranks remain static.

    Key mechanisms include:

    • Concrete Size Hints and Straight-Line Specialization: For every symbolic dimension, TorchDynamo records a concrete size hint from the initial input. When control flow branches on shape conditions, TorchDynamo consults the hint, selects the active branch, and adds a symbolic shape guard over input dimensions rather than introducing conditional branches into the compiled IR.
    • Meta Function Propagation: Propagates symbolic shape and stride formulas from input tensors to intermediate outputs without executing compute kernels across 2,657 PyTorch operators.
    • 0/1 Specialization: Dimensions with values of 0 or 1 are specialized into static constants rather than assigned symbolic variables, ensuring correct resolution of PyTorch broadcasting semantics and optimization paths. For general symbolic variables ss, negative inference guarantees s∉{0,1}s \notin \{0, 1\}, eliminating unnecessary guard generation.
    • Unbacked Symbolic Integers: When tensor dimensions depend on data values (e.g., from .nonzero() or .item()), dynamic integers lack concrete hints. Control flow branching on unbacked integers forces a graph break, while shape bounds are bounded using constrain_range annotations.
  7. Knowl 7 — TorchDynamo Symbolic Evaluation and Guard Validation Model

    model/method

    TorchDynamo executes abstract interpretation on CPython bytecode frames using an internal class hierarchy rooted at VariableTracker.

    Key VariableTracker subclasses model Python runtime semantics:

    • TensorVariable: Tracks tensor metadata via FakeTensor and links to an fx.Proxy node inside the partially built FX graph.
    • ConstDictVariable and DataClassVariable: Model dictionary structures and dataclasses with constant keys, enabling compile-time constant propagation.
    • ListVariable and TupleVariable: Model mutable and immutable sequences of tracked symbolic variables.
    • UserFunctionVariable and UserMethodVariable: Represent callables to be symbolically inlined into the caller frame.
    • UserDefinedClassVariable and UserDefinedObjectVariable: Model user-defined classes and instances by lazily specializing on attribute access and tracking mutation.

    Every VariableTracker collects guards during symbolic evaluation and combines them via set union across operations. Guards represent dynamic assertions (over 30 distinct types, including tensor shapes, strides, data types, device identity, object types, module dictionary states, and global PyTorch execution modes). The accumulated guards are compiled into a unified Python guard function stored alongside the compiled bytecode in _PyCode_SetExtra. On subsequent calls to the frame, the guard function runs first; if any condition evaluates to false, the cache misses and re-compilation or fallback occurs.

  8. Knowl 8 — Graph Capture Robustness and Runtime Overheads of TorchDynamo Compared to Prior PyTorch Capture Methods

    data/table

    Across 188 open-source benchmark models spanning TorchBench (80 diverse models), HuggingFace (46 transformer models), and TIMM (62 vision models), TorchDynamo captures complete or partial graphs across 97% of models, whereas TorchScript fails on over half of real-world models due to unsupported Python features. TorchDynamo captures whole-program single graphs without graph breaks for 81.9% of working models, and adds less than 5% steady-state runtime overhead during capture.

    Metric TorchBench HuggingFace TIMM
    Model Count 80 46 62
    Works with TorchDynamo 74 (93%) 46 (100%) 62 (100%)
    Compare with TorchScript 36 (45%) 0 (0%) 61 (98%)
    Operators Captured 91.8% 99.8% 100%
    Mean Operators per Graph 252.8 612.6 450.7
    Mean Graphs per Model 21.1 7.7 1.0
    Models with 0 graph breaks 52 (70%) 41 (89%) 62 (100%)
    Models with 1 to 9 graph breaks 6 (8%) 1 (2%) 0 (0%)
    Models with 10+ graph breaks 16 (22%) 4 (9%) 0 (0%)

    Runtime capture overhead measured as a percentage slowdown relative to eager PyTorch execution (using identical eager kernels, float32 on NVIDIA V100 GPU on TorchBench):

    Capture Framework Inference Overhead Training Overhead
    TorchDynamo 5% 1%
    Lazy Tensors 38% 90%
    Lazy Tensors + cross-iteration pipelining 31% 86%
  9. Knowl 9 — Speedup Comparison of TorchInductor and Alternative Compiler Backends Over Eager PyTorch

    data/table

    Geometric mean speedups over eager PyTorch execution were measured using TorchDynamo frontend graph capture across TorchBench, HuggingFace, and TIMM on an NVIDIA A100 GPU and Intel Xeon 8275CL CPU. TorchInductor achieves a geometric mean speedup of 2.27×2.27\times for GPU float32 inference and 1.41×1.41\times for training across all 180+ models, outperforming alternative deep learning compilers (nvFuser, NNC, PyTorch/XLA, ONNX Runtime, TVM, and Hidet).

    TorchBench HuggingFace TIMM
    Configuration / Backend Inference Training Inference Training Inference Training
    NVIDIA A100 (float32)
    TorchDynamo-only (None) 0.95× 0.99× 1.01× 0.98× 1.00× 1.00×
    TorchInductor 2.73× 1.38× 1.47× 1.24× 2.48× 1.38×
    nvFuser 1.23× 1.04× 1.09× 1.09× 1.16× 1.03×
    NNC 1.12× 1.03× 0.98× 0.94× 1.02× 0.96×
    PyTorch/XLA 0.80× 0.73× 1.03× 0.98× 1.24× 1.11×
    ONNX Runtime 0.86× N/A 0.84× N/A 0.92× N/A
    TVM 0.16× N/A 0.09× N/A 0.13× N/A
    Hidet 0.54× N/A N/A N/A 0.30× N/A
    NVIDIA A100 (float16)
    TorchDynamo-only (None) 0.95× 0.99× 1.00× 0.97× 1.00× 1.00×
    TorchInductor 2.59× 1.50× 1.91× 1.45× 2.77× 1.50×
    nvFuser 1.27× 1.04× 1.07× 1.04× 1.13× 1.01×
    NNC 1.14× 1.03× 0.98× 0.94× 1.00× 0.95×
    PyTorch/XLA 0.82× 0.80× 1.16× 0.24× 1.59× 1.27×
    Intel Xeon CPU (float32)
    TorchDynamo-only (None) 1.00× 0.99× 1.00× 1.00× 1.06× 1.00×
    TorchInductor 1.39× 1.35× 2.54× 1.36× 2.55× 1.42×
    NNC 1.14× 0.99× 1.15× 0.92× 1.04× 0.85×
    PyTorch/XLA 0.27× 0.21× 0.38× 0.34× 0.38× 0.25×
  10. Knowl 10 — Ablation of Compiler Optimizations in TorchInductor

    data/table

    The relative speedup contributions of individual optimization passes within TorchInductor were evaluated on the HuggingFace benchmark suite (float16 precision, NVIDIA A100 GPU). Kernel fusion and lowering-time inlining represent the dominant sources of performance gain; when both are disabled, TorchInductor experiences a net slowdown relative to eager execution because fine-grained operator decompositions generate high memory traffic without subsequent fusion.

    Configuration Inference Speedup Training Speedup
    All TorchInductor optimizations 1.91× 1.45×
    Without loop/layout reordering 1.91× (-0.00) 1.28× (-0.17)
    Without matmul templates 1.85× (-0.06) 1.41× (-0.04)
    Without parameter freezing 1.85× (-0.06) 1.45× (-0.00)
    Without pattern matching 1.83× (-0.08) 1.45× (-0.00)
    Without CUDA Graphs 1.81× (-0.10) 1.37× (-0.08)
    Without fusion 1.68× (-0.23) 1.27× (-0.18)
    Without inlining 1.58× (-0.33) 1.31× (-0.14)
    Without fusion and inlining 0.80× (-1.11) 0.59× (-0.86)

    Optimizations evaluated include:

    • Fusion and Inlining: Inlining duplicates pointwise operations directly into consumer kernels during lowering; fusion combines adjacent kernels and horizontal consumer kernels during scheduling.
    • Loop/Layout Reordering: Applies a voting algorithm across kernel loop nests to align memory physical strides with iterate dimensions.
    • Matmul Templates: Replaces cuBLAS/cuDNN calls with autotuned Triton matrix multiplication kernels with fused pointwise epilogues.
    • Parameter Freezing: In-place constant folding of parameter-dependent subgraphs (inference only).
    • Pattern Matching: FX graph level peephole transformations before IR lowering.
    • CUDA Graphs: Direct kernel launch replay at the CUDA driver level, eliminating Python/C++ wrapper overheads.

Coverage note — None was omitted. All major technical contributions, algorithms, design components (TorchDynamo bytecode capture, guards, symbolic shapes, IR, scheduling/fusion, AOTAutograd), empirical comparisons, and ablation data have been fully captured.

References

  1. 1.Martin Abadi et al. 2016. TensorFlow: A system for large-scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), 265–283. https://www.usenix.org/system/files/conference/osdi16/osdi16-abadi.pdf.
  2. 2.[SW] Martín Abadi et al., TensorFlow, Large-scale machine learning on heterogeneous systems Nov. 2015. doi: 10.5281/zenodo.4724125.
  3. 3.Hameer Abbasi, Edward Z Yang, and Ralf Gommers. 2020. Improving subclassing Tensor by propagating subclass instances. https://github.com/pytorch/rfcs/blob/master/RFC-0001-torch-function-for-methods.md. (Aug. 2020).
  4. 4.Akshay Agrawal et al. 2019. TensorFlow Eager: A Multi-Stage, Python-Embedded DSL for Machine Learning. CoRR, abs/1903.01855. http://arxiv.org/abs/1903.01855 arXiv: 1903.01855.
  5. 5.Rami Al-Rfou et al. 2016. Theano: A Python framework for fast computation of mathematical expressions. CoRR, abs/1605.02688. http://arxiv.org/abs/1605.02688 arXiv: 1605.02688.
  6. 6.Jason Ansel, Shoaib Kamil, Kalyan Veeramachaneni, Jonathan Ragan-Kelley, Jeffrey Bosboom, Una-May O’Reilly, and Saman Amarasinghe. 2014. Opentuner: an extensible framework for program autotuning. In Proceedings of the 23rd International Conference on Parallel Architectures and Compilation (PACT ’14). Association for Computing Machinery, Edmonton, AB, Canada, 303–316. isbn: 9781450328098. doi: 10.1145/2628071.2628092.
  7. 7.Riyadh Baghdadi, Jessica Ray, Malek Ben Romdhane, Emanuele Del Sozzo, Abdurrahman Akkas, Yunming Zhang, Patricia Suriana, Shoaib Kamil, and Saman Amarasinghe. 2019. Tiramisu: a polyhedral compiler for expressing fast and portable code. In Proceedings of the 2019 IEEE/ACM International Symposium on Code Generation and Optimization (CGO 2019). IEEE Press, Washington, DC, USA, 193–205. isbn: 9781728114361.
  8. 8.[SW] James Bradbury et al., JAX: composable transformations of Python+NumPy programs version 0.3.13, 2018. url: http://github.com/google/jax.
  9. 9.Dino Viehland Brett Cannon. 2016. PEP 523: adding a frame evaluation API to CPython. https://peps.python.org/pep-0523/. (2016).
  10. 10.Jack Cao. 2022. PyTorch/XLA 2022 Q4 dev update. https://dev-discuss.pytorch.org/t/pytorch-xla-2022-q4-dev-update/961. (2022).
  11. 11.Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. Learning to optimize tensor programs. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (NIPS’18). Curran Associates Inc., Montréal, Canada, 3393–3404.
  12. 12.Tianqi Chen et al. 2018. TVM: an automated End-to-End optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsbad, CA, (Oct. 2018), 578–594. isbn: 978-1-939133-08-3. https://www.usenix.org/conference/osdi18/presentation/chen.
  13. 13.Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cuDNN: efficient primitives for deep learning. (2014). arXiv: 1410.0759 [cs.NE].
  14. 14.Will Constable et al. 2020. TorchBench: a collection of open source benchmarks for PyTorch performance and usability evaluation. https://github.com/pytorch/benchmark. (Sept. 2020).
  15. 15.Leonardo Dagum and Ramesh Menon. 1998. OpenMP: an industry standard API for shared-memory programming. Computational Science & Engineering, IEEE, 5, 1, 46–55.
  16. 16.ONNX Runtime developers. 2021. ONNX runtime. https://www.onnxruntime.ai. (2021).
  17. 17.Zachary DeVito et al. 2018. TorchScript. https://pytorch.org/docs/1.9.0/jit.html. (Sept. 2018).
  18. 18.Yaoyao Ding, Cody Hao Yu, Bojian Zheng, Yizhi Liu, Yida Wang, and Gennady Pekhimenko. 2023. Hidet: task-mapping programming paradigm for deep learning tensor programs. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (ASPLOS 2023). Association for Computing Machinery, Vancouver, BC, Canada, 370–384. isbn: 9781450399166. doi: 10.1145/3575693.3575702.
  19. 19.Siyuan Feng et al. 2022. TensorIR: an abstraction for automatic tensorized program optimization. (2022). arXiv: 2207.04296 [cs.LG].
  20. 20.Alan Gray. 2019. Getting started with CUDA graphs. https://developer.nvidia.com/blog/cuda-graphs/. (2019).
  21. 21.Charles R. Harris et al. 2020. Array programming with NumPy. Nature, 585, 357–362. doi: 10.1038/s41586-020-2649-2.
  22. 22.Horace He. 2019. The state of machine learning frameworks in 2019. https://thegradient.pub/state-of-ml-frameworks-2019-pytorch-dominates-research-tensorflow-dominates-industry/. (2019).
  23. 23.Mike Innes et al. 2017. On machine learning and programming languages. https://julialang.org/blog/2017/12/ml-pl/. (Dec. 2017).
  24. 24.ISO. 1998. ISO/IEC 14882:1998: Programming languages — C++. (Sept. 1998), 732. http://webstore.ansi.org/ansidocstore/product.asp?sku=ISO%2FIEC+14882%2D1998.
  25. 25.Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross B. Girshick, Sergio Guadarrama, and Trevor Darrell. 2014. Caffe: Convolutional Architecture for Fast Feature Embedding. CoRR, abs/1408.5093. http://arxiv.org/abs/1408.5093 arXiv: 1408.5093.
  26. 26.Norman P. Jouppi et al. 2017. In-datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA ’17). Association for Computing Machinery, Toronto, ON, Canada, 1–12. isbn: 9781450348928. doi: 10.1145/3079856.3080246.
  27. 27.Chris Lattner et al. 2021. MLIR: scaling compiler infrastructure for domain specific computation. In 2021 IEEE/ACM International Symposium on Code Generation and Optimization (CGO), 2–14. doi: 10.1109/CGO51591.2021.9370308.
  28. 28.Aaron Meurer et al. 2017. SymPy: symbolic computing in Python. PeerJ Computer Science, 3, (Jan. 2017). doi: 10.7717/peerj-cs.103.
  29. 29.Adrian Mönnich, Armin Ronacher, David Lord, Grey Li, Joshua Bronson, Markus Unterwaditzer, and Philip Jones. 2023. Jinja project. https://github.com/pallets/jinja. (2023).
  30. 30.NVIDIA, Péter Vingelmann, and Frank H.P. Fitzek. 2023. CUDA. https://developer.nvidia.com/cuda-toolkit. (2023).
  31. 31.
    1. ONNX. https://onnx.ai/. (2023).
  32. 32.
    1. Pytorch: an imperative style, high-performance deep learning library. Proceedings of the 33rd International Conference on Neural Information Processing Systems. Curran Associates Inc., Red Hook, NY, USA, 12 pages.
  33. 33.Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013. Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines. In Proceedings of the 34th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI ’13). Association for Computing Machinery, Seattle, Washington, USA, 519–530. isbn: 9781450320146. doi: 10.1145/2491956.2462176.
  34. 34.James Reed, Zachary DeVito, Horace He, Ansley Ussery, and Jason Ansel. 2022. Torch.fx: practical program capture and transformation for deep learning in python. In Proceedings of Machine Learning and Systems. D. Marculescu, Y. Chi, and C. Wu, (Eds.) Vol. 4, 638–651. https://proceedings.mlsys.org/paper/2022/file/ca46c1b9512a7a8315fa3c5a946e8265-Paper.pdf.
  35. 35.Elvis Saravia. 2021. Papers with Code 2021: a year in review. https://medium.com/paperswithcode/papers-with-code-2021-a-year-in-review-de75d5a77b8b. (2021).
  36. 36.Christian Sarofeen, Piotr Bialecki, Jie Jiang, Kevin Stephano, Masaki Kozuki, Neal Vaidya, and Stas Bekman. 2022. Introducing nvFuser, a deep learning compiler for PyTorch. https://pytorch.org/blog/introducing-nvfuser-a-deep-learning-compiler-for-pytorch/. (2022).
  37. 37.Frank Seide and Amit Agarwal. 2016. CNTK: microsoft’s open-source deep-learning toolkit. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ’16). Association for Computing Machinery, San Francisco, California, USA, 2135. isbn: 9781450342322. doi: 10.1145/2939672.2945397.
  38. 38.Junru Shao et al. 2022. Tensor program optimization with probabilistic programs. (2022). arXiv: 2205.13603 [cs.LG].
  39. 39.Alex Suhan, Davide Libenzi, Ailing Zhang, Parker Schuh, Brennan Saeta, Jie Young Sohn, and Denys Shabalin. 2021. LazyTensor: combining eager execution with domain-specific compilers. arXiv preprint arXiv:2102.13267.
  40. 40.PyTorch Team. 2023. TorchDynamo Benchmarking Code. https://github.com/pytorch/pytorch/tree/main/benchmarks/dynamo. (2023).
  41. 41.PyTorch Team. 2023. TorchInductor Performance Dashboard. https://hud.pytorch.org/benchmark/compilers. (2023).
  42. 42.PyTorch XLA Team. 2023. PyTorch/XLA. https://github.com/pytorch/xla. (2023).
  43. 43.Vijay Thakkar et al. 2023. CUTLASS. https://github.com/NVIDIA/cutlass. Version 3.0.0. (Jan. 2023).
  44. 44.[SW] The IREE Authors, IREE Sept. 2019. url: https://github.com/openxla/iree.
  45. 45.The XLA Team. 2017. XLA - Tensorflow, compiled. https://developers.googleblog.com/2017/03/xla-tensorflow-compiled.html. (Mar. 2017).
  46. 46.Philippe Tillet, H. T. Kung, and David Cox. 2019. Triton: an intermediate language and compiler for tiled neural network computations. In (MAPL 2019). Association for Computing Machinery, Phoenix, AZ, USA, 10–19. isbn: 9781450367196. doi: 10.1145/3315508.3329973.
  47. 47.Seiya Tokui et al. 2019. Chainer: A Deep Learning Framework for Accelerating the Research Cycle. CoRR, abs/1908.00213. http://arxiv.org/abs/1908.00213 arXiv: 1908.00213.
  48. 48.Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S. Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018. Tensor comprehensions: framework-agnostic high-performance machine learning abstractions. (2018). arXiv: 1802.04730 [cs.PL].
  49. 49.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS’17). Curran Associates Inc., Long Beach, California, USA, 6000–6010. isbn: 9781510860964.
  50. 50.B. P. Welford. 1962. Note on a method for calculating corrected sums of squares and products. Technometrics, 4, 3, 419–420. doi: 10.1080/00401706.1962.10490022.
  51. 51.Jian Weng, Animesh Jain, Jie Wang, Leyuan Wang, Yida Wang, and Tony Nowatzki. 2021. Unit: unifying tensorized instruction compilation. In 2021 IEEE/ACM International Symposium on Code Generation and Optimization (CGO), 77–89. doi: 10.1109/CGO51591.2021.9370330.
  52. 52.Ross Wightman. 2019. PyTorch image models. https://github.com/rwightman/pytorch-image-models. (2019). doi: 10.5281/zenodo.4414861.
  53. 53.Thomas Wolf et al. 2020. Transformers: State-of-the-Art Natural Language Processing. In Association for Computational Linguistics, (Oct. 2020), 38–45. https://www.aclweb.org/anthology/2020.emnlp-demos.6.
  54. 54.Jiarong Xing, Leyuan Wang, Shang Zhang, Jack Chen, Ang Chen, and Yibo Zhu. 2022. Bolt: bridging the gap between auto-tuners and hardware-native performance. In Proceedings of Machine Learning and Systems. D. Marculescu, Y. Chi, and C. Wu, (Eds.) Vol. 4, 204–216. https://proceedings.mlsys.org/paper_files/paper/2022/file/38b3eff8baf56627478ec76a704e9b52-Paper.pdf.
  55. 55.Shangdi Yu and Horace He. 2023. Transcending runtime-memory tradeoffs in checkpointing by being fusion aware. In Proceedings of Machine Learning and Systems.
  56. 56.Bojian Zheng et al. 2022. Dietcode: automatic optimization for dynamic tensor programs. In Proceedings of Machine Learning and Systems. D. Marculescu, Y. Chi, and C. Wu, (Eds.) Vol. 4, 848–863. https://proceedings.mlsys.org/paper_files/paper/2022/file/fa7cdfad1a5aaf8370ebeda47a1ff1c3-Paper.pdf.
  57. 57.Lianmin Zheng et al. 2020. Ansor: generating high-performance tensor programs for deep learning. In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation (OSDI’20) Article 49. USENIX Association, USA, 17 pages. isbn: 978-1-939133-19-9.
  58. 58.Size Zheng et al. 2022. AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction. In Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA ’22). Association for Computing Machinery, New York, New York, 874–887. isbn: 9781450386104. doi: 10.1145/3470496.3527440.
  59. 59.Hongyu Zhu et al. 2022. ROLLER: fast and efficient tensor compilation for deep learning. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). USENIX Association, Carlsbad, CA, (July 2022), 233–248. isbn: 978-1-939133-28-1. https://www.usenix.org/conference/osdi22/presentation/zhu.
  60. 60.Mikhail Zolotukhin. 2021. NNC walkthrough: how PyTorch ops get fused. https://dev-discuss.pytorch.org/t/nnc-walkthrough-how-pytorch-ops-get-fused/125. (2021).

Citation

MLA
Ansel, J., et al. “PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation”. Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, 2024, pp. 929–47, https://doi.org/10.1145/3620665.3640366.
APA
Ansel, J., Yang, E., He, H., Gimelshein, N., Jain, A., Voznesensky, M., Bao, B., Bell, P., Berard, D., Burovski, E., Chauhan, G., Chourdia, A., Constable, W., Desmaison, A., DeVito, Z., Ellison, E., Feng, W., Gong, J., Gschwind, M., … Chintala, S. (2024). PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, 929–947. https://doi.org/10.1145/3620665.3640366
Chicago
Ansel, J., E. Yang, H. He, et al. 2024. “PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation”. Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, 929–47. https://doi.org/10.1145/3620665.3640366.
Harvard
Ansel, J. et al. (2024) “PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation”, Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. ACM, pp. 929–947. Available at: https://doi.org/10.1145/3620665.3640366.
Vancouver
1. Ansel J, Yang E, He H, et al (2024) PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. In: Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. ACM, pp 929–947

BibTeX

@inproceedings{Ansel_2024, series={ASPLOS ’24}, title={PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation}, url={http://dx.doi.org/10.1145/3620665.3640366}, DOI={10.1145/3620665.3640366}, booktitle={Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2}, publisher={ACM}, author={Ansel, Jason and Yang, Edward and He, Horace and Gimelshein, Natalia and Jain, Animesh and Voznesensky, Michael and Bao, Bin and Bell, Peter and Berard, David and Burovski, Evgeni and Chauhan, Geeta and Chourdia, Anjali and Constable, Will and Desmaison, Alban and DeVito, Zachary and Ellison, Elias and Feng, Will and Gong, Jiong and Gschwind, Michael and Hirsh, Brian and Huang, Sherlock and Kalambarkar, Kshiteej and Kirsch, Laurent and Lazos, Michael and Lezcano, Mario and Liang, Yanbo and Liang, Jason and Lu, Yinghai and Luk, C. K. and Maher, Bert and Pan, Yunjie and Puhrsch, Christian and Reso, Matthias and Saroufim, Mark and Siraichi, Marcos Yukio and Suk, Helen and Zhang, Shunting and Suo, Michael and Tillet, Phil and Zhao, Xu and Wang, Eikan and Zhou, Keren and Zou, Richard and Wang, Xiaodong and Mathews, Ajit and Wen, William and Chanan, Gregory and Wu, Peng and Chintala, Soumith}, year={2024}, month=Apr, pages={929–947}, collection={ASPLOS ’24} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF