FiT: Flexible Vision Transformer for Diffusion Model

Zeyu LuZidong WangDi HuangChengyue WuXihui LiuWanli OuyangLei Bai

article2024ICML96 citations

Proposes a flexible vision transformer architecture that treats images as variable-length token sequences using 2D rotary positional embeddings, enabling diffusion models to generate high-fidelity images across arbitrary resolutions and aspect ratios without cropping.

Listen

Modern generative artificial intelligence relies heavily on vision transformers within diffusion models to synthesize images. However, conventional models require inputs with fixed, square dimensions and rely on aggressive cropping and resizing during training. This introduces pervasive data biases into generated outputs, leading to cropped subjects, blurring, and severe quality degradation whenever users request non-standard aspect ratios or higher resolutions. Overcoming these rigid constraints is essential for deploying versatile generative media systems.

The article demonstrates and evaluates the Flexible Vision Transformer (FiT), a generative diffusion architecture designed to synthesize high-quality images across unrestricted resolutions and arbitrary aspect ratios without requiring fixed-grid constraints.

To achieve this flexibility, the approach reconceptualizes images not as static pixel grids, but as sequences of dynamically sized visual tokens. During training, native image aspect ratios are preserved by scaling images to a maximum token budget (up to 256 tokens) and using padding tokens with masked self-attention to process variable-length batches safely. Architecturally, the model replaces standard absolute position embeddings with decoupled two-dimensional rotary position embeddings (2D RoPE) and integrates Swish-gated linear units. The authors evaluated this framework on standard ImageNet benchmarks and text-to-image datasets, comparing baseline performance across multiple in-distribution and out-of-distribution aspect ratios against state-of-the-art generative baselines.

The findings demonstrate substantial improvements across various settings. First, when tested on aspect ratios within the training token budget (such as 160x320 and 128x384), the premier model variant (FiT-XL/2) achieved quality scores (FID) of 5.74 and 16.81, outperforming leading prior models such as DiT-XL/2 (which scored 20.14 and 107.2) and U-ViT. Second, on out-of-distribution resolutions that exceed the training token limit (such as 320x320, 224x448, and 160x480), the model maintained strong generation fidelity, achieving top scores across all tested dimensions (5.42, 7.90, and 15.72 FID, respectively). Third, the newly formulated training-free interpolation methods (VisionNTK and VisionYaRN) successfully resolved dimension mismatch issues, reducing error metrics on extreme aspect ratios by more than 40 points compared to standard extrapolation methods. Finally, text-to-image experiments confirmed similar large advantages over baseline transformers on the CC3M dataset.

These results show that generative systems do not need to be locked into rigid square dimensions to maintain high image quality. Eliminating image cropping during training directly prevents visual artifacts and distortion in production. Furthermore, the architecture achieved superior extrapolation performance after only 1.8 million training steps, compared to baseline models trained for up to 7 million steps, representing meaningful compute and operational cost savings. Senior leaders should note, however, that deploying such systems carries common societal risks related to realistic disinformation and intellectual property considerations regarding training data.

Organizations developing generative vision systems should consider adopting dynamic sequence modeling and 2D rotary position embeddings rather than fixed-grid architectures. For immediate deployment, teams can leverage training-free interpolation techniques to extend existing models across varied aspect ratios without incurring expensive retraining costs. Further research and engineering pilots should focus on extending baseline pre-training to larger token budgets (e.g., 1024 tokens), exploring fine-tuning strategies for long sequences, and adapting the framework to video synthesis and image editing.

Confidence in these findings is high for standard image benchmarks within the tested boundary conditions. However, empirical limits remain: generation quality degrades as aspect ratios exceed 1:7 or total resolutions exceed 512x512 under current training bounds. Deployments targeting resolutions or modalities beyond these experimental parameters should proceed with structured validation.

whlzy/FiTLu et al (2024).pdf
Cover for FiT: Flexible Vision Transformer for Diffusion Model

Abstract

Nature is infinitely resolution-free. In the context of this reality, existing diffusion models, such as Diffusion Transformers, often face challenges when processing image resolutions outside of their trained domain. To overcome this limitation, we present the Flexible Vision Transformer (FiT), a transformer architecture specifically designed for generating images with unrestricted resolutions and aspect ratios. Unlike traditional methods that perceive images as static-resolution grids, FiT conceptualizes images as sequences of dynamically-sized tokens. This perspective enables a flexible training strategy that effortlessly adapts to diverse aspect ratios during both training and inference phases, thus promoting resolution generalization and eliminating biases induced by image cropping. Enhanced by a meticulously adjusted network structure and the integration of training-free extrapolation techniques, FiT exhibits remarkable flexibility in resolution extrapolation generation. Comprehensive experiments demonstrate the exceptional performance of FiT across a broad range of resolutions. Repository available at https://github.com/whlzy/FiT.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Flexible Vision Transformer for Diffusion
  • 3.1. Preliminary
  • 3.2. Flexible Training and Inference Pipeline
  • 3.3. Flexible Vision Transformer Architecture
  • 3.4. Training Free Resolution Extrapolation
  • 4. Experiments
  • 4.1. FiT Implementation
  • 4.2. FiT Architecture Design
  • 4.3. FiT Resolution Extrapolation Design
  • 4.4. FiT In-Distribution Resolution Results
  • 4.5. FiT Out-Of-Distribution Resolution Results
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Experimental Setups
  • B. Network Flops Analysis
  • C. Extreme Aspect Ratios and Resolutions Analysis
  • D. Text-to-Image Experiments
  • E. Detailed Attention Score with 2D RoPE and decoupled 2D-RoPE.
  • F. Limitations and Future Work
  • G. More Model Samples

Knowls

  1. Knowl 1 — Flexible Vision Transformer Architecture

    model/method

    The Flexible Vision Transformer (FiT) is a diffusion backbone designed to synthesize images at unrestricted resolutions and aspect ratios by treating images as variable-length token sequences rather than fixed-dimension grids.

    Evolving from the Diffusion Transformer (DiT), each FiT block incorporates three core architectural adjustments:

    1. 2D Rotary Positional Embedding (2D RoPE): Replaces fixed absolute positional embeddings with rotary embeddings evaluated separately across height and width coordinate axes, providing structural support for length and resolution generalization.
    2. Masked Multi-Head Self-Attention (Masked MHSA): Replaces standard MHSA by applying an additive sequence mask MM that prevents real noisy image tokens from attending to padding tokens inserted during batching.
    3. Swish-Gated Linear Unit (SwiGLU): Replaces the conventional Multi-Layer Perceptron (MLP) within the Feed-Forward Network (FFN) blocks to enhance capacity and training stability, defined as:

    SwiGLU(x,W,V)=SiLU(xW)⊗(xV)\text{SwiGLU}(x, W, V) = \text{SiLU}(xW) \otimes (xV)

    FFN(x)=SwiGLU(x,W1,W2)W3\text{FFN}(x) = \text{SwiGLU}(x, W_1, W_2)W_3

    where W1,W2,W3W_1, W_2, W_3 are un-biased parameter weight matrices, ⊗\otimes represents the Hadamard element-wise product, and SiLU(x)=x⋅σ(x)\text{SiLU}(x) = x \cdot \sigma(x).

  2. Knowl 2 — Flexible Training and Inference Pipeline for Diffusion Transformers

    model/method

    To eliminate artifacts and information loss caused by fixed center-cropping and distortion during preprocessing, FiT introduces a flexible training and inference pipeline:

    • Data Preprocessing & Encoding: Training images are resized such that their original aspect ratio is strictly preserved while ensuring total resolution satisfies H×W≤2562H \times W \le 256^2. A pre-trained Variational Autoencoder (VAE) downsamples the image by a factor of 8 into latent feature maps of dimension 4×H8×W84 \times \frac{H}{8} \times \frac{W}{8}. The latent codes are divided into non-overlapping patches of size p=2p=2, resulting in a 1D sequence of L=H16×W16L = \frac{H}{16} \times \frac{W}{16} latent tokens where L≤Lmax=256L \le L_{\text{max}} = 256.
    • Batch Packing and Padding: Sequences of variable token length LL within a batch are right-padded to the fixed maximum token capacity Lmax=256L_{\text{max}} = 256 using padding tokens. Positional indices are padded accordingly.
    • Loss Formulation: The diffusion denoising loss is computed exclusively on the valid LL token predictions; all padded positions are masked out from loss evaluation and gradient updates.
    • Inference Pipeline: Given an arbitrary target resolution (Htest,Wtest)(H_{\text{test}}, W_{\text{test}}), initial Gaussian noise tokens are mapped to their 2D coordinate positions, padded to LmaxL_{\text{max}}, denoised iteratively over KK diffusion sampling steps through the Masked MHSA transformer blocks, un-padded, reshaped to the grid coordinates, and decoded by the VAE decoder.
  3. Knowl 3 — 2D Rotary Positional Embedding for Vision Transformers

    equation

    To inject relative spatial information across arbitrary 2D coordinate grids, 2D Rotary Positional Embedding (2D RoPE) decomposes the feature dimension ∣D∣|D| of query qm∈R∣D∣q_m \in \mathbb{R}^{|D|} and key kn∈R∣D∣k_n \in \mathbb{R}^{|D|} into orthogonal subspaces of dimension ∣D∣2\frac{|D|}{2} corresponding to vertical coordinate h∈{1,…,H}h \in \{1, \dots, H\} and horizontal coordinate w∈{1,…,W}w \in \{1, \dots, W\}:

    fq(qm,hm,wm)=[eihmΘqm(h)  ∥  eiwmΘqm(w)]f_q(q_m, h_m, w_m) = \left[ e^{i h_m \Theta} q_m^{(h)} \;\parallel\; e^{i w_m \Theta} q_m^{(w)} \right]

    fk(kn,hn,wn)=[eihnΘkn(h)  ∥  eiwnΘkn(w)]f_k(k_n, h_n, w_n) = \left[ e^{i h_n \Theta} k_n^{(h)} \;\parallel\; e^{i w_n \Theta} k_n^{(w)} \right]

    where ∥\parallel denotes vector concatenation along the channel dimension, qm(h),qm(w)∈R∣D∣/2q_m^{(h)}, q_m^{(w)} \in \mathbb{R}^{|D|/2} represent the partitioned query vector, and Θ=Diag(θ1,…,θ∣D∣/4)\Theta = \text{Diag}(\theta_1, \dots, \theta_{|D|/4}) is the diagonal rotary frequency matrix with entries θd=b−2d/∣D∣\theta_d = b^{-2d/|D|} (base b=10000b = 10000).

    The resulting attention score between tokens mm and nn decomposes additively without height-width cross-terms:

    Am,n=Re⟨fq(qm,hm,wm),fk(kn,hn,wn)⟩A_{m,n} = \text{Re}\langle f_q(q_m, h_m, w_m), f_k(k_n, h_n, w_n) \rangle

    =∑j=0∣D∣/4−1[(q2j(h)k2j(h)+q2j+1(h)k2j+1(h))cos⁡((hm−hn)θj)+(q2j(h)k2j+1(h)−q2j+1(h)k2j(h))sin⁡((hm−hn)θj)]= \sum_{j=0}^{|D|/4 - 1} \left[ (q_{2j}^{(h)} k_{2j}^{(h)} + q_{2j+1}^{(h)} k_{2j+1}^{(h)}) \cos((h_m - h_n)\theta_j) + (q_{2j}^{(h)} k_{2j+1}^{(h)} - q_{2j+1}^{(h)} k_{2j}^{(h)}) \sin((h_m - h_n)\theta_j) \right]

    +∑j=0∣D∣/4−1[(q2j(w)k2j(w)+q2j+1(w)k2j+1(w))cos⁡((wm−wn)θj)+(q2j(w)k2j+1(w)−q2j+1(w)k2j(w))sin⁡((wm−wn)θj)]+ \sum_{j=0}^{|D|/4 - 1} \left[ (q_{2j}^{(w)} k_{2j}^{(w)} + q_{2j+1}^{(w)} k_{2j+1}^{(w)}) \cos((w_m - w_n)\theta_j) + (q_{2j}^{(w)} k_{2j+1}^{(w)} - q_{2j+1}^{(w)} k_{2j}^{(w)}) \sin((w_m - w_n)\theta_j) \right]

  4. Knowl 4 — VisionNTK Positional Embedding Interpolation

    definition

    VisionNTK is a training-free resolution extrapolation method for decoupled 2D RoPE diffusion models generating images at inference resolutions (Htest,Wtest)(H_{\text{test}}, W_{\text{test}}) that exceed the training maximum context Ltrain=LmaxL_{\text{train}} = \sqrt{L_{\text{max}}}.

    Independent scale factors for height and width are defined as:

    sh=max⁡(HtestLtrain,1.0),sw=max⁡(WtestLtrain,1.0)s_h = \max\left(\frac{H_{\text{test}}}{L_{\text{train}}}, 1.0\right), \quad s_w = \max\left(\frac{W_{\text{test}}}{L_{\text{train}}}, 1.0\right)

    VisionNTK scales the base frequency b=10000b = 10000 along height and width dimensions separately:

    bh=b⋅sh∣D∣∣D∣−2,bw=b⋅sw∣D∣∣D∣−2b_h = b \cdot s_h^{\frac{|D|}{|D|-2}}, \quad b_w = b \cdot s_w^{\frac{|D|}{|D|-2}}

    yielding coordinate-specific frequency sets Θh={θdh=bh−2d/∣D∣,1≤d≤∣D∣/2}\Theta_h = \{\theta_d^h = b_h^{-2d/|D|}, 1 \le d \le |D|/2\} and Θw={θdw=bw−2d/∣D∣,1≤d≤∣D∣/2}\Theta_w = \{\theta_d^w = b_w^{-2d/|D|}, 1 \le d \le |D|/2\}. When the aspect ratio is 1:11:1 (sh=sws_h = s_w), VisionNTK reduces to standard LLM NTK-aware interpolation.

  5. Knowl 5 — VisionYaRN Positional Embedding Interpolation

    definition

    VisionYaRN is a training-free positional interpolation method extending decoupled 2D RoPE for out-of-distribution vision generation. For a given training scale Ltrain=LmaxL_{\text{train}} = \sqrt{L_{\text{max}}}, height/width scale factors sh=max⁡(Htest/Ltrain,1.0)s_h = \max(H_{\text{test}}/L_{\text{train}}, 1.0) and sw=max⁡(Wtest/Ltrain,1.0)s_w = \max(W_{\text{test}}/L_{\text{train}}, 1.0), and dimension-dependent ratio r(d)=Ltrain2πb2d/∣D∣r(d) = \frac{L_{\text{train}}}{2\pi b^{2d/|D|}}, the modified rotary frequencies are:

    θdh=(1−γ(r(d)))θdsh+γ(r(d))θd\theta_d^h = (1 - \gamma(r(d)))\frac{\theta_d}{s_h} + \gamma(r(d))\theta_d

    θdw=(1−γ(r(d)))θdsw+γ(r(d))θd\theta_d^w = (1 - \gamma(r(d)))\frac{\theta_d}{s_w} + \gamma(r(d))\theta_d

    where γ(r)\gamma(r) is a ramp function parameterised by threshold hyperparameters α,β\alpha, \beta (set to α=1,β=32\alpha=1, \beta=32):

    γ(r)={0if r<α1if r>βr−αβ−αotherwise\gamma(r) = \begin{cases} 0 & \text{if } r < \alpha \\ 1 & \text{if } r > \beta \\ \frac{r-\alpha}{\beta-\alpha} & \text{otherwise} \end{cases}

    VisionYaRN also scales the query and key representations by 1t=0.1ln⁡(s)+1\frac{1}{\sqrt{t}} = 0.1 \ln(s) + 1.

  6. Knowl 6 — Masked Multi-Head Self-Attention Formulation

    equation

    To prevent information leakage between valid image tokens and padding tokens within variable-length batches, FiT replaces standard MHSA with Masked Multi-Head Self-Attention. For head ii with queries QiQ_i, keys KiK_i, values ViV_i, and key channel dimension dkd_k, the attention map is computed as:

    Masked Attn(Qi,Ki,Vi)=Softmax(QiKiTdk+M)Vi\text{Masked Attn}(Q_i, K_i, V_i) = \text{Softmax}\left( \frac{Q_i K_i^T}{\sqrt{d_k}} + M \right) V_i

    where M∈RLmax×LmaxM \in \mathbb{R}^{L_{\text{max}} \times L_{\text{max}}} is the attention mask defined such that Mj,k=0M_{j, k} = 0 if both token jj and token kk are valid latent tokens, and Mj,k=−∞M_{j, k} = -\infty if token kk is a padding token.

  7. Knowl 7 — Class-Conditional Image Generation Benchmark on In-Distribution Resolutions

    data/table

    FiT-XL/2 (trained for 1.8M steps with batch size 256 and patch size p=2p=2, combined with VisionNTK extrapolation) was evaluated on ImageNet at three in-distribution resolutions (L≤256L \le 256 tokens) against state-of-the-art generative models with classifier-free guidance:

    Method Train Cost 256×\times256 (1:1) 160×\times320 (1:2) 128×\times384 (1:3)
    FID↓\downarrow sFID↓\downarrow IS↑\uparrow FID↓\downarrow sFID↓\downarrow IS↑\uparrow FID↓\downarrow sFID↓\downarrow IS↑\uparrow
    BigGAN-deep - 6.95 7.36 171.4 - - - - - -
    StyleGAN-XL - 2.30 4.02 265.12 - - - - - -
    MaskGIT 1387k×\times256 6.18 - 182.1 - - - - - -
    CDM - 4.88 - 158.71 - - - - - -
    U-ViT-H/2-G 500k×\times1024 2.35 5.68 265.02 6.93 12.64 175.08 196.84 95.90 7.54
    ADM-G,U 1980k×\times256 3.94 6.14 215.84 10.26 12.28 126.99 56.52 43.21 32.19
    LDM-4-G 178k×\times1200 3.60 5.12 247.67 10.04 11.47 119.56 29.67 26.33 57.71
    MDT-G 6500k×\times256 1.79 4.57 283.01 135.6 73.08 9.35 124.9 70.69 13.38
    DiT-XL/2-G 7000k×\times256 2.27 4.60 278.24 20.14 30.50 97.28 107.2 68.89 15.48
    FiT-XL/2-G 1800k×\times256 4.27 9.99 249.72 5.74 10.05 190.14 16.81 20.62 110.93

    While models trained exclusively on fixed square grids degrade severely when generating non-square aspect ratios (e.g., DiT-XL/2 FID jumps to 20.14 at 1:2 and 107.2 at 1:3), FiT-XL/2 maintains high fidelity, achieving state-of-the-art FID of 5.74 at 160×320160 \times 320 and 16.81 at 128×384128 \times 384.

  8. Knowl 8 — Out-of-Distribution Resolution Extrapolation Benchmark

    data/table

    FiT-XL/2 was benchmarked on ImageNet at out-of-distribution resolutions exceeding the training token limit (L>256L > 256 tokens) using classifier-free guidance:

    Method Train Cost 320×\times320 (1:1, 400 tok) 224×\times448 (1:2, 392 tok) 160×\times480 (1:3, 300 tok)
    FID↓\downarrow sFID↓\downarrow IS↑\uparrow FID↓\downarrow sFID↓\downarrow IS↑\uparrow FID↓\downarrow sFID↓\downarrow IS↑\uparrow
    U-ViT-H/2-G 500k×\times1024 7.65 16.30 208.01 67.10 42.92 45.54 95.56 44.45 24.01
    ADM-G,U 1980k×\times256 9.39 9.01 161.95 11.34 14.50 146.00 23.92 25.55 80.73
    LDM-4-G 178k×\times1200 6.24 13.21 220.03 8.55 17.62 186.25 19.24 20.25 99.34
    MDT-G 6500k×\times256 383.5 136.5 4.24 365.9 142.8 4.91 276.7 138.1 7.20
    DiT-XL/2-G 7000k×\times256 9.98 23.57 225.72 94.94 56.06 35.75 140.2 79.60 14.70
    FiT-XL/2-G 1800k×\times256 5.42 15.41 252.65 7.90 19.63 215.29 15.72 22.57 132.76

    FiT-XL/2 with VisionNTK achieves the best FID and Inception Score across all three out-of-distribution resolutions, outperforming both prior transformer models (DiT, MDT, U-ViT) and CNN-based models (ADM, LDM-4).

  9. Knowl 9 — Ablation of FiT Architectural Components

    data/table

    Ablation experiments evaluated the incremental impact of flexible training, SwiGLU, and 2D RoPE on base models (FiT-B/2) trained for 400K steps on ImageNet without classifier-free guidance:

    Model Variant Pos. Embed. FFN Training 256×\times256 FID↓\downarrow 160×\times320 FID↓\downarrow 224×\times448 FID↓\downarrow
    DiT-B Abs. PE MLP Fixed 44.83 91.32 109.10
    Config A Abs. PE MLP Flexible 43.34 50.51 52.55
    Config B Abs. PE SwiGLU Flexible 41.75 48.66 52.34
    Config C Abs. PE + 2D RoPE MLP Flexible 39.11 46.71 46.60
    Config D 2D RoPE MLP Flexible 37.29 45.06 46.16
    FiT-B 2D RoPE SwiGLU Flexible 36.36 43.96 44.67
    • Flexible Training: Reduces FID from 91.32 to 50.51 on 160×320160 \times 320 and from 109.10 to 52.55 on 224×448224 \times 448.
    • 2D RoPE vs. Abs. PE: Replacing absolute positional embeddings with 2D RoPE (Config D vs. Config A) decreases FID by 6.05 at 256×256256 \times 256, 5.45 at 160×320160 \times 320, and 6.39 at 224×448224 \times 448. Combining absolute PE with 2D RoPE (Config C) underperforms pure 2D RoPE (Config D).
    • SwiGLU vs. MLP: SwiGLU consistently improves FID across all tested resolutions (Config B vs. Config A and FiT-B vs. Config D).
  10. Knowl 10 — Ablation of Positional Interpolation Methods for Resolution Extrapolation

    data/table

    Comparison of positional interpolation methods on DiT-B/2 and FiT-B/2 at 400K training steps without classifier-free guidance on out-of-distribution resolutions:

    Method 320×\times320 (1:1) 224×\times448 (1:2) 160×\times480 (1:3)
    FID↓\downarrow IS↑\uparrow FID↓\downarrow IS↑\uparrow FID↓\downarrow IS↑\uparrow
    DiT-B 95.47 18.38 109.10 14.00 143.80 8.93
    DiT-B + EI 81.48 20.97 133.20 11.11 160.40 7.30
    DiT-B + PI 72.47 24.15 133.40 11.73 156.50 7.80
    FiT-B (Direct) 61.35 31.01 44.67 37.10 56.81 25.25
    FiT-B + PI 65.76 29.32 175.42 8.45 224.83 5.89
    FiT-B + YaRN 44.76 44.70 82.19 29.68 104.06 20.76
    FiT-B + NTK 57.31 33.97 45.24 38.84 59.19 26.01
    FiT-B + VisionYaRN 44.76 44.70 41.92 45.87 62.84 27.84
    FiT-B + VisionNTK 57.31 33.97 43.84 39.22 56.76 26.40

    Vanilla YaRN and Position Interpolation (PI) degrade significantly when the aspect ratio deviates from 1:11:1. Decoupled formulations (VisionYaRN and VisionNTK) rectify this aspect-ratio imbalance, outperforming direct extrapolation across all three out-of-distribution resolutions.

  11. Knowl 11 — Compute and Capacity Scaling Analysis of FiT

    empirical result

    FiT performance scales predictably with model parameter capacity and training compute (FLOPs):

    • At a fixed training budget of 400K steps, scaling model capacity monotonically decreases ImageNet-256 FID: FiT-B/2 (130M parameters, 29.065×10929.065\times 10^9 inference FLOPs, 8.93×10158.93\times 10^{15} training FLOPs) achieves 36.36 FID; FiT-L/2 (0.103×10120.103\times 10^{12} inference FLOPs, 3.16×10163.16\times 10^{16} training FLOPs) achieves 22.32 FID; and FiT-XL/2 (675M parameters, 0.153×10120.153\times 10^{12} inference FLOPs, 4.70×10164.70\times 10^{16} training FLOPs) achieves 20.11 FID.
    • For FiT-XL/2, scaling training compute from 400K steps (4.70×10164.70\times 10^{16} FLOPs) to 2500K steps (2.94×10172.94\times 10^{17} FLOPs) reduces FID from 20.11 to 10.30 (without guidance).
  12. Knowl 12 — Extrapolation Boundaries and Stated Limitations of FiT

    limitation

    FiT exhibits specific operating constraints and empirical failure boundaries:

    1. Extrapolation Limits: When tested on resolutions and aspect ratios far beyond the training domain (H×W≤2562H \times W \le 256^2), FID degrades substantially from 4.27 at 256×256256 \times 256 to 13.37 at 384×384384 \times 384, 47.46 at 448×448448 \times 448, and 94.58 at 512×512512 \times 512. For extreme aspect ratios, FID degrades to 35.30 at 1:4 (128×512128 \times 512), 69.89 at 1:5 (120×600120 \times 600), and 113.95 at 1:6 (120×720120 \times 720), placing practical generation limits around 512×512512 \times 512 resolution and 1:71:7 aspect ratio.
    2. Training Context Bound: Training and evaluation were capped at maximum sequence length Lmax=256L_{\text{max}} = 256 tokens.
    3. Training-Free Restriction: Only training-free extrapolation techniques were explored; supervised long-context fine-tuning methods were not evaluated.
    4. Modality Constraint: The architecture was verified exclusively on 2D image synthesis, leaving video and higher-dimensional generative extensions to future work.

Coverage note — Text-to-image experiments on CC3M (Appendix D) were omitted as they serve as a minor validation of the core class-conditional ImageNet contributions.

References

  1. 1.Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., and Zhu, J. All are worth words: A vit backbone for diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
  2. 2.Brock, A., Donahue, J., and Simonyan, K. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018.
  3. 3.Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A. Video generation models as world simulators. 2024. Accessed: 2024-5-1.
  4. 4.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 2020.
  5. 5.Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T. Maskgit: Masked generative image transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  6. 6.Chen, S., Wong, S., Chen, L., and Tian, Y. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595, 2023.
  7. 7.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 2023a.
  8. 8.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., and et al, P. B. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 2023b.
  9. 9.Dehghani, M., Mustafa, B., Djolonga, J., Heek, J., Minderer, M., Caron, M., Steiner, A., Puigcerver, J., Geirhos, R., Alabdulmohsin, I., et al. Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution. arXiv preprint arXiv:2307.06304, 2023.
  10. 10.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2009.
  11. 11.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 2021.
  12. 12.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  13. 13.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  14. 14.Gao, S., Zhou, P., Cheng, M.-M., and Yan, S. Masked diffusion transformer is a strong image synthesizer. arXiv preprint arXiv:2303.14389, 2023.
  15. 15.Gong, J., Bai, L., Ye, P., Xu, W., Liu, N., Dai, J., Yang, X., and Ouyang, W. Cascast: Skillful high-resolution precipitation nowcasting via cascaded modelling. arXiv preprint arXiv:2402.04290, 2024.
  16. 16.Hatamizadeh, A., Song, J., Liu, G., Kautz, J., and Vahdat, A. Diffit: Diffusion vision transformers for image generation. arXiv preprint arXiv:2312.02139, 2023.
  17. 17.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in Neural Information Processing Systems, 2017.
  18. 18.Ho, J. and Salimans, T. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021.
  19. 19.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 2020.
  20. 20.Ho, J., Saharia, C., Chan, W., Fleet, D. J., Norouzi, M., and Salimans, T. Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research, 2022.
  21. 21.Hyv¨arinen, A. and Dayan, P. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research, 2005.
  22. 22.Kynk¨aanniemi, T., Karras, T., Laine, S., and Lehtinen, J.and Aila, T. Improved precision and recall metric for assessing generative models. Advances in Neural Information Processing Systems, 2019.
  23. 23.Li, C., Huang, D., Lu, Z., Xiao, Y., Pei, Q., and Bai, L. A survey on long video generation: Challenges, methods, and prospects. arXiv preprint arXiv:2403.16407, 2024.
  24. 24.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll´ar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014.
  25. 25.Ling, F., Lu, Z., Luo, J.-J., Bai, L., Behera, S. K., Jin, D., Pan, B., Jiang, H., and Yamagata, T. Diffusion model-based probabilistic downscaling for 180-year east asian climate reconstruction. npj Climate and Atmospheric Science, 2024.
  26. 26.Liu, X., Yan, H., Zhang, S., An, C., Qiu, X., and Lin, D. Scaling laws of rope-based extrapolation. arXiv preprint arXiv:2310.05209, 2023.
  27. 27.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  28. 28.Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  29. 29.LocalLLaMA. Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation. https://www.reddit.com/r/LocalLLaMA/comments/14lz7j5/ntkaware_scaled_rope_allows_llama_models_to_have/. Accessed: 2024-2-1.
  30. 30.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  31. 31.Lu, Z., Jiang, J., Huang, J., Wu, G., and Liu, X. Glama: Joint spatial and frequency loss for general image inpainting. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  32. 32.Lu, Z., Huang, D., Bai, L., Qu, J., Wu, C., Liu, X., and Ouyang, W. Seeing is not always believing: Benchmarking human and model perception of ai-generated images. Advances in Neural Information Processing Systems, 2024a.
  33. 33.Lu, Z., Wu, C., Chen, X., Wang, Y., Bai, L., Qiao, Y., and Liu, X. Hierarchical diffusion autoencoders and disentangled image manipulation. In IEEE/CVF Winter Conference on Applications of Computer Vision, 2024b.
  34. 34.Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S. Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073, 2021.
  35. 35.Nash, C., Menick, J., Dieleman, S., and Battaglia, P. W. Generating images with sparse representations. arXiv preprint arXiv:2103.03841, 2021.
  36. 36.Peebles, W. and Xie, S. Scalable diffusion models with transformers. In IEEE/CVF International Conference on Computer Vision, 2023.
  37. 37.Peng, B., Quesnelle, J., Fan, H., and Shippole, E. Yarn: Efficient context window extension of large language models. arXiv preprint arXiv:2309.00071, 2023.
  38. 38.Poole, B., Jain, A., Barron, J. T., and Mildenhall, B. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022.
  39. 39.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 2021.
  40. 40.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  41. 41.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  42. 42.Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
  43. 43.Ruoss, A., Del´etang, G., Genewein, T., Grau-Moya, J., Csord´as, R., Bennani, M., Legg, S., and Veness, J. Randomized positional encodings boost length generalization of transformers. arXiv preprint arXiv:2305.16843, 2023.
  44. 44.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 2022.
  45. 45.Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. Improved techniques for training gans. Advances in Neural Information Processing Systems, 2016.
  46. 46.Sauer, A., Schwarz, K., and Geiger, A. Stylegan-xl: Scaling stylegan to large diverse datasets. In ACM SIGGRAPH 2022 conference proceedings, 2022.
  47. 47.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Association for Computational Linguistics, 2018.
  48. 48.Shazeer, N. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  49. 49.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020a.
  50. 50.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020b.
  51. 51.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 2024.
  52. 52.Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., , and Wei, F. A length-extrapolatable transformer. arXiv preprint arXiv:2212.10554, 2022.
  53. 53.Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  54. 54.Touvron, H., Vedaldi, A., Douze, M., and J´egou, H. Fixing the train-test resolution discrepancy. Advances in Neural Information Processing Systems, 2019.
  55. 55.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J´egou, H. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, 2021.
  56. 56.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., and et al, B. R. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  57. 57.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., and et al, N. B. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  58. 58.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in Neural Information Processing Systems, 2017.
  59. 59.Wang, X., Chen, G., Qian, G., Gao, P., Wei, X.-Y., Wang, Y., Tian, Y., and Gao, W. Large-scale multi-modal pre-trained models: A comprehensive survey. Machine Intelligence Research, 2023.
  60. 60.Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., et al. Effective long-context scaling of foundation models. arXiv preprint arXiv:2309.16039, 2023.
  61. 61.Yu, J., Lin, Z., Yang, J., Shen, X., Lu, X., and Huang, T. S. Generative image inpainting with contextual attention. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.

Citation

MLA
Lu, Z., et al. “FiT: Flexible Vision Transformer for Diffusion Model”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.12376.
APA
Lu, Z., Wang, Z., Huang, D., Wu, C., Liu, X., Ouyang, W., & Bai, L. (2024). FiT: Flexible Vision Transformer for Diffusion Model. arXiv. https://doi.org/10.48550/arxiv.2402.12376
Chicago
Lu, Z., Z. Wang, D. Huang, et al. 2024. “FiT: Flexible Vision Transformer for Diffusion Model”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.12376.
Harvard
Lu, Z. et al. (2024) “FiT: Flexible Vision Transformer for Diffusion Model”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.12376.
Vancouver
1. Lu Z, Wang Z, Huang D, Wu C, Liu X, Ouyang W, Bai L (2024) FiT: Flexible Vision Transformer for Diffusion Model. https://doi.org/10.48550/arxiv.2402.12376

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.12376,
  doi = {10.48550/ARXIV.2402.12376},
  url = {https://arxiv.org/abs/2402.12376},
  author = {Lu, Zeyu and Wang, Zidong and Huang, Di and Wu, Chengyue and Liu, Xihui and Ouyang, Wanli and Bai, Lei},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {FiT: Flexible Vision Transformer for Diffusion Model},
  publisher = {arXiv},
  year = {2024},
  copyright = {Creative Commons Attribution Non Commercial Share Alike 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/