FiT: Flexible Vision Transformer for Diffusion Model
Zeyu LuZidong WangDi HuangChengyue WuXihui LiuWanli OuyangLei Bai
Proposes a flexible vision transformer architecture that treats images as variable-length token sequences using 2D rotary positional embeddings, enabling diffusion models to generate high-fidelity images across arbitrary resolutions and aspect ratios without cropping.
Modern generative artificial intelligence relies heavily on vision transformers within diffusion models to synthesize images. However, conventional models require inputs with fixed, square dimensions and rely on aggressive cropping and resizing during training. This introduces pervasive data biases into generated outputs, leading to cropped subjects, blurring, and severe quality degradation whenever users request non-standard aspect ratios or higher resolutions. Overcoming these rigid constraints is essential for deploying versatile generative media systems.
The article demonstrates and evaluates the Flexible Vision Transformer (FiT), a generative diffusion architecture designed to synthesize high-quality images across unrestricted resolutions and arbitrary aspect ratios without requiring fixed-grid constraints.
To achieve this flexibility, the approach reconceptualizes images not as static pixel grids, but as sequences of dynamically sized visual tokens. During training, native image aspect ratios are preserved by scaling images to a maximum token budget (up to 256 tokens) and using padding tokens with masked self-attention to process variable-length batches safely. Architecturally, the model replaces standard absolute position embeddings with decoupled two-dimensional rotary position embeddings (2D RoPE) and integrates Swish-gated linear units. The authors evaluated this framework on standard ImageNet benchmarks and text-to-image datasets, comparing baseline performance across multiple in-distribution and out-of-distribution aspect ratios against state-of-the-art generative baselines.
The findings demonstrate substantial improvements across various settings. First, when tested on aspect ratios within the training token budget (such as 160x320 and 128x384), the premier model variant (FiT-XL/2) achieved quality scores (FID) of 5.74 and 16.81, outperforming leading prior models such as DiT-XL/2 (which scored 20.14 and 107.2) and U-ViT. Second, on out-of-distribution resolutions that exceed the training token limit (such as 320x320, 224x448, and 160x480), the model maintained strong generation fidelity, achieving top scores across all tested dimensions (5.42, 7.90, and 15.72 FID, respectively). Third, the newly formulated training-free interpolation methods (VisionNTK and VisionYaRN) successfully resolved dimension mismatch issues, reducing error metrics on extreme aspect ratios by more than 40 points compared to standard extrapolation methods. Finally, text-to-image experiments confirmed similar large advantages over baseline transformers on the CC3M dataset.
These results show that generative systems do not need to be locked into rigid square dimensions to maintain high image quality. Eliminating image cropping during training directly prevents visual artifacts and distortion in production. Furthermore, the architecture achieved superior extrapolation performance after only 1.8 million training steps, compared to baseline models trained for up to 7 million steps, representing meaningful compute and operational cost savings. Senior leaders should note, however, that deploying such systems carries common societal risks related to realistic disinformation and intellectual property considerations regarding training data.
Organizations developing generative vision systems should consider adopting dynamic sequence modeling and 2D rotary position embeddings rather than fixed-grid architectures. For immediate deployment, teams can leverage training-free interpolation techniques to extend existing models across varied aspect ratios without incurring expensive retraining costs. Further research and engineering pilots should focus on extending baseline pre-training to larger token budgets (e.g., 1024 tokens), exploring fine-tuning strategies for long sequences, and adapting the framework to video synthesis and image editing.
Confidence in these findings is high for standard image benchmarks within the tested boundary conditions. However, empirical limits remain: generation quality degrades as aspect ratios exceed 1:7 or total resolutions exceed 512x512 under current training bounds. Deployments targeting resolutions or modalities beyond these experimental parameters should proceed with structured validation.
- Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). It introduces the Diffusion Transformer (DiT) baseline architecture that FiT directly builds upon and adapts to handle flexible resolutions and arbitrary aspect ratios.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). It establishes the foundational Vision Transformer (ViT) paradigm of tokenizing 2D image patches into sequence representations processed by self-attention.
- Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). It pioneered two-stage high-resolution generative modeling combining discrete visual representations with transformer-based sequence architectures.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). It scales multimodal flow-based transformers (MM-DiT) for high-resolution image generation, advancing the family of transformer architectures for continuous generative modeling.
- Paper: Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think, Sihyun Yu et al. (2025). It introduces representation alignment techniques to accelerate the training and enhance the visual fidelity of diffusion transformers like DiT and its flexible variants.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). It expands dynamic resolution processing and multidimensional positional embeddings to large vision-language transformer models for unrestricted input handling.
- Paper: DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation, Makoto Shing et al. (2026). It presents a memory-efficient block-wise training formulation that applies diffusion model principles directly across transformer architectures including diffusion transformers.
