Built independently by an author, for readers. Read the story and support ChapterPal

keyword

rotary positional embedding

Rotary positional embedding is a method used in transformer neural networks to encode the position of tokens within a sequence by rotating their feature representations in multi-dimensional space. Unlike traditional positional encodings that add fixed or learned vectors directly to token embeddings, rotary positional embedding applies a rotation matrix to the query and key representations within the attention mechanism, scaling the rotation angle proportionally to each token position. This mathematical formulation naturally incorporates relative distance directly into the attention computation, allowing the model to focus on the spatial or sequential gaps between elements. Consequently, rotary positional embedding offers strong relative position awareness, enables smooth decay of attention weights over distance, and supports effective generalization to sequence lengths or spatial dimensions beyond those observed during training.

1 item

FiT: Flexible Vision Transformer for Diffusion Model

FiT: Flexible Vision Transformer for Diffusion Model

Zeyu Lu, Zidong Wang, Di Huang, Chengyue Wu, Xihui Liu, Wanli Ouyang, Lei Bai

OrganizationsShanghai Artificial Intelligence LaboratoryShanghai Jiao Tong UniversityTsinghua UniversityUniversity of Hong KongUniversity of Sydney

Why you should read this

Proposes a flexible vision transformer architecture that treats images as variable-length token sequences using 2D rotary positional embeddings, enabling diffusion models to generate high-fidelity images across arbitrary resolutions and aspect ratios without cropping.

Nature is infinitely resolution-free. In the context of this reality, existing diffusion models, such as Diffusion Transformers, often face challenges when processing image resolutions outside of their trained domain. To overcome this limitation, we present the Flexible Vision Transformer (FiT), a transformer architecture specifically designed for generating images with unrestricted resolutions and aspect ratios. Unlike traditional methods that perceive images as static-resolution grids, FiT conceptualizes images as sequences of dynamically-sized tokens. This perspective enables a flexible training strategy that effortlessly adapts to diverse aspect ratios during both training and inference phases, thus promoting resolution generalization and eliminating biases induced by image cropping. Enhanced by a meticulously adjusted network structure and the integration of training-free extrapolation techniques, FiT exhibits remarkable flexibility in resolution extrapolation generation. Comprehensive experiments demonstrate the exceptional performance of FiT across a broad range of resolutions. Repository available at https://github.com/whlzy/FiT.

Added

2026-09-26