FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
Roman BachmannJesse AllardiceDavid MizrahiEnrico FiniOguzhan Fatih KarElmira AmirlooAlaaeldin El-NoubyAmir ZamirAfshin Dehghan
Introduces an image tokenizer that encodes 2D visual content into ordered, variable-length 1D token sequences paired with a rectified flow decoder, enabling autoregressive models to generate high-quality images from coarse to fine details using as few as eight tokens.
Modern generative vision systems rely on image tokenization to compress high-dimensional pixels into discrete representations that models can predict. Standard approaches encode images into rigid two-dimensional grids where the sequence length is determined strictly by the image dimensions rather than visual complexity. Consequently, models must always predict a full, fixed budget of tokens regardless of whether the target image is a basic object or an intricate scene, creating computational inefficiencies and rigid generation pipelines.
The article introduces and evaluates FlexTok, a unified framework that resamples images into ordered, variable-length one-dimensional token sequences. FlexTok allows an image to be represented and reconstructed using anywhere from 1 to 256 discrete tokens within a single model architecture.
To achieve this, the authors combine a Vision Transformer encoder utilizing learnable register tokens with finite scalar quantization, causal attention masking, and nested dropout to enforce a coarse-to-fine token ordering. The tokens condition an end-to-end rectified flow decoder that reconstructs the image from intermediate latents, enhanced with an inductive bias feature loss to accelerate training. The approach was systematically evaluated across standard reconstruction benchmarks and generative tasks, including class-conditioned generation on ImageNet-1k and text-to-image synthesis on large-scale captioned datasets using autoregressive models scaled up to 3 billion parameters.
The evaluation yielded several key findings. First, FlexTok produces visually plausible reconstructions across extreme compression rates, achieving a reconstruction Fréchet Inception Distance below 2.0 with as few as 8 tokens and reaching 1.08 with 256 tokens. Second, high-level semantic and geometric concepts naturally emerge in the earliest tokens, while fine details are added in later tokens. Third, in class-conditioned generation, the top-performing FlexTok setup paired with an autoregressive model matched or outperformed existing baselines such as TiTok and 2D grid models, achieving a generation Fréchet Inception Distance of 1.86 at 32 tokens while requiring roughly one-eighth the sequence length of traditional 256-token grids. Finally, downstream token demand scales with prompt complexity: simple tasks require only 4 to 16 tokens to fulfill conditioning, whereas complex open-ended text prompts benefit from using up to 256 tokens.
These findings demonstrate that generative models do not require uniform compute budgets for every image. By structuring visual information into a coarse-to-fine visual vocabulary, systems can terminate generation early on simple tasks, potentially yielding significant inference speedups and lower computational costs. Furthermore, the class-conditional autoregressive models achieve optimal performance without requiring classifier-free guidance, eliminating the traditional twofold compute overhead associated with generating unconditioned passes.
Practitioners should explore early-stopping mechanisms during inference to dynamically adjust generation length based on prompt complexity. Organizations deploying autoregressive generative vision models should evaluate one-dimensional flexible-length tokenization as an alternative to fixed two-dimensional grids to lower computational costs. When high fidelity across diverse sequence lengths is required, practitioners should adopt rectified flow decoding alongside representation alignment losses.
However, some limitations remain. While the flow-based decoder maintains high perceptual quality with few tokens, semantic variance across different random seeds is higher when using very low token counts. In addition, the adaptive conditioning components in the decoder currently add substantial parameter overhead. The findings provide high confidence for standard single-image generation benchmarks at 256x256 resolution, but further research is required to evaluate scaling behavior on higher-resolution imagery, dense document rendering, and other domains such as video.
- Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). Read this foundational VQGAN work to understand how discrete visual codes enable autoregressive image generation, the fixed-grid tokenization FlexTok replaces.
- Paper: Generating Diverse High-Fidelity Images with VQ-VAE-2, Ali Razavi et al. (2019). Its hierarchical VQ-VAE approach establishes the discrete image-token representations that motivate FlexTok’s more flexible, ordered compression.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This study develops the rectified-flow generation framework that FlexTok adapts for decoding images from variable-length token sequences.
- Paper: FiT: Flexible Vision Transformer for Diffusion Model, Zeyu Lu et al. (2024). FiT introduces flexible-length visual token sequences for generation, providing a direct architectural precedent for FlexTok’s variable-budget representation.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). GlobalCom2 extends adaptive visual-token budgeting to high-resolution vision-language inputs by allocating compression according to crop and token importance.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). FiCoCo carries flexible token reduction into multimodal inference, combining redundancy filtering and information transfer to cut visual context costs.
- Paper: ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation, Tianchen Zhao et al. (2025). ViDiT-Q continues the efficiency agenda for visual generation by compressing diffusion-transformer computation through token- and channel-aware quantization.
- Paper: Accelerating Diffusion Transformers with Token-wise Feature Caching, Chang Zou et al. (2025). This work extends generation-time efficiency beyond token budgets by caching token-wise features across diffusion steps.
- Paper: Variational Flow Maps: Make Some Noise for One-Step Conditional Generation, Abbas Mammadov et al. (2026). Variational Flow Maps push efficient conditional image generation toward one or a few steps, complementing FlexTok’s reduction of the token budget.
