FlexTok: Resampling Images into 1D Token Sequences of Flexible Length

Roman BachmannJesse AllardiceDavid MizrahiEnrico FiniOguzhan Fatih KarElmira AmirlooAlaaeldin El-NoubyAmir ZamirAfshin Dehghan

article2025ICML116 citations

Introduces an image tokenizer that encodes 2D visual content into ordered, variable-length 1D token sequences paired with a rectified flow decoder, enabling autoregressive models to generate high-quality images from coarse to fine details using as few as eight tokens.

Listen

Modern generative vision systems rely on image tokenization to compress high-dimensional pixels into discrete representations that models can predict. Standard approaches encode images into rigid two-dimensional grids where the sequence length is determined strictly by the image dimensions rather than visual complexity. Consequently, models must always predict a full, fixed budget of tokens regardless of whether the target image is a basic object or an intricate scene, creating computational inefficiencies and rigid generation pipelines.

The article introduces and evaluates FlexTok, a unified framework that resamples images into ordered, variable-length one-dimensional token sequences. FlexTok allows an image to be represented and reconstructed using anywhere from 1 to 256 discrete tokens within a single model architecture.

To achieve this, the authors combine a Vision Transformer encoder utilizing learnable register tokens with finite scalar quantization, causal attention masking, and nested dropout to enforce a coarse-to-fine token ordering. The tokens condition an end-to-end rectified flow decoder that reconstructs the image from intermediate latents, enhanced with an inductive bias feature loss to accelerate training. The approach was systematically evaluated across standard reconstruction benchmarks and generative tasks, including class-conditioned generation on ImageNet-1k and text-to-image synthesis on large-scale captioned datasets using autoregressive models scaled up to 3 billion parameters.

The evaluation yielded several key findings. First, FlexTok produces visually plausible reconstructions across extreme compression rates, achieving a reconstruction Fréchet Inception Distance below 2.0 with as few as 8 tokens and reaching 1.08 with 256 tokens. Second, high-level semantic and geometric concepts naturally emerge in the earliest tokens, while fine details are added in later tokens. Third, in class-conditioned generation, the top-performing FlexTok setup paired with an autoregressive model matched or outperformed existing baselines such as TiTok and 2D grid models, achieving a generation Fréchet Inception Distance of 1.86 at 32 tokens while requiring roughly one-eighth the sequence length of traditional 256-token grids. Finally, downstream token demand scales with prompt complexity: simple tasks require only 4 to 16 tokens to fulfill conditioning, whereas complex open-ended text prompts benefit from using up to 256 tokens.

These findings demonstrate that generative models do not require uniform compute budgets for every image. By structuring visual information into a coarse-to-fine visual vocabulary, systems can terminate generation early on simple tasks, potentially yielding significant inference speedups and lower computational costs. Furthermore, the class-conditional autoregressive models achieve optimal performance without requiring classifier-free guidance, eliminating the traditional twofold compute overhead associated with generating unconditioned passes.

Practitioners should explore early-stopping mechanisms during inference to dynamically adjust generation length based on prompt complexity. Organizations deploying autoregressive generative vision models should evaluate one-dimensional flexible-length tokenization as an alternative to fixed two-dimensional grids to lower computational costs. When high fidelity across diverse sequence lengths is required, practitioners should adopt rectified flow decoding alongside representation alignment losses.

However, some limitations remain. While the flow-based decoder maintains high perceptual quality with few tokens, semantic variance across different random seeds is higher when using very low token counts. In addition, the adaptive conditioning components in the decoder currently add substantial parameter overhead. The findings provide high confidence for standard single-image generation benchmarks at 256x256 resolution, but further research is required to evaluate scaling behavior on higher-resolution imagery, dense document rendering, and other domains such as video.

Bachmann et al (2025).pdf
Cover for FlexTok: Resampling Images into 1D Token Sequences of Flexible Length

Abstract

We introduce FlexTok, a tokenizer that projects 2D images into variable-length, ordered 1D token sequences. For example, a 256×256 image can be resampled into anywhere from 1 to 256 discrete tokens, hierarchically and semantically compressing its information. By training a rectified flow model as the decoder and using nested dropout, FlexTok produces plausible reconstructions regardless of the chosen token sequence length. We evaluate our approach in an autoregressive generation setting using a simple GPT-style Transformer. On ImageNet, this approach achieves an FID < 2 across 8 to 128 tokens, outperforming TiTok and matching state-of-the-art methods with far fewer tokens. We further extend the model to support to text-conditioned image generation and examine how FlexTok relates to traditional 2D tokenization. A key finding is that FlexTok enables next-token prediction to describe images in a coarse-to-fine “visual vocabulary”, and that the number of tokens to generate depends on the complexity of the generation task.

Citation

MLA
Bachmann, R., et al. “FlexTok: Resampling Images into 1D Token Sequences of Flexible Length”. arXiv, 2025, https://doi.org/10.48550/arxiv.2502.13967.
APA
Bachmann, R., Allardice, J., Mizrahi, D., Fini, E., Kar, O. F., Amirloo, E., El-Nouby, A., Zamir, A., & Dehghan, A. (2025). FlexTok: Resampling Images into 1D Token Sequences of Flexible Length. arXiv. https://doi.org/10.48550/arxiv.2502.13967
Chicago
Bachmann, R., J. Allardice, D. Mizrahi, et al. 2025. “FlexTok: Resampling Images into 1D Token Sequences of Flexible Length”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2502.13967.
Harvard
Bachmann, R. et al. (2025) “FlexTok: Resampling Images into 1D Token Sequences of Flexible Length”. arXiv. Available at: https://doi.org/10.48550/arxiv.2502.13967.
Vancouver
1. Bachmann R, Allardice J, Mizrahi D, Fini E, Kar OF, Amirloo E, El-Nouby A, Zamir A, Dehghan A (2025) FlexTok: Resampling Images into 1D Token Sequences of Flexible Length. https://doi.org/10.48550/arxiv.2502.13967

BibTeX

@misc{https://doi.org/10.48550/arxiv.2502.13967,
  doi = {10.48550/ARXIV.2502.13967},
  url = {https://arxiv.org/abs/2502.13967},
  author = {Bachmann, Roman and Allardice, Jesse and Mizrahi, David and Fini, Enrico and Kar, Oğuzhan Fatih and Amirloo, Elmira and El-Nouby, Alaaeldin and Zamir, Amir and Dehghan, Afshin},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {FlexTok: Resampling Images into 1D Token Sequences of Flexible Length},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/