Any-Order Flexible Length Masked Diffusion

Jaeyeon KimC. LeeCarles Domingo-EnrichYilun DuS. KakadeTimothy NgotiaocoSitan ChenM. S. Albergo

article2025arXiv51 citations

Introduces Flexible Masked Diffusion Models (FlexMDMs) to eliminate the fixed-length constraint of discrete diffusion models by enabling dynamic token insertion alongside any-order generation, substantially improving mathematical reasoning and code infilling when fine-tuned on large language models.

Listen

Modern generative artificial intelligence models for discrete sequences, such as text and computer code, are increasingly exploring masked diffusion architectures as an alternative to traditional left-to-right autoregressive systems. Masked diffusion models offer substantial advantages in parallel processing and arbitrary-order text generation. However, they face a critical operational constraint: they operate strictly on fixed-length sequences and cannot dynamically insert tokens during generation. To generate variable-length outputs, current implementations must rely on static padding heuristics, which limit their adaptability in structured tasks such as multi-step planning, code editing, and variable-length reasoning.

To address this limitation, the article introduces Flexible Masked Diffusion Models (FlexMDM), a discrete generative framework designed to model sequences of variable length natively while preserving the ability to generate tokens in any order. The central objective of the article is to demonstrate that an explicit mathematical framework based on continuous-time Markov chains and joint stochastic interpolants can train models to simultaneously predict which tokens to unmask and estimate how many tokens must be inserted, all starting from an empty string.

To evaluate this architecture, the researchers performed theoretical derivations alongside a progression of empirical benchmarks across multiple domains and scales. They trained small-scale models from scratch on raw text from the OpenWebText corpus and tested them on a synthetic discrete maze-planning task where models had to connect intermediate subgoals. To establish real-world scalability and cost efficiency, the team adapted a pretrained 8-billion-parameter masked diffusion model (LLaDA-8B) into a flexible model using parameter-efficient fine-tuning (LoRA) on 16 GPUs over a three-day period, subsequent to which the model was evaluated on standard mathematical reasoning (GSM8K) and code-infilling (HumanEval) benchmarks.

Key findings show that the proposed framework delivers marked improvements in structural fidelity and task performance without degrading generation quality. First, the flexible model matched the text fluency of standard masked diffusion baselines while accurately capturing the true underlying length distribution of the data, which standard diffusion models failed to calibrate even with high computational budgets. Second, on the discrete maze-planning task, the model achieved a 90% success rate on the hardest configuration, outperforming the baseline model by approximately 60 percentage points because it did not require preallocating token slots for subgoals. Third, when scaled to the 8-billion-parameter level, the adapted model improved mathematical problem-solving accuracy from 58% to 67% and code infilling success from 52% to 65%, with performance continuing to rise as more inference compute was allocated.

The implications of these results are significant for engineering teams seeking to deploy non-autoregressive generative models. By eliminating fixed-canvas constraints and enabling dynamic token insertions, the framework expands the applicability of diffusion architectures to complex reasoning, localized text editing, and automated code completion. Importantly, the ability to retrofit existing pretrained models within days using standard compute hardware indicates low deployment barriers and minimal capital expenditure compared to training foundation models from scratch.

For technical leaders and practitioners, adopting this flexible framework is recommended when developing applications that require structured planning, code infilling, or non-linear document editing. Engineering teams working with masked diffusion backbones should prioritize adding scalar insertion prediction heads and adapting pretrained weights rather than relying on inefficient padding schemes. Future development should focus on expanding instruction fine-tuning across broader, more diverse datasets to assess general-purpose language capabilities beyond specialized math and coding benchmarks.

Limitations of the study include its primary reliance on domain-specific fine-tuning sets and the use of zero-shot evaluations on selected benchmarks. While theoretical guarantees hold under exact posterior estimation, real-world deployments remain subject to discretization errors and the capacity of the underlying neural network. Nevertheless, the consistency of gains across synthetic planning, pretraining, and scaled benchmarks provides high confidence in the framework's viability as a general enhancement for discrete diffusion modeling.

  • Paper: Looped Diffusion Language Models, Sanghyun Lee et al. (2026). LoopMDM carries masked diffusion language modeling into parameter-efficient iterative computation, extending the source’s exploration of improved MDM capabilities.
  • Paper: DiffusionGemma Technical Report, DiffusionGemma Team et al.. DiffusionGemma applies diffusion language modeling at large-model scale, continuing the source’s path toward practical, fast text generation.
  • Paper: Unlocking Lossless Speedups in LLMs via Discrete Diffusion, Subham Sekhar Sahoo et al. (2026). Uno builds on diffusion-based parallel text generation to deliver verified, lossless speedups for autoregressive models.
  • Paper: Context-weighted Discrete Flow Matching, Daniil Cherniavskii et al. (2026). Context-weighted discrete flow matching advances any-order generation by making denoising updates depend on available context.
  • Paper: Flow Reasoning Models: Turning Flows Into Efficient Recurrent Reasoners, Alec Helbling et al. (2026). Flow Reasoning Models extend discrete flow generation into recurrent refinement for structured reasoning tasks.
  • Paper: ELF: Embedded Language Flows, Keya Hu et al. (2026). ELF continues diffusion language modeling through continuous embedding-space flows, exploring a complementary route to efficient text generation.
Cover for Any-Order Flexible Length Masked Diffusion

Abstract

Masked diffusion models (MDMs) have recently emerged as a promising alternative to autoregressive models over discrete domains. MDMs generate sequences in an any-order, parallel fashion, enabling fast inference and strong performance on non-causal tasks. However, a crucial limitation is that they do not support token insertions and are thus limited to fixed-length generations. To this end, we introduce Flexible Masked Diffusion Models (FlexMDMs), a discrete diffusion paradigm that simultaneously can model sequences of flexible length while provably retaining MDMs' flexibility of any-order inference. Grounded in an extension of the stochastic interpolant framework, FlexMDMs generate sequences by inserting mask tokens and unmasking them. Empirically, we show that FlexMDMs match MDMs in perplexity while modeling length statistics with much higher fidelity. On a synthetic maze planning task, they achieve ≈60%\approx 60 \% higher success rate than MDM baselines. Finally, we show pretrained MDMs can easily be retrofitted into FlexMDMs: on 16 H100s, it takes only three days to fine-tune LLaDA-8B into a FlexMDM, achieving superior performance on math (GSM8K, 58%→67%58\% \to 67\%) and code infilling performance (52%→65%52\% \to 65\%).

Table of Contents

  • 1 Introduction
  • 2 Preliminaries: Continuous-Time Markov Chains and Masked Diffusions
  • 2.1 Masked Diffusion Models
  • 3 Variable Length Masked Diffusions: Training
  • 4 Variable Length Masked Diffusions: Inference
  • 5 Experiment
  • 5.1 Pretraining
  • 5.1.1 Pretraining on text data
  • 5.1.2 Planning task
  • 5.2 Scaling up FlexMDM
  • 6 Conclusion
  • References
  • A Related Works
  • B Notation
  • C Discrete Stochastic Interpolants: Definitions and Propositions for Section
  • C.1 Discrete Stochastic Interpolant
  • C.2 The Masked Diffusion Interpolant
  • D Joint Discrete Stochastic Interpolants: Definitions and Propositions for Section
  • D.1 Joint Interpolant
  • D.2 Flexible-Length Masked Diffusion
  • E Details for Section
  • E.1 Precise detail on the inference algorithms
  • E.2 Proof of FlexMDM’s any-order inference capability
  • E.2.1 Proof preliminaries
  • E.3 Formal guarantee for adaptive inference
  • E.4 Proof of Theorem
  • E.5 Proof of Lemma
  • E.6 Proof of Lemma
  • F Experimental details
  • F.1 Pretraining on OpenWebText
  • F.2 Pretraining on the Maze planning task
  • F.3 Weight Initialization training from LLaDA

Knowls

  1. Knowl 1 — FlexMDM constructs variable-length sequences by inserting and unmasking tokens

    model/method

    Let a clean data sequence be x1=(x10,…,x1n−1)x_1=(x_1^0,\ldots,x_1^{n-1}), sampled from a distribution that may contain different sequence lengths. Choose differentiable, nondecreasing schedules αt,βt∈[0,1]\alpha_t,\beta_t\in[0,1] for t∈[0,1]t\in[0,1], with α0=β0=0\alpha_0=\beta_0=0 and α1=β1=1\alpha_1=\beta_1=1. For each source position ii, independently sample an insertion time T1iT_1^i with density α˙t\dot\alpha_t and, conditional on T1i=uT_1^i=u, an unmasking time T2i≥uT_2^i\ge u with density β˙t/(1−βu)\dot\beta_t/(1-\beta_u). At time tt, source token x1ix_1^i is absent if t<T1it<T_1^i, represented by a mask token mm if T1i≤t<T2iT_1^i\le t<T_2^i, and visible as x1ix_1^i if t≥T2it\ge T_2^i. Concatenate the non-absent tokens in their original order to form xtx_t. The auxiliary ordered index set sts_t records which source positions have been inserted, so it tracks the alignment between xtx_t and x1x_1. At t=0t=0 the sequence is empty, while at t=1t=1 it equals x1x_1. This joint state (xt,st)(x_t,s_t) supports variable-length generation without requiring the model to know the target length at initialization.

  2. Knowl 2 — FlexMDM rates combine a clean-token posterior with an insertion expectation

    equation

    For a current sequence xx and its latent ordered alignment ss to a clean sequence, FlexMDM has two kinds of transitions. If current position ii contains a mask, replacing it with vocabulary token vv has rate Rt(x,x[xi←v])=β˙t1−βt Pr⁡(x1s[i]=v∣xt=x)R_t(x,x[x^i\leftarrow v])=\frac{\dot\beta_t}{1-\beta_t}\,\Pr(x_1^{s[i]}=v\mid x_t=x). Inserting a mask at gap ii, written x◃imx\triangleleft_i m, has rate Rt(x,x◃im)=α˙t1−αt E[s[i]−s[i−1]−1∣xt=x]R_t(x,x\triangleleft_i m)=\frac{\dot\alpha_t}{1-\alpha_t}\,\mathbb{E}[s[i]-s[i-1]-1\mid x_t=x]. Here x1s[i]x_1^{s[i]} is the clean token aligned to current position ii; s[i]−s[i−1]−1s[i]-s[i-1]-1 counts source tokens not yet inserted in that gap; and the boundary indices are s[−1]=−1s[-1]=-1 and s[∣x∣]=ns[|x|]=n, where n=∣x1∣n=|x_1|. The insertion formula applies to each of the ∣x∣+1|x|+1 gaps, including the two sequence boundaries. Thus the model needs a posterior distribution over clean tokens at masked positions and a scalar expected number of still-missing tokens per gap. These rates induce the marginal path distribution of the FlexMDM interpolant.

  3. Knowl 3 — The FlexMDM loss learns the two sufficient rate quantities and bounds terminal error

    theoretical result

    Given an interpolant sample (xt,st)(x_t,s_t) drawn from a clean sequence x1x_1, let fθ(xt,t)[i,⋅]f_\theta(x_t,t)[i,\cdot] predict the clean-token distribution at a masked position and let gθ(xt,t)[i]>0g_\theta(x_t,t)[i]>0 predict the missing-token count in gap ii. With ki=s[i]−s[i−1]−1k_i=s[i]-s[i-1]-1 and ϕ(k,g)=g−klog⁡g\phi(k,g)=g-k\log g, the training objective is Lθ=∫01E[β˙t1−βt∑i:xti=m−log⁡fθ(xt,t)[i,x1s[i]]+α˙t1−αt∑i=0∣xt∣ϕ(ki,gθ(xt,t)[i])]dt\mathcal{L}_\theta=\int_0^1\mathbb{E}\left[\frac{\dot\beta_t}{1-\beta_t}\sum_{i:x_t^i=m}-\log f_\theta(x_t,t)[i,x_1^{s[i]}]+\frac{\dot\alpha_t}{1-\alpha_t}\sum_{i=0}^{|x_t|}\phi(k_i,g_\theta(x_t,t)[i])\right]dt. The expectation is over clean data and the insertion/unmasking randomness of the interpolant. The objective is uniquely minimized by fθ(x,t)[i,v]=Pr⁡(x1s[i]=v∣xt=x)f_\theta(x,t)[i,v]=\Pr(x_1^{s[i]}=v\mid x_t=x) and gθ(x,t)[i]=E[ki∣xt=x]g_\theta(x,t)[i]=\mathbb{E}[k_i\mid x_t=x]. If p1p_1 is the data distribution and p1θp_1^\theta is the terminal distribution generated using the learned rates, then DKL(p1∥p1θ)≤Lθ−L⋆D_{\mathrm{KL}}(p_1\Vert p_1^\theta)\le\mathcal{L}_\theta-\mathcal{L}_\star, where L⋆\mathcal{L}_\star is the global minimum. Therefore, exact minimization recovers the target distribution under the stated interpolant and rate construction.

  4. Knowl 4 — FlexMDM permits any-order unmasking while retaining exact sampling

    theoretical result

    Assume access to the ground-truth conditional distribution of clean tokens at masked positions and the ground-truth expected insertion counts. A sampler may choose any currently masked position, including through an adaptive or confidence-based rule, and reveal a token sampled from that position's correct posterior marginal. It may interleave these unmasking steps arbitrarily with insertion steps governed by the FlexMDM insertion-only continuous-time Markov chain, whose rate at gap ii is α˙t1−αtE[s[i]−s[i−1]−1∣xt=x]\frac{\dot\alpha_t}{1-\alpha_t}\mathbb{E}[s[i]-s[i-1]-1\mid x_t=x]. After the insertion path reaches t=1t=1, any remaining masks can be unmasked using the target-compatible posterior. The resulting sequence is an exact sample from the data distribution, regardless of the chosen unmasking order. In particular, the unmasking order need not follow the unmasking schedule used to construct the training interpolant; the guarantee assumes correct posterior and insertion quantities and exact continuous-time insertion dynamics.

  5. Knowl 5 — A joint interpolant yields a valid marginal continuous-time Markov chain

    theoretical result

    A joint discrete stochastic interpolant consists of coupled random variables (Xt,St)(X_t,S_t) connecting a base state (X0,S0)(X_0,S_0) to a data state (X1,S1)(X_1,S_1), together with a conditional continuous-time Markov chain on the joint state space. Let Ktx0,x1((x,s),(y,s′))K_t^{x_0,x_1}((x,s),(y,s')) be its transition rate when the endpoints are x0,x1x_0,x_1. A rate matrix on the sequence alone is obtained by averaging over the auxiliary state and endpoints conditional on the observed sequence: Rt(x,y)=E[∑s′KtX0,X1((x,St),(y,s′))∣Xt=x]R_t(x,y)=\mathbb{E}\left[\sum_{s'}K_t^{X_0,X_1}((x,S_t),(y,s'))\mid X_t=x\right]. The continuous-time Markov chain with this marginal rate matrix has sequence marginals equal to Law⁡(Xt)\operatorname{Law}(X_t). The auxiliary state can therefore retain alignment or other path information needed to define simple transitions, even when that information is not available in the generated sequence itself.

  6. Knowl 6 — Vanilla FlexMDM inference uses tau-leaping for unmasking and insertion

    algorithm

    The paper's vanilla sampler discretizes time and batches events over each interval. Its inputs are learned posterior and insertion heads fθ,gθf_\theta,g_\theta, schedules α,β\alpha,\beta, a grid 0=t1<⋯<tN=10=t_1<\cdots<t_N=1, and the empty initial sequence; it returns the sequence at t=1t=1. For each interval, compute τ=tk+1−tk\tau=t_{k+1}-t_k and use rates evaluated at tkt_k. For every masked position ii and clean token vv, draw an independent count cv∼Poisson⁡(τβ˙tk1−βtkfθ(x,tk)[i,v])c_v\sim\operatorname{Poisson}(\tau\frac{\dot\beta_{t_k}}{1-\beta_{t_k}}f_\theta(x,t_k)[i,v]); unmask position ii only if exactly one token has count one and every other token has count zero, assigning that token. Then, for every gap ii, draw ℓi∼Poisson⁡(τα˙tk1−αtkgθ(x,tk)[i])\ell_i\sim\operatorname{Poisson}(\tau\frac{\dot\alpha_{t_k}}{1-\alpha_{t_k}}g_\theta(x,t_k)[i]) and insert ℓi\ell_i mask tokens in that gap. Repeat through the grid and return the final sequence. The paper uses tau-leaping, which batches transitions within each interval; as the grid is refined, the discretization approaches the continuous-time chain.

  7. Knowl 7 — Known context can be held fixed while FlexMDM generates a variable-length target

    model/method

    For conditional generation, FlexMDM can exclude known context tokens from the insertion/unmasking interpolant, keeping those tokens unchanged at every time while applying the process to the unknown target. In maze planning, the subgoal tokens are fixed and the model inserts and fills in tokens for the connecting path. For math instruction fine-tuning, instruction tokens are fixed while the answer is generated. For code infilling, the prefix, suffix, and separators are fixed while the middle span is generated; the FlexMDM initial state contains the two span boundaries but no masks revealing the desired middle-span length. This avoids giving the model target-length information through a preallocated masked span. The same conditioning principle is used for the fixed-length MDM comparisons, with each model's context held outside its respective generation process.

  8. Knowl 8 — OpenWebText experiments show improved length calibration without a fluency penalty

    empirical result

    The from-scratch text experiment used variable-length OpenWebText paragraphs, with sequences capped at 1,024 tokens, and compared 175M-parameter FlexMDM and MDM models trained for 500K iterations with batch size 1,024. The MDM baseline handled variable-length data by padding to a fixed maximum length. Evaluation varied the number of sampling steps and measured generated-text perplexity as a fluency proxy and the generated sequence-length distribution. FlexMDM tracked the training-data length distribution closely with 256 sampling steps, whereas the MDM length distribution remained miscalibrated even with 1,024 steps. The models had comparable generative perplexity, and perplexity improved for both as sampling steps increased. The reported comparison therefore indicates better length-statistics fidelity for FlexMDM without an observed generative-perplexity penalty.

  9. Knowl 9 — FlexMDM substantially improves subgoal-conditioned maze success

    empirical result

    The planning benchmark used 41 × 41 mazes containing invalid cells. Given ordered subgoals, a model had to generate a path that visited them in order without entering invalid cells; task difficulty was varied using K∈{2,7,12}K\in\{2,7,12\} subgoals. Both models were evaluated with 256 sampling steps, and success required satisfying both path validity and ordered-subgoal constraints. The success rates were: Easy, MDM 68.4% and FlexMDM 92.3%; Medium, MDM 29.3% and FlexMDM 90.4%; Hard, MDM 24.2% and FlexMDM 90.0%. FlexMDM maintained roughly 90% success as difficulty increased, while the fixed-length MDM baseline degraded sharply. The result illustrates the advantage of inserting path tokens when the locations and number of steps needed between subgoals are not known in advance.

  10. Knowl 10 — An 8B pretrained MDM can be retrofitted into FlexMDM with downstream gains

    empirical result

    The scaling experiment initialized FlexMDM from the LLaDA-8B masked diffusion model, adding a time-embedding pathway and a scalar insertion-expectation head, and training LoRA adapters; approximately 400M parameters were trainable. The base conversion used a 50:50 mixture of OpenWebText and Proof-Pile-2 and took 200,000 gradient steps with batch size 64 on 16 H100 GPUs, approximately three days. The resulting model was instruction-fine-tuned and compared with LLaDA initialized and fine-tuned using the same number of task-specific pairs. On GSM8K, reported Pass@1 improved from 58% for LLaDA to 67% for FlexMDM; on HumanEval single-line code infilling, success improved from 52% to 65%. FlexMDM performance continued to improve as more sampling steps were allocated, while the LLaDA baseline was comparatively flat. These results demonstrate transfer from an MDM initialization and improved performance in the evaluated math and infilling tasks.

Coverage note — The paper's detailed derivations and proof-only intermediate lemmas are omitted because they do not add standalone contributed results beyond the stated rate, training, and inference guarantees; exhaustive optimizer and architecture hyperparameters are omitted where they do not change the principal findings.

References

  1. 1.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  2. 2.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  3. 3.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning, pages 2256–2265. pmlr, 2015.
  4. 4.Jiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet, and Michalis Titsias. Simplified and generalized masked diffusion for discrete data. Advances in neural information processing systems, 37:103131–103167, 2024.
  5. 5.Subham Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin Chiu, Alexander Rush, and Volodymyr Kuleshov. Simple and effective masked diffusion language models. Advances in Neural Information Processing Systems, 37:130136–130184, 2024.
  6. 6.Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky TQ Chen, Gabriel Synnaeve, Yossi Adi, and Yaron Lipman. Discrete flow matching. Advances in Neural Information Processing Systems, 37:133345–133385, 2024.
  7. 7.Jiacheng Ye, Jiahui Gao, Shansan Gong, Lin Zheng, Xin Jiang, Zhenguo Li, and Lingpeng Kong. Beyond autoregression: Discrete diffusion for complex reasoning and planning. arXiv preprint arXiv:2410.14157, 2024.
  8. 8.Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, and Chongxuan Li. Large language diffusion models. arXiv preprint arXiv:2502.09992, 2025.
  9. 9.Jiacheng Ye, Zhihui Xie, Lin Zheng, Jiahui Gao, Zirui Wu, Xin Jiang, Zhenguo Li, and Lingpeng Kong. Dream 7b, 2025. URL https://hkunlp.github.io/blog/2025/dream.
  10. 10.Shen Nie, Fengqi Zhu, Chao Du, Tianyu Pang, Qian Liu, Guangtao Zeng, Min Lin, and Chongxuan Li. Scaling up masked diffusion models on text. arXiv preprint arXiv:2410.18514, 2024.
  11. 11.Michael S. Albergo and Eric Vanden-Eijnden. Building normalizing flows with stochastic interpolants, 2022.
  12. 12.Michael S Albergo, Nicholas M Boffi, and Eric Vanden-Eijnden. Stochastic interpolants: A unifying framework for flows and diffusions. arXiv preprint arXiv:2303.08797, 2023a.
  13. 13.Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling, 2022.
  14. 14.Kaiwen Zheng, Yongxin Chen, Hanzi Mao, Ming-Yu Liu, Jun Zhu, and Qinsheng Zhang. Masked diffusion models are secretly time-agnostic masked models and exploit inaccurate categorical sampling. arXiv preprint arXiv:2409.02908, 2024.
  15. 15.Jingyang Ou, Shen Nie, Kaiwen Xue, Fengqi Zhu, Jiacheng Sun, Zhenguo Li, and Chongxuan Li. Your absorbing discrete diffusion secretly models the conditional distributions of clean data. arXiv preprint arXiv:2406.03736, 2024.
  16. 16.Zirui Wu, Lin Zheng, Zhihui Xie, Jiacheng Ye, Jiahui Gao, Yansong Feng, Zhenguo Li, Victoria W., Guorui Zhou, and Lingpeng Kong. Dreamon: Diffusion language models for code infilling beyond fixed-size canvas, 2025a. URL https://hkunlp.github.io/blog/2025/dreamon.
  17. 17.Marton Havasi, Brian Karrer, Itai Gat, and Ricky TQ Chen. Edit flows: Flow matching with edit operations. arXiv preprint arXiv:2506.09018, 2025.
  18. 18.Andrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth, George Deligiannidis, and Arnaud Doucet. A continuous time framework for discrete denoising models. Advances in Neural Information Processing Systems, 35:28266–28279, 2022.
  19. 19.Jaeyeon Kim, Kulin Shah, Vasilis Kontonis, Sham Kakade, and Sitan Chen. Train for the worst, plan for the best: Understanding token ordering in masked diffusions. arXiv preprint arXiv:2502.06768, 2025.
  20. 20.Michael S. Albergo, Nicholas M. Boffi, Michael Lindsey, and Eric Vanden-Eijnden. Multimarginal generative modeling with stochastic interpolants, 2023b. URL https://arxiv.org/abs/2310.03695.
  21. 21.Hugo Negrel, Florentin Coeurdoux, Michael S. Albergo, and Eric Vanden-Eijnden. Multitask learning with stochastic interpolants, 2025. URL https://arxiv.org/abs/2508.04605.
  22. 22.William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023.
  23. 23.Shansan Gong, Shivam Agarwal, Yizhe Zhang, Jiacheng Ye, Lin Zheng, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han, et al. Scaling diffusion language models via adaptation from autoregressive models. arXiv preprint arXiv:2410.17891, 2024.
  24. 24.Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus, 2019.
  25. 25.Michael Janner, Yilun Du, Joshua Tenenbaum, and Sergey Levine. Planning with diffusion for flexible behavior synthesis. In International Conference on Machine Learning, 2022.
  26. 26.Michael Samuel Albergo, Mark Goldstein, Nicholas Matthew Boffi, Rajesh Ranganath, and Eric Vanden-Eijnden. Stochastic interpolants with data-dependent couplings. In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=FFILRGD0jG.
  27. 27.Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics, 2023.
  28. 28.Siyan Zhao, Devaansh Gupta, Qinqing Zheng, and Aditya Grover. d1: Scaling reasoning in diffusion large language models via reinforcement learning. arXiv preprint arXiv:2504.12216, 2025.
  29. 29.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  30. 30.Siming Huang, Tianhao Cheng, Jason Klein Liu, Jiaran Hao, Liuyihan Song, Yang Xu, J. Yang, J. H. Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Zhaoxiang Zhang, Jie Fu, Qian Liu, Ge Zhang, Zili Wang, Yuan Qi, Yinghui Xu, and Wei Chu. Opencoder: The open cookbook for top-tier code large language models. 2024. URL https://arxiv.org/pdf/2411.04905.
  31. 31.Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. Efficient training of language models to fill in the middle. arXiv preprint arXiv:2207.14255, 2022.
  32. 32.Xinyin Ma, Runpeng Yu, Gongfan Fang, and Xinchao Wang. dkv-cache: The cache for diffusion language models. arXiv preprint arXiv:2505.15781, 2025a.
  33. 33.Subham Sekhar Sahoo, Zhihan Yang, Yash Akhauri, Johnna Liu, Deepansha Singh, Zhoujun Cheng, Zhengzhong Liu, Eric Xing, John Thickstun, and Arash Vahdat. Esoteric language models. arXiv preprint arXiv:2506.01928, 2025.
  34. 34.Chengyue Wu, Hao Zhang, Shuchen Xue, Zhijian Liu, Shizhe Diao, Ligeng Zhu, Ping Luo, Song Han, and Enze Xie. Fast-dllm: Training-free acceleration of diffusion llm by enabling kv cache and parallel decoding. arXiv preprint arXiv:2505.22618, 2025b.
  35. 35.Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forre, and Max Welling. Argmax flows ´ and multinomial diffusion: Learning categorical distributions. Advances in neural information processing systems, 34:12454–12465, 2021.
  36. 36.Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tarlow, and Rianne Van Den Berg. Structured denoising diffusion models in discrete state-spaces. Advances in neural information processing systems, 34:17981–17993, 2021.
  37. 37.Aaron Lou, Chenlin Meng, and Stefano Ermon. Discrete diffusion modeling by estimating the ratios of the data distribution. arXiv preprint arXiv:2310.16834, 2023.
  38. 38.Andrew Campbell, Jason Yim, Regina Barzilay, Tom Rainforth, and Tommi Jaakkola. Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design. arXiv preprint arXiv:2402.04997, 2024.
  39. 39.Neta Shaul, Itai Gat, Marton Havasi, Daniel Severo, Anuroop Sriram, Peter Holderrieth, Brian Karrer, Yaron Lipman, and Ricky TQ Chen. Flow matching with general discrete paths: A kinetic-optimal perspective. arXiv preprint arXiv:2412.03487, 2024.
  40. 40.Fred Zhangzhi Peng, Zachary Bezemek, Sawan Patel, Jarrid Rector-Brooks, Sherwood Yao, Avishek Joey Bose, Alexander Tong, and Pranam Chatterjee. Path planning for masked diffusion model sampling. arXiv preprint arXiv:2502.03540, 2025.
  41. 41.Litu Rout, Constantine Caramanis, and Sanjay Shakkottai. Anchored diffusion language model. arXiv preprint arXiv:2505.18456, 2025.
  42. 42.Shansan Gong, Ruixiang Zhang, Huangjie Zheng, Jiatao Gu, Navdeep Jaitly, Lingpeng Kong, and Yizhe Zhang. Diffucoder: Understanding and improving masked diffusion models for code generation. arXiv preprint arXiv:2506.20639, 2025.
  43. 43.Yuxuan Song, Zheng Zhang, Cheng Luo, Pengyang Gao, Fan Xia, Hao Luo, Zheng Li, Yuehang Yang, Hongli Yu, Xingwei Qu, et al. Seed diffusion: A large-scale diffusion language model with high-speed inference. arXiv preprint arXiv:2508.02193, 2025.
  44. 44.Inception Labs, Samar Khanna, Siddhant Kharbanda, Shufan Li, Harshit Varma, Eric Wang, Sawyer Birnbaum, Ziyang Luo, Yanis Miraoui, Akash Palrecha, et al. Mercury: Ultra-fast language models based on diffusion. arXiv preprint arXiv:2506.17298, 2025.
  45. 45.Heli Ben-Hamu, Itai Gat, Daniel Severo, Niklas Nolte, and Brian Karrer. Accelerated sampling from masked diffusion models via entropy bounded unmasking. arXiv preprint arXiv:2505.24857, 2025.
  46. 46.Alexander Swerdlow, Mihir Prabhudesai, Siddharth Gandhi, Deepak Pathak, and Katerina Fragkiadaki. Unified multimodal discrete diffusion. arXiv preprint arXiv:2503.20853, 2025.
  47. 47.Google DeepMind. Gemini diffusion, 2025. URL https://blog.google/technology/google-deepmind/gemini-diffusion/.
  48. 48.Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman. Maskgit: Masked generative image transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11315–11325, 2022.
  49. 49.Long Ma, Fangwei Zhong, and Yizhou Wang. Reinforced context order recovery for adaptive reasoning and planning. arXiv preprint arXiv:2508.13070, 2025b.
  50. 50.Zhe Wang, Jiaxin Shi, Nicolas Heess, Arthur Gretton, and Michalis K Titsias. Learning-order autoregressive models with application to molecular graph generation. arXiv preprint arXiv:2503.05979, 2025.
  51. 51.Alec Radford and Jeffrey Wu. Rewon child, david luan, dario amodei, and ilya sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  52. 52.Juechu Dong, Boyuan Feng, Driss Guessous, Yanbo Liang, and Horace He. Flex attention: A programming model for generating optimized attention kernels. arXiv preprint arXiv:2412.05496, 2024.
  53. 53.Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Re. Flashattention: Fast and memory- ´ efficient exact attention with io-awareness. Advances in neural information processing systems, 35:16344–16359, 2022.
  54. 54.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv: 2307.09288, 2023.
  55. 55.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  56. 56.Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024.
  57. 57.Qwen. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024.
  58. 58.Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  59. 59.Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow, 2022. URL https://arxiv.org/abs/2209.03003.
  60. 60.Peter Holderrieth, Marton Havasi, Jason Yim, Neta Shaul, Itai Gat, Tommi Jaakkola, Brian Karrer, Ricky T. Q. Chen, and Yaron Lipman. Generator matching: Generative modeling with arbitrary markov processes. In The Thirteenth International Conference on Learning Representations, 2025a. URL https://openreview.net/forum?id=RuP17cJtZo.
  61. 61.Joe Benton, Yuyang Shi, Valentin De Bortoli, George Deligiannidis, and Arnaud Doucet. From denoising diffusions to denoising markov models, 2024. URL https://arxiv.org/abs/2211.03595.
  62. 62.Andrew Campbell, William Harvey, Christian Weilbach, Valentin De Bortoli, Tom Rainforth, and Arnaud Doucet. Trans-dimensional generative modeling via jump diffusion models, 2023. URL https://arxiv.org/abs/2305.16261.
  63. 63.Francisco Vargas, Shreyas Padhy, Denis Blessing, and Nikolas Nusken. Transport meets variational inference: Controlled monte carlo diffusions, 2025. URL https://arxiv.org/abs/2307.01050.
  64. 64.Julius Berner, Lorenz Richter, and Karen Ullrich. An optimal control perspective on diffusion-based generative modeling. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview.net/forum?id=oYIjw37pTP.
  65. 65.Stefano Peluchetti. Non-denoising forward-time diffusions, 2022. URL https://openreview.net/forum?id=oVfIKuhqfC.
  66. 66.Peter Holderrieth, Michael Samuel Albergo, and Tommi Jaakkola. LEAPS: A discrete neural sampler via locally equivariant networks. In Forty-second International Conference on Machine Learning, 2025b. URL https://openreview.net/forum?id=Hq2RniQAET.
  67. 67.Neta Shaul, Itai Gat, Marton Havasi, Daniel Severo, Anuroop Sriram, Peter Holderrieth, Brian Karrer, Yaron Lipman, and Ricky T. Q. Chen. Flow matching with general discrete paths: A kinetic-optimal perspective. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=tcvMzR2NrP.

Citation

MLA
Kim, J., et al. “Any-Order Flexible Length Masked Diffusion”. arXiv, 2025, http://arxiv.org/abs/2509.01025v2.
APA
Kim, J., Cheuk-Kit, L., Domingo-Enrich, C., Du, Y., Kakade, S., Ngotiaoco, T., Chen, S., & Albergo, M. (2025). Any-Order Flexible Length Masked Diffusion. arXiv. http://arxiv.org/abs/2509.01025v2
Chicago
Kim, J., L. Cheuk-Kit, C. Domingo-Enrich, et al. 2025. “Any-Order Flexible Length Masked Diffusion”. arXiv. http://arxiv.org/abs/2509.01025v2.
Harvard
Kim, J. et al. (2025) “Any-Order Flexible Length Masked Diffusion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2509.01025v2.
Vancouver
1. Kim J, Cheuk-Kit L, Domingo-Enrich C, Du Y, Kakade S, Ngotiaoco T, Chen S, Albergo M (2025) Any-Order Flexible Length Masked Diffusion. arXiv

BibTeX

@article{kim2025any,
  title = {Any-Order Flexible Length Masked Diffusion},
  author = {Kim, Jaeyeon and Cheuk-Kit, Lee and Domingo-Enrich, Carles and Du, Yilun and Kakade, Sham and Ngotiaoco, Timothy and Chen, Sitan and Albergo, Michael},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2509.01025v2},
  eprint = {2509.01025}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/