Any-Order Flexible Length Masked Diffusion
Jaeyeon KimC. LeeCarles Domingo-EnrichYilun DuS. KakadeTimothy NgotiaocoSitan ChenM. S. Albergo
Introduces Flexible Masked Diffusion Models (FlexMDMs) to eliminate the fixed-length constraint of discrete diffusion models by enabling dynamic token insertion alongside any-order generation, substantially improving mathematical reasoning and code infilling when fine-tuned on large language models.
Modern generative artificial intelligence models for discrete sequences, such as text and computer code, are increasingly exploring masked diffusion architectures as an alternative to traditional left-to-right autoregressive systems. Masked diffusion models offer substantial advantages in parallel processing and arbitrary-order text generation. However, they face a critical operational constraint: they operate strictly on fixed-length sequences and cannot dynamically insert tokens during generation. To generate variable-length outputs, current implementations must rely on static padding heuristics, which limit their adaptability in structured tasks such as multi-step planning, code editing, and variable-length reasoning.
To address this limitation, the article introduces Flexible Masked Diffusion Models (FlexMDM), a discrete generative framework designed to model sequences of variable length natively while preserving the ability to generate tokens in any order. The central objective of the article is to demonstrate that an explicit mathematical framework based on continuous-time Markov chains and joint stochastic interpolants can train models to simultaneously predict which tokens to unmask and estimate how many tokens must be inserted, all starting from an empty string.
To evaluate this architecture, the researchers performed theoretical derivations alongside a progression of empirical benchmarks across multiple domains and scales. They trained small-scale models from scratch on raw text from the OpenWebText corpus and tested them on a synthetic discrete maze-planning task where models had to connect intermediate subgoals. To establish real-world scalability and cost efficiency, the team adapted a pretrained 8-billion-parameter masked diffusion model (LLaDA-8B) into a flexible model using parameter-efficient fine-tuning (LoRA) on 16 GPUs over a three-day period, subsequent to which the model was evaluated on standard mathematical reasoning (GSM8K) and code-infilling (HumanEval) benchmarks.
Key findings show that the proposed framework delivers marked improvements in structural fidelity and task performance without degrading generation quality. First, the flexible model matched the text fluency of standard masked diffusion baselines while accurately capturing the true underlying length distribution of the data, which standard diffusion models failed to calibrate even with high computational budgets. Second, on the discrete maze-planning task, the model achieved a 90% success rate on the hardest configuration, outperforming the baseline model by approximately 60 percentage points because it did not require preallocating token slots for subgoals. Third, when scaled to the 8-billion-parameter level, the adapted model improved mathematical problem-solving accuracy from 58% to 67% and code infilling success from 52% to 65%, with performance continuing to rise as more inference compute was allocated.
The implications of these results are significant for engineering teams seeking to deploy non-autoregressive generative models. By eliminating fixed-canvas constraints and enabling dynamic token insertions, the framework expands the applicability of diffusion architectures to complex reasoning, localized text editing, and automated code completion. Importantly, the ability to retrofit existing pretrained models within days using standard compute hardware indicates low deployment barriers and minimal capital expenditure compared to training foundation models from scratch.
For technical leaders and practitioners, adopting this flexible framework is recommended when developing applications that require structured planning, code infilling, or non-linear document editing. Engineering teams working with masked diffusion backbones should prioritize adding scalar insertion prediction heads and adapting pretrained weights rather than relying on inefficient padding schemes. Future development should focus on expanding instruction fine-tuning across broader, more diverse datasets to assess general-purpose language capabilities beyond specialized math and coding benchmarks.
Limitations of the study include its primary reliance on domain-specific fine-tuning sets and the use of zero-shot evaluations on selected benchmarks. While theoretical guarantees hold under exact posterior estimation, real-world deployments remain subject to discretization errors and the capacity of the underlying neural network. Nevertheless, the consistency of gains across synthetic planning, pretraining, and scaled benchmarks provides high confidence in the framework's viability as a general enhancement for discrete diffusion modeling.
- Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). Its absorbing-mask discrete diffusion framework provides the core modeling background for the source’s masked-token denoising approach.
- Paper: Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution, Aaron Lou et al. (2024). SEDD develops discrete diffusion language modeling with absorbing masks, preparing readers for the source’s advances to masked diffusion generation.
- Paper: Large Language Diffusion Models, Shen Nie et al. (2025). LLaDA establishes the pretrained masked diffusion language model that the source demonstrates can be retrofitted into a flexible-length model.
- Paper: DiffusER: Discrete Diffusion via Edit-based Reconstruction, Machel Reid et al. (2023). Its edit-based discrete diffusion introduces token insertion and deletion, making it useful context for the source’s insertion-based solution to fixed-length generation.
- Paper: Looped Diffusion Language Models, Sanghyun Lee et al. (2026). LoopMDM carries masked diffusion language modeling into parameter-efficient iterative computation, extending the source’s exploration of improved MDM capabilities.
- Paper: DiffusionGemma Technical Report, DiffusionGemma Team et al.. DiffusionGemma applies diffusion language modeling at large-model scale, continuing the source’s path toward practical, fast text generation.
- Paper: Unlocking Lossless Speedups in LLMs via Discrete Diffusion, Subham Sekhar Sahoo et al. (2026). Uno builds on diffusion-based parallel text generation to deliver verified, lossless speedups for autoregressive models.
- Paper: Context-weighted Discrete Flow Matching, Daniil Cherniavskii et al. (2026). Context-weighted discrete flow matching advances any-order generation by making denoising updates depend on available context.
- Paper: Flow Reasoning Models: Turning Flows Into Efficient Recurrent Reasoners, Alec Helbling et al. (2026). Flow Reasoning Models extend discrete flow generation into recurrent refinement for structured reasoning tasks.
- Paper: ELF: Embedded Language Flows, Keya Hu et al. (2026). ELF continues diffusion language modeling through continuous embedding-space flows, exploring a complementary route to efficient text generation.
