Masked Generative Nested Transformers with Decode Time Scaling

Sahil GoyalDebapriya TulaGagan JainPradeep ShenoyPrateek JainSujoy Paul

article2025ICML4 citations

Presents a decode-time scaling framework for masked generative transformers that shares parameters across nested sub-models and reuses cached computations, reducing visual generation inference compute nearly threefold without sacrificing quality.

Listen

Visual generation models, such as diffusion and discrete masked transformers, generate high-quality images and videos but face severe inference bottlenecks. Standard parallel decoding algorithms repeatedly pass the entire model across all decoding iterations. Running a uniform, large model throughout the entire sequence causes substantial computational redundancy because initial generation stages primarily require coarse structuring rather than fine-grained refinement.

The article introduces Masked Generative Nested Transformers with Decode Time Scaling (MaGNeTS) to accelerate inference in visual generation. The authors evaluate whether dynamic model scaling paired with key-value caching across decoding iterations can significantly lower computational requirements without compromising generation quality.

The researchers designed an architecture that scales nested sub-models extracted from a single transformer backbone, sharing all parameters without increasing memory footprint. Generation starts using the smallest nested model and progressively switches to larger models as iterations proceed. The approach incorporates key-value caching to store and reuse representations of unmasked tokens across iterations, combined with periodic cache refreshes during model transitions to preserve accuracy. The framework was trained using progressive knowledge distillation alongside ground-truth losses and evaluated across standard visual benchmarks, including class-conditional ImageNet image generation, UCF101 video generation, and Kinetics600 frame prediction.

The experimental findings show substantial efficiency gains across all evaluated tasks. First, MaGNeTS reduces computational operations by 2.5 to 3.7 times compared to strong baselines, cutting ImageNet generation compute by approximately 3x (from 1,800 to 608 GFLOPs in a 16-step setup) while maintaining competitive image quality. Second, on hardware evaluations using a single accelerator chip, real-time throughput increased by 2.5 times, rising from 22.5 to 56.3 images per second. Third, scaling model sizes upward from small to large outperformed reverse or static schedules, confirming that larger capacity is strictly necessary only in final refinement steps. Fourth, while basic caching degraded output quality, periodic cache refreshes during model transitions restored generation fidelity while retaining major computational savings.

These results demonstrate that uniform compute allocation across iterative generation is inefficient. By front-loading computation onto smaller nested structures and caching static tokens, organizations can drastically reduce infrastructure costs, operational energy consumption, and user-facing latency in generative media pipelines. The framework achieves state-of-the-art efficiency while avoiding the overhead of deploying separate, disjoint models for each stage.

Decision-makers should consider adopting dynamic compute schedules and nested architectures for parallel decoding workflows. Development teams can implement progressive model scaling schedules immediately on compatible masked architectures, while ensuring guidance mechanisms are applied only during final iterations to maximize efficiency. Future work should focus on investigating token-dependent dynamic schedules and optimizing cache-refresh mechanisms to eliminate remaining compute overheads.

The primary limitation of this method is its reliance on discrete token representations, meaning key-value caching cannot be directly transferred to continuous-space models without modification. Furthermore, overall visual fidelity remains dependent on the quality of underlying visual tokenizers, and baseline visual artifacts can persist in the output. Within the scope of discrete masked transformer generation, confidence in the reported computational savings and throughput improvements is high.

arXiv: 2502.00382

No sufficiently relevant recommendations were found.

Cover for Masked Generative Nested Transformers with Decode Time Scaling

Abstract

Recent advances in visual generation have made significant strides in producing content of exceptional quality. However, most methods suffer from a fundamental problem - a bottleneck of inference computational efficiency. Most of these algorithms involve multiple passes over a transformer model to generate tokens or denoise inputs. However, the model size is kept consistent throughout all iterations, which makes it computationally expensive. In this work, we aim to address this issue primarily through two key ideas - (a) not all parts of the generation process need equal compute, and we design a decode time model scaling schedule to utilize compute effectively, and (b) we can cache and reuse some of the computation. Combining these two ideas leads to using smaller models to process more tokens while large models process fewer tokens. These different-sized models do not increase the parameter size, as they share parameters. We rigorously experiment with ImageNet256×\times256 , UCF101, and Kinetics600 to showcase the efficacy of the proposed method for image/video generation and frame prediction. Our experiments show that with almost 3×3\times less compute than baseline, our model obtains competitive performance.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 Method
  • 5 Experiments and Results
  • 5.1 Image Generation
  • 5.2 Video Generation
  • 5.3 Ablation Studies
  • 6 Conclusion
  • References
  • A Motivation for Decode Time Model Scaling
  • B Hyper-parameter Details
  • C Additional Ablations
  • D Compute Gains
  • E Limitations.

Knowls

  1. Knowl 1 — Decode-time scaling allocates larger nested models to later generation steps

    model/method

    MaGNeTS reduces the inference cost of iterative parallel decoding by changing model size during generation instead of using the full transformer at every step. For a sequence of discrete tokens XkX^k at iteration kk, it predicts and samples tokens with the scheduled model mkm_k, then masks the tokens that remain to be generated:

    Xk+1=Mask⁡(Sample⁡(mk(Xk,c),k)).X^{k+1}=\operatorname{Mask}(\operatorname{Sample}(m_k(X^k,c),k)).

    Here cc is the conditioning input, KK is the total number of decoding iterations, and the MaskGIT-style sampling operation selects tokens according to predicted confidence and the unmasking schedule. A schedule comprises nested models mpim_{p_i} used for kik_i iterations each, with ∑iki=K\sum_i k_i=K. The downscaling factor pp specifies model size: larger pp means a smaller model. The paper uses p∈{8,4,2,1}p\in\{8,4,2,1\} and finds that progressively increasing model size, typically in the order 8→4→2→18\to4\to2\to1, gives a favorable compute–quality trade-off. Because the models are nested and share parameters, using several sizes does not add their parameter counts.

  2. Knowl 2 — Nested models obtain smaller transformers by slicing shared computations

    model/method

    MaGNeTS extracts four transformer sizes from one full model by performing partial computations in both self-attention and the MLP. For downscaling factor p∈{1,2,4,8}p\in\{1,2,4,8\}, the attention query, key, and value features use 1/p1/p of the full head-feature dimension; the attention output is then projected back to the full model dimension using partial computation. The MLP is reduced in the same way. The resulting models have approximately linearly reduced parameter count and inference compute as pp increases, while sharing the full model's parameters. In the attention layer, the feature size is reduced along the head-feature dimension, rather than by reducing the number of attention heads.

    This parameter sharing avoids the extra parameters required by independently trained models at every size: for the four relative sizes 1,1/2,1/4,1/81,1/2,1/4,1/8, separate models would total 1.8751.875 times the full-model parameter count. In the reported size-L comparison, nested and standalone models both achieved ImageNet FID 3.13.1; on UCF101, the nested model achieved FVD 96.496.4 versus 115.0115.0 for standalone models.

  3. Knowl 3 — Parallel decoding caches key–value features of committed tokens and refreshes at model changes

    algorithm

    MaGNeTS reuses transformer attention key–value (KV) features for tokens that have been unmasked and committed, so later iterations need not recompute those tokens. The cache is indexed by token position and stores the KV features needed by the transformer layers. A change in nested-model size triggers a cache clear: the cached features have dimensions associated with the previous model, and the paper finds that simply padding them to fit a larger model can harm performance. Newly unmasked tokens are added to the cache for use in the following iteration. The decoding procedure is:

    Input: Initial token sequence X0, iteration count K, nested-model schedule, condition c
    Initialize an empty KV cache
    For k from 0 to K - 1:
        Select the scheduled nested model for iteration k
        If the selected model differs from the model used at iteration k - 1:
            Clear the KV cache
        Process uncached token positions with the selected model, using cached KV features
        Obtain prediction logits and KV features for the processed positions
        Apply MaskGIT sampling with the logits and the current iteration's unmasking schedule
        Keep previously committed token IDs unchanged
        Add KV features for tokens newly unmasked by this iteration to the cache
    Return the token sequence after K iterations

    The paper uses a cosine schedule for the number of tokens unmasked per iteration. Cache clearing at model-size transitions is also the cache-refresh operation: it requires recomputing features, but limits the quality loss caused by stale or dimension-mismatched cached features.

  4. Knowl 4 — Progressive ground-truth and distillation losses train the nested models

    equation

    MaGNeTS trains smaller nested models with a mixture of ground-truth supervision and online distillation from the next larger nested model. The full model receives only ground-truth supervision. With NN nested models, input XX, target token sequence YY, supervised loss L\mathcal{L}, and mixture weight αi\alpha_i for model mim_i, the training objective is:

    Ltrain=1N[L(mN(X),Y)+∑i=1N−1(αiL(mi(X),Y)+(1−αi)L(mi(X),mi+1(X)))].\mathcal{L}_{\mathrm{train}}=\frac{1}{N}\left[\mathcal{L}(m_N(X),Y)+\sum_{i=1}^{N-1}\left(\alpha_i\mathcal{L}(m_i(X),Y)+(1-\alpha_i)\mathcal{L}(m_i(X),m_{i+1}(X))\right)\right].

    For each smaller model mim_i, the teacher is its next larger neighbor mi+1m_{i+1}, and gradients are stopped through the teacher output. The weights αi\alpha_i are linearly annealed from 11 to 00 over training, shifting the smaller models from ground-truth supervision toward distillation. The reported ablation found that ground-truth-only training was weaker on UCF101 and for the smallest ImageNet model, while distillation-only training diverged.

  5. Knowl 5 — ImageNet generation approaches MaskGIT++ quality at substantially lower compute

    data/table

    For class-conditional ImageNet generation at 256×256256\times256, MaGNeTS uses the pretrained MaskGIT tokenizer, which maps images to a 16×1616\times16 sequence of discrete tokens from a codebook of size 1024. The comparison below uses 303M-parameter models; g4g4 and g6g6 denote classifier-free guidance on only the final four or six decoding iterations. The ImageNet results table on page 6 reports lower FID and inference GFLOPs for MaskGIT++ and MaGNeTS as follows:

    Method FID ↓\downarrow IS ↑\uparrow Precision ↑\uparrow Recall ↑\uparrow Steps GFLOPs
    MaskGIT++ g4g4 2.5 260.3 83 54 12 1.3k
    MaGNeTS g4g4 3.1 254.8 85 50 12 490
    MaskGIT++ g6g6 2.3 280.6 84 51 16 1.8k
    MaGNeTS g6g6 2.9 253.1 84 51 16 608

    Thus, at matching step counts and guidance windows, MaGNeTS uses about 2.65×2.65\times less compute for 12 steps and about 3×3\times less for 16 steps, with a modestly higher FID. The paper also reports that guidance on only a few final iterations gives similar generation quality to guidance on every iteration. On a single TPUv5 chip, the page 14 throughput table reports 56.3 images/s for MaGNeTS versus 22.5 images/s for MaskGIT++ using size-XL models.

  6. Knowl 6 — MaGNeTS reduces compute for class-conditional UCF101 video generation

    data/table

    On class-conditional UCF101 video generation, MaGNeTS uses a 306M-parameter model for 12 decoding steps and generates 16-frame videos at 128×128128\times128 resolution. The page 7 results table compares it with MAGVIT-L, using the same parameter count and number of steps; no classifier-free guidance is used for this task.

    Method FVD ↓\downarrow IS ↑\uparrow Steps GFLOPs
    MAGVIT-L 74.4±274.4\pm2 89.54±0.2189.54\pm0.21 12 ∼4.3\sim4.3k
    MaGNeTS 96.4±296.4\pm2 88.53±0.2088.53\pm0.20 12 ∼1.7\sim1.7k

    MaGNeTS therefore uses roughly 2.5 times less compute while retaining a competitive inception score; its FVD is higher, indicating worse video fidelity than MAGVIT-L under the reported metric.

  7. Knowl 7 — Frame prediction on Kinetics600 trades some FVD for roughly 3.6-fold lower compute

    data/table

    For Kinetics600 frame prediction conditioned on five frames, the paper reports results for 306M-parameter MAGVIT-L and MaGNeTS models, each using 12 decoding steps. The g2g2 variants apply classifier-free guidance only during the final two steps. The page 7 results table gives:

    Method FVD ↓\downarrow IS ↑\uparrow Guidance GFLOPs
    MAGVIT-L 7.2±0.17.2\pm0.1 16.48±0.0116.48\pm0.01 None ∼4.3\sim4.3k
    MAGVIT-L g2g2 6.6±0.16.6\pm0.1 16.29±0.0116.29\pm0.01 Final 2 steps ∼5.1\sim5.1k
    MaGNeTS 10.8±0.110.8\pm0.1 16.25±0.0216.25\pm0.02 None ∼1.2\sim1.2k
    MaGNeTS g2g2 9.6±0.19.6\pm0.1 16.25±0.0116.25\pm0.01 Final 2 steps ∼1.4\sim1.4k

    Relative to the corresponding MAGVIT-L setting, MaGNeTS uses about 3.6 times less compute. Its FVD is higher, while its inception score is close to the reported MAGVIT-L scores.

  8. Knowl 8 — Gradual upscaling outperforms abrupt or downward model schedules

    empirical result

    The scheduling ablation evaluates a size-L model with nested sizes ordered by downscaling factors p=(8,4,2,1)p=(8,4,2,1) over 12 ImageNet decoding iterations. The compute–FID plots on page 8 show that, at a given compute budget, a schedule that changes model size can outperform using a single nested size for all iterations. Schedules with smoother transitions, such as (3,3,3,3)(3,3,3,3) or (0,0,8,4)(0,0,8,4) iterations at the four sizes, perform better than abrupt jumps such as (6,0,0,6)(6,0,0,6) or (3,0,0,9)(3,0,0,9). The same ablation shows that progressively increasing model size performs better than running large-to-small: larger models are more useful in later decoding iterations.

  9. Knowl 9 — Cache refresh recovers quality lost by KV reuse

    empirical result

    On ImageNet 256×256256\times256 with size-L models, the cache ablation compares uniform decoding and a four-stage upscaling schedule with three iterations at each of p=(8,4,2,1)p=(8,4,2,1). FLOP gain is measured relative to the uncached, unscheduled baseline. The page 8 ablation table reports:

    Baseline Baseline + cache Baseline + cache + refresh Scheduled Scheduled + cache Scheduled + cache + refresh
    FID ↓\downarrow 2.5 3.4 2.6 3.1 4.8 3.1
    FLOP gain 1.0 1.3 1.2 2.1 3.5 3.0

    Caching alone degrades FID, particularly when combined with model scheduling; refreshing the cache at transitions recovers that quality loss. Scheduled decoding with caching and refresh gives a 3.0-fold compute gain at FID 3.1, compared with FID 2.5 for the uncached baseline.

  10. Knowl 10 — The method retains tokenizer and artifact limitations

    limitation

    MaGNeTS can produce visible artifacts, including artifacts also found in MaskGIT++; the paper illustrates failure cases on page 16. Its effectiveness depends directly on the quality of the pretrained visual tokenizer, and improving tokenizers is outside the work's scope. The authors state that nested modeling and decode-time scaling do not require a particular model architecture or sampling scheme, but the KV-caching component requires discrete tokens.

Coverage note — Detailed sampling-temperature settings, the GFLOPs accounting implementation, and additional qualitative galleries are omitted as supplementary implementation or illustration material; the selective-guidance finding is included with the ImageNet result.

References

  1. 1.Berthelot, D., Autef, A., Lin, J., Yap, D. A., Zhai, S., Hu, S., Zheng, D., Talbott, W., and Gu, E. Tract: Denoising diffusion models with transitive closure time-distillation, 2023. URL https://arxiv.org/abs/2303.04248.
  2. 2.Beyer, L., Zhai, X., Royer, A., Markeeva, L., Anil, R., and Kolesnikov, A. Knowledge distillation: A good teacher is patient and consistent, 2022. URL https://arxiv.org/abs/2106.05237.
  3. 3.Brock, A. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018.
  4. 4.Cai, M., Yang, J., Gao, J., and Lee, Y. J. Matryoshka multimodal models. arXiv preprint arXiv:2405.17430, 2024a.
  5. 5.Cai, R., Muralidharan, S., Heinrich, G., Yin, H., Wang, Z., Kautz, J., and Molchanov, P. Flextron: Many-in-one flexible large language model. arXiv preprint arXiv:2406.10260, 2024b.
  6. 6.Carreira, J., Noland, E., Banki-Horvath, A., Hillier, C., and Zisserman, A. A short note about kinetics-600, 2018. URL https://arxiv.org/abs/1808.01340.
  7. 7.Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T. Maskgit: Masked generative image transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11315–11325, 2022.
  8. 8.Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.-H., Murphy, K., Freeman, W. T., Rubinstein, M., et al. Muse: Text-to-image generation via masked generative transformers. arXiv preprint arXiv:2301.00704, 2023.
  9. 9.Choi, J., Lee, J., Shin, C., Kim, S., Kim, H., and Yoon, S. Perception prioritized training of diffusion models, 2022. URL https://arxiv.org/abs/2204.00227.
  10. 10.Clark, A., Donahue, J., and Simonyan, K. Adversarial video generation on complex datasets. arXiv preprint arXiv:1907.06571, 2019.
  11. 11.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  12. 12.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding, 2019. URL https://arxiv.org/abs/1810.04805.
  13. 13.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34:8780–8794, 2021.
  14. 14.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In CVPR, pp. 12873–12883, 2021.
  15. 15.Esser, P., Kulal, S., Blattmann, A., Entezari, R., Muller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., Lacey, K., Goodwin, A., Marek, Y., and Rombach, R. Scaling rectified flow transformers for high-resolution image synthesis, 2024. URL https://arxiv.org/abs/2403.03206.
  16. 16.Feng, W., Yang, C., An, Z., Huang, L., Diao, B., Wang, F., and Xu, Y. Relational diffusion distillation for efficient image generation, 2024. URL https://arxiv.org/abs/2410.07679.
  17. 17.Gao, S., Zhou, P., Cheng, M.-M., and Yan, S. Masked diffusion transformer is a strong image synthesizer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 23164–23173, 2023.
  18. 18.Ge, S., Hayes, T., Yang, H., Yin, X., Pang, G., Jacobs, D., Huang, J.-B., and Parikh, D. Long video generation with time-agnostic vqgan and time-sensitive transformer, 2022. URL https://arxiv.org/abs/2204.03638.
  19. 19.Gu, S., Chen, D., Bao, J., Wen, F., Zhang, B., Chen, D., Yuan, L., and Guo, B. Vector quantized diffusion model for text-to-image synthesis, 2022. URL https://arxiv.org/abs/2111.14822.
  20. 20.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  21. 21.Ho, J. and Salimans, T. Classifier-free diffusion guidance, 2022. URL https://arxiv.org/abs/2207.12598.
  22. 22.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  23. 23.Ho, J., Saharia, C., Chan, W., Fleet, D. J., Norouzi, M., and Salimans, T. Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research, 23(47):1–33, 2022a.
  24. 24.Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J. Video diffusion models. Advances in Neural Information Processing Systems, 35:8633–8646, 2022b.
  25. 25.Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022.
  26. 26.Hoogeboom, E., Heek, J., and Salimans, T. simple diffusion: End-to-end diffusion for high resolution images. In International Conference on Machine Learning, pp. 13213–13232. PMLR, 2023.
  27. 27.Hoppe, T., Mehrjou, A., Bauer, S., Nielsen, D., and Dittadi, A. Diffusion models for video prediction and infilling. arXiv preprint arXiv:2206.07696, 2022.
  28. 28.Hu, W., Dou, Z.-Y., Li, L. H., Kamath, A., Peng, N., and Chang, K.-W. Matryoshka query transformer for large vision-language models. arXiv preprint arXiv:2405.19315, 2024.
  29. 29.Jabri, A., Fleet, D., and Chen, T. Scalable adaptive computation for iterative generation. arXiv preprint arXiv:2212.11972, 2022.
  30. 30.Jain, G., Hegde, N., Kusupati, A., Nagrani, A., Buch, S., Jain, P., Arnab, A., and Paul, S. Mixture of nested experts: Adaptive processing of visual tokens, 2024. URL https://arxiv.org/abs/2407.19985.
  31. 31.Kingma, D. P. and Gao, R. Understanding the diffusion objective as a weighted integral of elbos. arXiv preprint arXiv:2303.00848, 2, 2023.
  32. 32.Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al. Videopoet: A large language model for zero-shot video generation. ICML, 2024.
  33. 33.Kong, Z. and Ping, W. On fast sampling of diffusion probabilistic models, 2021. URL https://arxiv.org/abs/2106.00132.
  34. 34.Kudugunta, S., Kusupati, A., Dettmers, T., Chen, K., Dhillon, I., Tsvetkov, Y., Hajishirzi, H., Kakade, S., Farhadi, A., Jain, P., et al. Matformer: Nested transformer for elastic inference. arXiv preprint arXiv:2310.07707, 2023.
  35. 35.Kusupati, A., Bhatt, G., Rege, A., Wallingford, M., Sinha, A., Ramanujan, V., Howard-Snyder, W., Chen, K., Kakade, S., Jain, P., et al. Matryoshka representation learning. Advances in Neural Information Processing Systems, 35:30233–30249, 2022.
  36. 36.Le Moing, G., Ponce, J., and Schmid, C. Ccvs: Context-aware controllable video synthesis. Advances in Neural Information Processing Systems, 34:14042–14055, 2021.
  37. 37.Lee, H., Lee, H., Gye, S., and Kim, J. Beta sampling is all you need: Efficient image generation strategy for diffusion models using stepwise spectral analysis, 2024. URL https://arxiv.org/abs/2407.12173.
  38. 38.Lee, Y., Kim, J.-Y., Go, H., Jeong, M., Oh, S., and Choi, S. Multi-architecture multi-expert diffusion models, 2023. URL https://arxiv.org/abs/2306.04990.
  39. 39.Li, L., Li, H., Zheng, X., Wu, J., Xiao, X., Wang, R., Zheng, M., Pan, X., Chao, F., and Ji, R. Autodiffusion: Training-free optimization of time steps and architectures for automated diffusion model acceleration, 2023. URL https://arxiv.org/abs/2309.10438.
  40. 40.Li, T., Tian, Y., Li, H., Deng, M., and He, K. Autoregressive image generation without vector quantization. arXiv preprint arXiv:2406.11838, 2024a.
  41. 41.Li, T., Tian, Y., Li, H., Deng, M., and He, K. Autoregressive image generation without vector quantization, 2024b. URL https://arxiv.org/abs/2406.11838.
  42. 42.Liu, L., Ren, Y., Lin, Z., and Zhao, Z. Pseudo numerical methods for diffusion models on manifolds, 2022. URL https://arxiv.org/abs/2202.09778.
  43. 43.Lou, A., Meng, C., and Ermon, S. Discrete diffusion modeling by estimating the ratios of the data distribution. In Forty-first International Conference on Machine Learning.
  44. 44.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps, 2022. URL https://arxiv.org/abs/2206.00927.
  45. 45.Luc, P., Clark, A., Dieleman, S., Casas, D. d. L., Doron, Y., Cassirer, A., and Simonyan, K. Transformation-based adversarial video prediction on large-scale data. arXiv preprint arXiv:2003.04035, 2020.
  46. 46.Meng, C., Rombach, R., Gao, R., Kingma, D. P., Ermon, S., Ho, J., and Salimans, T. On distillation of guided diffusion models, 2023. URL https://arxiv.org/abs/2210.03142.
  47. 47.Mirzadeh, S.-I., Farajtabar, M., Li, A., Levine, N., Matsukawa, A., and Ghasemzadeh, H. Improved knowledge distillation via teacher assistant, 2019. URL https://arxiv.org/abs/1902.03393.
  48. 48.Nash, C., Menick, J., Dieleman, S., and Battaglia, P. W. Generating images with sparse representations. arXiv preprint arXiv:2103.03841, 2021.
  49. 49.Nash, C., Carreira, J., Walker, J., Barr, I., Jaegle, A., Malinowski, M., and Battaglia, P. Transframer: Arbitrary frame prediction with generative models. arXiv preprint arXiv:2203.09494, 2022.
  50. 50.Nichol, A. and Dhariwal, P. Improved denoising diffusion probabilistic models, 2021a. URL https://arxiv.org/abs/2102.09672.
  51. 51.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In International conference on machine learning, pp. 8162–8171. PMLR, 2021b.
  52. 52.Park, Y.-H., Kwon, M., Choi, J., Jo, J., and Uh, Y. Understanding the latent space of diffusion models through the lens of riemannian geometry, 2023. URL https://arxiv.org/abs/2307.12868.
  53. 53.Peebles, W. and Xie, S. Scalable diffusion models with transformers, 2023. URL https://arxiv.org/abs/2212.09748.
  54. 54.Razavi, A., Van den Oord, A., and Vinyals, O. Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems, 32, 2019.
  55. 55.Rippel, O., Gelbart, M., and Adams, R. Learning ordered representations with nested dropout. In International Conference on Machine Learning, pp. 1746–1754. PMLR, 2014.
  56. 56.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models, 2022. URL https://arxiv.org/abs/2112.10752.
  57. 57.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems, 35:36479–36494, 2022.
  58. 58.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models, 2022. URL https://arxiv.org/abs/2202.00512.
  59. 59.Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. Improved techniques for training gans, 2016. URL https://arxiv.org/abs/1606.03498.
  60. 60.Sauer, A., Schwarz, K., and Geiger, A. Stylegan-xl: Scaling stylegan to large diverse datasets, 2022. URL https://arxiv.org/abs/2202.00273.
  61. 61.Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022.
  62. 62.Skorokhodov, I., Tulyakov, S., and Elhoseiny, M. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 3626–3636, 2022.
  63. 63.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models, 2022. URL https://arxiv.org/abs/2010.02502.
  64. 64.Song, Y., Dhariwal, P., Chen, M., and Sutskever, I. Consistency models, 2023. URL https://arxiv.org/abs/2303.01469.
  65. 65.Soomro, K., Zamir, A. R., and Shah, M. Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012. URL https://arxiv.org/abs/1212.0402.
  66. 66.Stanton, S., Izmailov, P., Kirichenko, P., Alemi, A. A., and Wilson, A. G. Does knowledge distillation really work?, 2021. URL https://arxiv.org/abs/2106.05945.
  67. 67.Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L. Visual autoregressive modeling: Scalable image generation via next-scale prediction, 2024. URL https://arxiv.org/abs/2404.02905.
  68. 68.Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. Fvd: A new metric for video generation. 2019.
  69. 69.Van Den Oord, A., Vinyals, O., et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017.
  70. 70.Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D. Phenaki: Variable length video generation from open domain textual descriptions. In International Conference on Learning Representations, 2022.
  71. 71.Wang, Y., Ren, S., Lin, Z., Han, Y., Guo, H., Yang, Z., Zou, D., Feng, J., and Liu, X. Parallelized autoregressive visual generation, 2024. URL https://arxiv.org/abs/2412.15119.
  72. 72.Wang, Z., Wang, J., Liu, Z., and Qiu, Q. Binary latent diffusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 22576–22585, 2023.
  73. 73.Watson, D., Chan, W., Ho, J., and Norouzi, M. Learning fast samplers for diffusion models by differentiating through sample quality, 2022. URL https://arxiv.org/abs/2202.05830.
  74. 74.Weber, M., Yu, L., Yu, Q., Deng, X., Shen, X., Cremers, D., and Chen, L.-C. Maskbit: Embedding-free image generation via bit tokens, 2024. URL https://arxiv.org/abs/2409.16211.
  75. 75.Yan, J. N., Gu, J., and Rush, A. M. Diffusion models without attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8239–8249, 2024.
  76. 76.Yang, S., Chen, Y., Wang, L., Liu, S., and Chen, Y. Denoising diffusion step-aware models, 2024. URL https://arxiv.org/abs/2310.03337.
  77. 77.Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y. Vector-quantized image modeling with improved vqgan. arXiv preprint arXiv:2110.04627, 2021.
  78. 78.Yu, L., Cheng, Y., Sohn, K., Lezama, J., Zhang, H., Chang, H., Hauptmann, A. G., Yang, M.-H., Hao, Y., Essa, I., et al. Magvit: Masked generative video transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10459–10469, 2023a.
  79. 79.Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Gupta, A., Gu, X., Hauptmann, A. G., et al. Language model beats diffusion–tokenizer is key to visual generation. arXiv preprint arXiv:2310.05737, 2023b.
  80. 80.Yu, Q., He, J., Deng, X., Shen, X., and Chen, L.-C. Randomized autoregressive visual generation, 2024a. URL https://arxiv.org/abs/2411.00776.
  81. 81.Yu, Q., Weber, M., Deng, X., Shen, X., Cremers, D., and Chen, L.-C. An image is worth 32 tokens for reconstruction and generation, 2024b. URL https://arxiv.org/abs/2406.07550.
  82. 82.Yu, S., Tack, J., Mo, S., Kim, H., Kim, J., Ha, J.-W., and Shin, J. Generating videos with dynamics-aware implicit generative adversarial networks. arXiv preprint arXiv:2202.10571, 2022.
  83. 83.Yu, S., Kwak, S., Jang, H., Jeong, J., Huang, J., Shin, J., and Xie, S. Representation alignment for generation: Training diffusion transformers is easier than you think, 2024c. URL https://arxiv.org/abs/2410.06940.
  84. 84.Yuan, Z., Zhang, H., Lu, P., Ning, X., Zhang, L., Zhao, T., Yan, S., Dai, G., and Wang, Y. Ditfastattn: Attention compression for diffusion transformer models. arXiv preprint arXiv:2406.08552, 2024.
  85. 85.Zheng, C., Vuong, L. T., Cai, J., and Phung, D. Movq: Modulating quantized vectors for high-fidelity image generation, 2022. URL https://arxiv.org/abs/2209.09002.
  86. 86.Zheng, H., Nie, W., Vahdat, A., and Anandkumar, A. Fast training of diffusion models with masked transformers. arXiv preprint arXiv:2306.09305, 2023.

Citation

MLA
Goyal, S., et al. “Masked Generative Nested Transformers with Decode Time Scaling”. arXiv, 2025, http://arxiv.org/abs/2502.00382v1.
APA
Goyal, S., Tula, D., Jain, G., Shenoy, P., Jain, P., & Paul, S. (2025). Masked Generative Nested Transformers with Decode Time Scaling. arXiv. http://arxiv.org/abs/2502.00382v1
Chicago
Goyal, S., D. Tula, G. Jain, P. Shenoy, P. Jain, and S. Paul. 2025. “Masked Generative Nested Transformers with Decode Time Scaling”. arXiv. http://arxiv.org/abs/2502.00382v1.
Harvard
Goyal, S. et al. (2025) “Masked Generative Nested Transformers with Decode Time Scaling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.00382v1.
Vancouver
1. Goyal S, Tula D, Jain G, Shenoy P, Jain P, Paul S (2025) Masked Generative Nested Transformers with Decode Time Scaling. arXiv

BibTeX

@article{goyal2025masked,
  title = {Masked Generative Nested Transformers with Decode Time Scaling},
  author = {Goyal, Sahil and Tula, Debapriya and Jain, Gagan and Shenoy, Pradeep and Jain, Prateek and Paul, Sujoy},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.00382v1},
  eprint = {2502.00382}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/