All-atom Diffusion Transformers: Unified generative modelling of molecules and materials

Chaitanya K. JoshiXiang FuYi-Lun LiaoVahe GharakhanyanBenjamin Kurt MillerAnuroop SriramZachary W. Ulissi

article2025ICML103 citations

Presents a unified latent diffusion framework that generates both periodic crystals and non-periodic molecules with a single standard Transformer architecture, achieving an order-of-magnitude faster sampling and state-of-the-art validity across chemical benchmarks.

Listen

Generative artificial intelligence has become a critical tool for discovering new drugs and advanced materials, yet existing approaches remain heavily fragmented. Although all physical matter obeys the same atomic principles, current generative models rely on specialized, complex architectures tailored separately for non-periodic systems (like small molecules) or periodic systems (like crystal lattices). This separation limits cross-domain knowledge transfer, introduces significant computational overhead, and slows down discovery pipelines.

The article introduces and evaluates the All-atom Diffusion Transformer (ADiT), a unified framework designed to jointly generate both periodic materials and non-periodic molecules using a single, shared architecture.

The researchers developed a two-stage generative process. In the first stage, a standard autoencoder maps diverse 3D atomic structures—representing atom types, 3D coordinates, fractional coordinates, and unit cell parameters—into a shared continuous latent space. In the second stage, a Diffusion Transformer learns to generate new latent embeddings that decode into valid chemical structures, guided by a simple class label indicating whether the output should be periodic or non-periodic. The approach relies on standard Transformers trained with data augmentations rather than complex, computationally expensive rotation-invariant neural networks. Credibility was established across major public datasets, including QM9 (130,000 small molecules), MP20 (over 45,000 crystal structures), GEOM-DRUGS (430,000 larger molecules), and QMOF (14,000 metal-organic frameworks), using standardized quantum chemical checks and physics-based validation suites.

The evaluation revealed several key findings:

  1. Joint training of molecules and materials in a single shared model outperforms training on individual domains alone, confirming effective transfer learning across distinct chemical systems.
  2. In crystal generation benchmarks on MP20, ADiT achieved a stable, unique, and novel (S.U.N.) discovery rate of 5.3% to 6.5%, delivering an approximate 25% improvement over prior state-of-the-art models.
  3. In molecular generation on QM9 and GEOM-DRUGS, ADiT matched or surpassed specialized baselines, generating physically realistic 3D geometries that passed stringent structural sanity checks.
  4. ADiT achieved dramatic computational speedups, generating 10,000 sample structures in under 20 minutes on a single GPU—an order of magnitude faster than baseline models that required up to 2.5 hours on equivalent hardware.
  5. Generative performance predictably scaled as model size increased from 32 million to 450 million parameters, demonstrating foundation-model scaling behavior in chemical generation.

These findings indicate that specialized geometric architectures are not required to generate valid atomic structures at scale. Standard, widely adopted Transformer architectures can handle complex atomic data efficiently when paired with latent diffusion. This shift lowers the computational cost, training time, and technical complexity required to develop generative chemistry pipelines, making high-throughput material and drug discovery substantially more accessible.

Organizations developing computational chemistry infrastructure should consider adopting unified latent diffusion models over domain-specific pipelines to reduce technical debt and take advantage of shared data representations. For immediate next steps, teams should scale training to larger cross-domain datasets (such as ZINC, Alexandria, and the Protein Data Bank) and integrate conditional steering mechanisms—such as conditioning on target physical properties, binding affinities, or molecular scaffolds—to enable practical inverse design applications.

The findings are supported by consistent validation metrics across multiple random seeds and quantum chemical calculations. However, readers should note that the primary models were trained on relatively small benchmark datasets containing smaller atomic structures (up to hundreds of atoms). Full scaling behavior on macro-scale biomolecules and metal-organic frameworks containing thousands of atoms remains to be validated, and the current models operate unconditionally rather than targeting specific desired chemical properties.

No sufficiently relevant recommendations were found.

Cover for All-atom Diffusion Transformers: Unified generative modelling of molecules and materials

Abstract

Diffusion models are the standard toolkit for generative modelling of 3D atomic systems. However, for different types of atomic systems – such as molecules and materials – the generative processes are usually highly specific to the target system despite the underlying physics being the same. We introduce the All-atom Diffusion Transformer (ADiT), a unified latent diffusion framework for jointly generating both periodic materials and non-periodic molecular systems using the same model: (1) An autoencoder maps a unified, all-atom representations of molecules and materials to a shared latent embedding space; and (2) A diffusion model is trained to generate new latent embeddings that the autoencoder can decode to sample new molecules or materials. Experiments on MP20, QM9 and GEOM-DRUGS datasets demonstrate that jointly trained ADiT generates realistic and valid molecules as well as materials, obtaining state-of-the-art results on par with molecule and crystal-specific models. ADiT uses standard Transformers with minimal inductive biases for both the autoencoder and diffusion model, resulting in significant speedups during training and inference compared to equivariant diffusion models. Scaling ADiT up to half a billion parameters predictably improves performance, representing a step towards broadly generalizable foundation models for generative chemistry. Open source code: https://github.com/facebookresearch/all-atom-diffusion-transformer

Table of Contents

  • 1. Introduction
  • 2. All-atom Diffusion Transformers
  • 2.1. Stage 1: Autoencoder for reconstruction
  • 2.2. Stage 2: Latent diffusion generative model
  • 3. Experimental Setup
  • 4. Results
  • 5. Discussions
  • Impact Statement
  • References
  • A. Related Work
  • B. Evaluation Metrics
  • C. Additional Results
  • D. Ablation Study
  • E. Visualizations

Knowls

  1. Knowl 1 — ADiT unifies molecular and crystal generation through a shared latent space

    model/method

    All-atom Diffusion Transformers (ADiT) generate non-periodic molecules and periodic crystals with one two-stage model. First, a variational autoencoder (VAE) encodes atom-wise descriptions of both domains into a shared latent space and learns to reconstruct the atomic systems. Second, a class-conditional Diffusion Transformer (DiT) learns to generate latent representations from noise; the VAE decoder turns generated latents into molecules or crystals. A periodic/non-periodic class label selects the appropriate decoding outputs and supports domain-specific generation while the model shares its main learned representation and parameters.

  2. Knowl 2 — A unified atom-wise VAE represents molecules and periodic crystals

    model/method

    An atomic system with NN atoms is represented by atom types aia_i, Cartesian coordinates xi∈R3x_i\in\mathbb{R}^3, fractional coordinates fi∈[0,1)3f_i\in[0,1)^3, and a lattice matrix L∈R3×3L\in\mathbb{R}^{3\times3}. For a crystal, fractional coordinates satisfy fi=L−1xif_i=L^{-1}x_i; the lattice describes the repeating unit cell, whose parameters are put in Niggli-reduced form. For a molecule, lattice and fractional-coordinate attributes are null. Cartesian coordinates are measured in nanometers.

    The encoder embeds atom types and projects the coordinate attributes, processes the per-atom features with a standard Transformer, and predicts a Gaussian mean and standard deviation for each atom's latent vector. The decoder processes those vectors with a standard Transformer and predicts atom types, Cartesian and fractional coordinates, and lattice parameters; lattice parameters are predicted from the mean-pooled atom features. The reported baseline uses latent dimension d=8d=8 and Transformer hidden dimension dmodel=512d_{\text{model}}=512 for the VAE.

    Reconstruction uses atom-type cross-entropy, coordinate or fractional-coordinate mean-squared error (MSE), and MSE on lattice side lengths and angles. Cartesian-coordinate MSE is evaluated after separately centering the input and reconstructed coordinates. Crystal side lengths are normalized by the cube root of the atom count, and lattice angles are converted to radians. The weighted reconstruction loss is

    Lrec=λALA+λXLX+λFLF+λLlLLl+λLaLLa.L_{\mathrm{rec}}=\lambda_A L_A+\lambda_X L_X+\lambda_F L_F+\lambda_{L_l}L_{L_l}+\lambda_{L_a}L_{L_a}.

    Here LA,LX,LF,LLl,LLaL_A,L_X,L_F,L_{L_l},L_{L_a} denote atom-type, Cartesian-coordinate, fractional-coordinate, lattice-length, and lattice-angle losses, respectively. For crystals, the weights (λA,λX,λF,λLl,λLa)(\lambda_A,\lambda_X,\lambda_F,\lambda_{L_l},\lambda_{L_a}) are (1,0,10,1,10)(1,0,10,1,10); for molecules they are (1,10,0,0,0)(1,10,0,0,0), so each domain is trained against its applicable attributes. The VAE also uses a per-channel KL penalty toward a standard normal latent distribution, a bottleneck, and denoising augmentation that masks 10% of atom types and perturbs coordinates with Gaussian noise of standard deviation 0.10.1. Non-equivariant VAE training additionally uses random rotations and translations. At generation time, molecular decoding uses atom types and Cartesian coordinates, while crystal decoding uses atom types, fractional coordinates, and lattice parameters.

  3. Knowl 3 — Latent flow matching trains the DiT to generate VAE embeddings

    equation

    After training the VAE, ADiT trains a denoiser on its frozen encoder's atom-wise latents. For a clean latent set Z(1)Z^{(1)} and a zero-centered Gaussian latent set Z(0)Z^{(0)}, it samples time tt and linearly interpolates between them:

    Z(t)=(1−t)Z(0)+tZ(1),ut(Z(t)∣Z(1))=Z(1)−Z(t)1−t.Z^{(t)}=(1-t)Z^{(0)}+tZ^{(1)},\qquad u_t(Z^{(t)}\mid Z^{(1)})=\frac{Z^{(1)}-Z^{(t)}}{1-t}.

    Here each Z(t)Z^{(t)} is a set of NN vectors in Rd\mathbb{R}^d, tt is sampled uniformly between 0 and 1, and νt\nu_t is the target vector field toward the clean latent. The DiT receives Z(t)Z^{(t)}, time, and a periodic/non-periodic label, and predicts a clean latent set Z′(1)Z'^{(1)}. Training minimizes the mean squared error between the predicted and target vector fields, equivalently

    Lfm=1N(1−t)2∑i=1N∥zi(1)−zi′(1)∥22,L_{\mathrm{fm}}=\frac{1}{N(1-t)^2}\sum_{i=1}^N\left\|z_i^{(1)}-z_i'^{(1)}\right\|_2^2,

    where zi(1)z_i^{(1)} and zi′(1)z_i'^{(1)} are the clean and predicted clean latent vectors for atom ii. In practice, the time value is kept at least 0.010.01 and clipped at 0.90.9 to limit numerical instability.

  4. Knowl 4 — Classifier-free guidance and Euler integration produce decoded samples

    algorithm

    ADiT sampling takes a desired class cc (periodic crystal or non-periodic molecule), an integration-step count TT, and classifier-free guidance scale γ\gamma. It starts from Gaussian atom-wise latents and uses the DiT to predict clean latents both with class conditioning and with a null label ϕ\phi. The guidance-combined prediction defines the vector field used for Euler integration; the final latent set is decoded by the VAE. The DiT uses adaptive layer normalization with zero initialization for conditioning, drops class labels with probability 0.1 during training, and uses self-conditioning by concatenating a previous prediction with 50% dropout. Training inputs are randomly rotated and translated before encoding.

    Input: Class label c, integration steps T, guidance scale gamma
    Output: Generated atomic system
    Set Z(0) to N atom-wise samples from N(0, I_d)
    Set delta_t = 1/T
    For t over T Euler steps from 0 toward 1:
        Z_cond = DiT(Z(t), t, c)
        Z_uncond = DiT(Z(t), t, phi)
        Z_pred = (1 - gamma) * Z_uncond + gamma * Z_cond
        v = (Z_pred - Z(t)) / (1 - t)
        Z(t + delta_t) = Z(t) + delta_t * v
    Decode the final latent set with the VAE decoder
    Return the decoded system

    The paper reports that T=500T=500 or 10001000 and γ=1.0\gamma=1.0 or 2.02.0 generally work well for both domains; the best setting can differ between molecules and crystals.

  5. Knowl 5 — Joint ADiT improves MP20 crystal validity and stability

    data/table

    The MP20 evaluation compares crystal-generation models on 10,000 samples using structural, compositional, and overall validity, metastability, stability, and combined stability/uniqueness/novelty rates. Stability is based on DFT energy above hull below 0.00.0 eV/atom; metastability uses below 0.10.1 eV/atom. Joint training means training on both QM9 molecules and MP20 crystals. The jointly trained ADiT has the highest overall validity among the ADiT variants and improves on MP20-only ADiT in composition validity, stable rate, and S.U.N. rate. Baseline DFT values marked with †\dagger were replicated using ADiT's DFT setup; MatterGen-MP values marked with ∗* use 1,024 samples.

    Model Structure valid (%) Composition valid (%) Overall valid (%) Metastable (%) Stable (%) M.S.U.N. (%) S.U.N. (%)
    CDVAE 100.00 86.70 – – 1.6 – –
    DiffCSP 100.00 83.25 – – 5.0 – 3.3
    UniMat 97.2 89.4 – – – – –
    FlowMM 96.85 83.19 80.30 30.6†^{\dagger} 4.6†^{\dagger} 22.5†^{\dagger} 2.8†^{\dagger}
    FlowLLM 99.94 90.84 90.81 66.9†^{\dagger} 13.9†^{\dagger} 26.3†^{\dagger} 4.7†^{\dagger}
    MatterGen-MP – – – 78∗^* 13∗^* 21∗^* –
    MP20-only ADiT 99.58 90.46 90.13 81.6 14.1 25.91 4.7
    Jointly trained ADiT 99.74 92.14 91.92 81.0 15.4 28.2 5.3
  6. Knowl 6 — Joint ADiT generates valid QM9 molecules with realistic 3D geometry

    data/table

    Molecule evaluation uses 10,000 QM9 samples. Validity is the fraction convertible to canonical SMILES with RDKit, and uniqueness is measured among valid molecules. An asterisk marks models that explicitly generate hydrogen atoms. Jointly trained ADiT reaches 97.43% validity and 96.92% uniqueness without explicit hydrogen generation; its explicit-hydrogen variant reaches 94.45% and 97.82%. PoseBusters checks assess physical plausibility; ADiT's pass rates are close to the compared models and are particularly high for ring flatness, double-bond flatness, internal energy, and steric clashes.

    Model Validity (%) Unique (%)
    Equivariant Diffusion 97.50 96.71
    Equivariant Diffusion∗^* 91.90 98.69
    GeoLDM∗^* 93.80 98.82
    Symphony∗^* 83.50 97.98
    QM9-only ADiT 96.02 97.76
    QM9-only ADiT∗^* 92.19 97.90
    Jointly trained ADiT 97.43 96.92
    Jointly trained ADiT∗^* 94.45 97.82
    PoseBusters check (% pass) Symphony Equivariant Diffusion ADiT
    Atoms connected 99.92 99.88 99.70
    Bond angles 99.56 99.98 99.85
    Bond lengths 98.72 100.00 99.41
    Ring flat 100.00 100.00 100.00
    Double bond flat 99.07 98.58 99.98
    Internal energy 95.65 94.88 95.86
    No steric clash 98.16 99.79 99.79
  7. Knowl 7 — ADiT remains competitive on large GEOM-DRUGS molecules

    empirical result

    On GEOM-DRUGS, which contains 430,000 molecules of up to 180 atoms, ADiT was evaluated on 10,000 generated molecules and compared with equivariant diffusion and flow-matching models. ADiT achieved 95.3% validity and 100.0% uniqueness, compared with 94.6% and 100.0% for EQGAT-diff and 93.9% and 100.0% for SemlaFlow. ADiT's PoseBusters-valid rate was 85.3%, below SemlaFlow's 87.5% but above EQGAT-diff's 59.7%; its individual checks included 93.0% connected atoms, 92.5% bond lengths, 95.4% ring flatness, and 95.3% double-bond flatness. The result shows that the standard-Transformer ADiT can generate large molecules competitively without explicitly predicting atomic bonds. The paper also reports favorable sampling-time scaling against SemlaFlow on a single A100 GPU, without tabulating exact runtime values.

  8. Knowl 8 — Transformer denoisers improve with scale and sample faster than equivariant baselines

    empirical result

    Scaling the ADiT denoiser from DiT-S (32M parameters) through DiT-B (130M) to DiT-L (450M) consistently reduced training loss and improved crystal and molecule validity in the reported experiments, despite training on about 130,000 QM9 and MP20 samples. At epoch 2,000, the reported Pearson correlations between log parameter count and training loss, crystal validity, and molecule validity were −1.00-1.00, 0.910.91, and 0.940.94, respectively; the corresponding Spearman correlations were −1.00-1.00, 1.001.00, and 1.001.00.

    For inference, the paper compares 10,000 generated samples on a single V100 GPU and reports better scaling with integration-step count than equivariant FlowMM for crystals and GeoLDM for molecules. The abstract reports that ADiT can generate 10,000 samples in under 20 minutes on one V100 GPU, while the compared baselines can take up to 2.5 hours. These measurements support the paper's claim that standard Transformer denoisers enable more practical model scaling and repeated diffusion inference.

  9. Knowl 9 — Larger ADiTs improve crystal stability but can reduce novelty

    data/table

    For 10,000 generated crystals, increasing model size generally raised stability and stable-and-unique rates, but the strict stable, unique, and novel (S.U.N.) rate fell for the larger models. The authors attribute this pattern to greater memorization of the 27,000-crystal MP20 training set. Even the 32M-parameter MP20-only model achieved a 6.5% S.U.N. rate, compared with the reported 2.8% for FlowMM and 4.7% for FlowLLM. Rates are percentages; SS denotes stable, S.U.S.U. stable and unique, and S.U.N.S.U.N. stable, unique, and novel. The metastable columns use the DFT threshold Ehull<0.1E_{\mathrm{hull}}<0.1 eV/atom, while the stable columns use Ehull<0.0E_{\mathrm{hull}}<0.0 eV/atom.

    Model S (%) S.U. (%) S.U.N. (%) M.S. (%) M.S.U. (%) M.S.U.N. (%)
    MP20-only ADiT-S (32M) 12.8 11.8 6.5 71.1 64.9 38.1
    MP20-only ADiT-B (130M) 14.1 12.5 4.7 81.6 67.3 25.9
    Joint ADiT-S (32M) 12.6 11.4 6.0 71.9 64.7 37.7
    Joint ADiT-B (130M) 15.4 13.4 5.3 81.0 70.2 28.2
    Joint ADiT-L (450M) 15.5 13.5 5.0 82.5 70.9 27.9
  10. Knowl 10 — Autoencoder ablations favor standard Transformers for latent generation

    empirical result

    The VAE and latent-diffusion ablations compare standard Transformer and Equiformer-V2 autoencoders, latent dimensions, KL weights, and denoiser sizes. Although reconstruction match rates vary by configuration—for example, on MP20 at latent dimension 8, Equiformer-V2 has 88.90% match versus 84.50% for the Transformer—the paper reports that Transformer latents generally produced stronger crystal validity when used for latent diffusion. For the joint Transformer VAE with latent dimension 8 and KL weight 10−510^{-5}, reconstruction match rates were 88.60% for crystals and 97.00% for molecules, with RMSDs of 0.0239 Å and 0.0399 Å, respectively. With that configuration and DiT-B, generated overall crystal validity was 91.92% and QM9 molecule validity was 97.43%. The ablations also show that larger latent dimensions and lower KL weights often improve reconstruction, while the best generation setting is not uniformly determined by reconstruction metrics alone.

  11. Knowl 11 — ADiT provides an initial joint-generation result for metal-organic frameworks

    empirical result

    The authors extended ADiT to QMOF metal-organic frameworks (MOFs), training on 14,000 examples of up to 150 atoms, either with QMOF alone or jointly with QM9 and MP20. Among 1,000 generated MOFs assessed with 15 MOFChecker sanity tests, the QMOF-only model passed all tests for 15.7% of samples, while the joint model passed all tests for 10.2%. The joint run had not fully converged, so the lower MOF validity is not presented as a settled limit. It nevertheless generated crystals and molecules at about 91% and 95% validity, respectively, while also generating MOFs. The authors note that their MOF results were evaluated without DFT relaxation and are not directly comparable to specialized MOF models evaluated after relaxation.

  12. Knowl 12 — ADiT has not yet established performance on very large systems or conditional design

    limitation

    The paper identifies three limitations. First, its training datasets are relatively small, which may constrain generalization. Second, although experiments include molecules and crystals with up to hundreds of atoms and an initial MOF-generation study, the approach has not been fully validated on much larger systems such as MOFs or biomolecules containing thousands of atoms. Third, the reported models perform unconditional generation; conditioning on experimental properties, motif scaffolds, or molecular infilling remains future work. These limits qualify the paper's broader aim of developing a general-purpose generative chemistry model.

Coverage note — Omitted the PCA visualizations of latent embeddings and example structure figures because they provide qualitative illustration rather than additional quantitative or methodological contributions.

References

  1. 1.Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., Ronneberger, O., Willmore, L., Ballard, A. J., Bambrick, J., et al. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature, 2024.
  2. 2.Axelrod, S. and Gomez-Bombarelli, R. Geom, energy-annotated molecular conformations for property prediction and molecular generation. Scientific Data, 2022.
  3. 3.Batatia, I., Benner, P., Chiang, Y., Elena, A. M., Kovács, D. P., Riebesell, J., Advincula, X. R., Asta, M., Avaylon, M., Baldwin, W. J., et al. A foundation model for atomistic materials chemistry. arXiv preprint arXiv:2401.00096, 2023.
  4. 4.Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al. Improving image generation with better captions, 2023.
  5. 5.Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., et al. Video generation models as world simulators. 2024.
  6. 6.Buttenschoen, M., Morris, G. M., and Deane, C. M. Posebusters: Ai-based docking methods fail to generate physically valid poses or generalise to novel sequences. Chemical Science, 2024.
  7. 7.Buttenschoen, M., Ziv, Y., Morris, G. M., and Deane, C. An evaluation of unconditional 3d molecular generation methods. In ICLR Workshop on Generative and Experimental Perspectives for Biomolecular Design, 2025.
  8. 8.Campbell, A., Yim, J., Barzilay, R., Rainforth, T., and Jaakkola, T. Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design. In Forty-first International Conference on Machine Learning, 2024.
  9. 9.Chu, A. E., Kim, J., Cheng, L., El Nesr, G., Xu, M., Shuai, R. W., and Huang, P.-S. An all-atom protein generative model. Proceedings of the National Academy of Sciences, 2024.
  10. 10.Corso, G., Jing, B., Barzilay, R., Jaakkola, T., et al. Diffdock: Diffusion steps, twists, and turns for molecular docking. In International Conference on Learning Representations, 2023.
  11. 11.Dai, X., Hou, J., Ma, C.-Y., Tsai, S., Wang, J., Wang, R., Zhang, P., Vandenhende, S., Wang, X., Dubey, A., et al. Emu: Enhancing image generation models using photogenic needles in a haystack. arXiv preprint arXiv:2309.15807, 2023.
  12. 12.Daigavane, A., Kim, S. E., Geiger, M., and Smidt, T. Symphony: Symmetry-equivariant point-centered spherical harmonics for 3d molecule generation. In ICLR, 2024.
  13. 13.Davies, D. W., Butler, K. T., Jackson, A. J., Skelton, J. M., Morita, K., and Walsh, A. Smact: Semiconducting materials by analogy and chemical theory. Journal of Open Source Software, 2019.
  14. 14.Deng, B., Zhong, P., Jun, K., Riebesell, J., Han, K., Bartel, C. J., and Ceder, G. Chgnet as a pretrained universal neural network potential for charge-informed atomistic modelling. Nature Machine Intelligence, 2023.
  15. 15.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 2021.
  16. 16.Duval, A., Mathis, S. V., Joshi, C. K., Schmidt, V., Miret, S., Malliaros, F. D., Cohen, T., Lio, P., Bengio, Y., and Bronstein, M. A hitchhiker’s guide to geometric gnns for 3d atomic systems. arXiv preprint arXiv:2312.07511, 2023.
  17. 17.Esser, P., Kulal, S., Blattmann, A., Entezari, R., Muller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al. Scaling rectified flow transformers for high-resolution image synthesis. In International Conference on Machine Learning, 2024.
  18. 18.Flam-Shepherd, D. and Aspuru-Guzik, A. Language models can generate molecules, materials, and protein binding sites directly in three dimensions as xyz, cif, and pdb files. arXiv preprint arXiv:2305.05708, 2023.
  19. 19.Fu, X., Xie, T., Rosen, A. S., Jaakkola, T. S., and Smith, J. A. MOFDiff: Coarse-grained diffusion for metal-organic framework design. In ICLR, 2024.
  20. 20.Gao, R., Hoogeboom, E., Heek, J., Bortoli, V. D., Murphy, K. P., and Salimans, T. Diffusion meets flow matching: Two sides of the same coin. 2024. URL https://diffusionflow.github.io/.
  21. 21.Grosse-Kunstleve, R. W., Sauter, N. K., and Adams, P. D. Numerically stable algorithms for the computation of reduced unit cells. Acta Crystallographica Section A: Foundations of Crystallography, 2004.
  22. 22.Gruver, N., Sriram, A., Madotto, A., Wilson, A. G., Zitnick, C. L., and Ulissi, Z. W. Fine-tuned language models generate stable inorganic materials as text. In The Twelfth International Conference on Learning Representations, 2024.
  23. 23.Harris, C., Didi, K., Jamasb, A. R., Joshi, C. K., Mathis, S. V., Lio, P., and Blundell, T. Benchmarking generated poses: How rational is structure-based drug design with generative models? arXiv preprint arXiv:2308.07413, 2023.
  24. 24.Ho, J. and Salimans, T. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  25. 25.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 2020.
  26. 26.Hoogeboom, E., Satorras, V. G., Vignac, C., and Welling, M. Equivariant diffusion for molecule generation in 3d. In International conference on machine learning, pp. 8867–8887. PMLR, 2022.
  27. 27.Ingraham, J. B., Baranov, M., Costello, Z., Barber, K. W., Wang, W., Ismail, A., Frappier, V., Lord, D. M., Ng-Thow-Hing, C., Van Vlack, E. R., et al. Illuminating protein space with a programmable generative model. Nature, 2023.
  28. 28.Irwin, R., Tibo, A., Janet, J. P., and Olsson, S. Semlaflow–efficient 3d molecular generation with latent attention and equivariant flow matching. In The 28th International Conference on Artificial Intelligence and Statistics, 2025.
  29. 29.Jablonka, K. M. MOFChecker v0.9.6, 2023. URL https://github.com/kjappelbaum/mofchecker.
  30. 30.Jain, A., Ong, S. P., Hautier, G., Chen, W., Richards, W. D., Dacek, S., Cholia, S., Gunter, D., Skinner, D., Ceder, G., et al. Commentary: The materials project: A materials genome approach to accelerating materials innovation. APL materials, 2013.
  31. 31.Jiao, R., Huang, W., Lin, P., Han, J., Chen, P., Lu, Y., and Liu, Y. Crystal structure prediction by joint equivariant diffusion. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  32. 32.Joshi, C. Transformers are graph neural networks. The Gradient, 2020.
  33. 33.Joshi, C. K., Bodnar, C., Mathis, S. V., Cohen, T., and Lio, P. On the expressive power of geometric graph neural networks. In International conference on machine learning, 2023.
  34. 34.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al. Highly accurate protein structure prediction with alphafold. Nature, 2021.
  35. 35.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. In International Conference on Learning Representations, 2014.
  36. 36.Le, T., Cremer, J., Noe, F., Clevert, D.-A., and Schütt, K. T. Navigating the design space of equivariant diffusion-based generative models for de novo 3d molecule generation. In The Twelfth International Conference on Learning Representations, 2024.
  37. 37.Liao, Y.-L., Wood, B. M., Das, A., and Smidt, T. Equiformerv2: Improved equivariant transformer for scaling to higher-degree representations. In International Conference on Learning Representations, 2024.
  38. 38.Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In International Conference on Learning Representations, 2023.
  39. 39.Ma, N., Goldstein, M., Albergo, M. S., Boffi, N. M., Vanden-Eijnden, E., and Xie, S. Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers. In ECCV, 2024.
  40. 40.Martinkus, K., Ludwiczak, J., Liang, W.-C., Lafrance-Vanasse, J., Hotzel, I., Rajpal, A., et al. Abdiffuser: full-atom generation of in-vitro functioning antibodies. Advances in Neural Information Processing Systems, 2024.
  41. 41.Miller, B. K., Chen, R. T., Sriram, A., and Wood, B. M. Flowmm: Generating materials with riemannian flow matching. In Forty-first International Conference on Machine Learning, 2024.
  42. 42.O Pinheiro, P. O., Rackers, J., Kleinhenz, J., Maser, M., Mahmood, O., Watkins, A., Ra, S., Sresht, V., and Saremi, S. 3d molecule generation by denoising voxel grids. Advances in Neural Information Processing Systems, 36: 69077–69097, 2023.
  43. 43.Ong, S. P., Richards, W. D., Jain, A., Hautier, G., Kocher, M., Cholia, S., Gunter, D., Chevrier, V. L., Persson, K. A., and Ceder, G. Python materials genomics (pymatgen): A robust, open-source python library for materials analysis. Computational Materials Science, 2013.
  44. 44.Peebles, W. and Xie, S. Scalable diffusion models with transformers. In International Conference on Computer Vision, 2023.
  45. 45.Riebesell, J., Goodall, R. E., Jain, A., Benner, P., Persson, K. A., and Lee, A. A. Matbench discovery–an evaluation framework for machine learning crystal stability prediction. arXiv preprint arXiv:2308.14920, 2023.
  46. 46.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In CVPR, 2022.
  47. 47.Rosen, A. S., Iyer, S. M., Ray, D., Yao, Z., Aspuru-Guzik, A., Gagliardi, L., Notestein, J. M., and Snurr, R. Q. Machine learning the quantum-chemical properties of metal–organic frameworks for accelerated materials discovery. Matter, 2021.
  48. 48.Satorras, V. G., Hoogeboom, E., and Welling, M. E (n) equivariant graph neural networks. In International conference on machine learning. PMLR, 2021.
  49. 49.Schneuing, A., Harris, C., Du, Y., Didi, K., Jamasb, A., Igashov, I., et al. Structure-based drug design with equivariant diffusion models. Nature Computational Science, 2024.
  50. 50.Shoghi, N., Kolluru, A., Kitchin, J. R., Ulissi, Z. W., Zitnick, C. L., and Wood, B. M. From molecules to materials: Pre-training large generalizable models for atomic property prediction. In ICLR, 2024.
  51. 51.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning, 2015.
  52. 52.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems, 2019.
  53. 53.Sriram, A., Miller, B. K., Chen, R. T. Q., and Wood, B. M. Flowllm: Flow matching for material generation with large language models as base distributions. In NeurIPS, 2024.
  54. 54.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  55. 55.Vahdat, A., Kreis, K., and Kautz, J. Score-based generative modeling in latent space. Advances in neural information processing systems, 2021.
  56. 56.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 2017.
  57. 57.Vignac, C., Osman, N., Toni, L., and Frossard, P. Midi: Mixed graph and 3d denoising diffusion for molecule generation. In ECML PKDD, 2023.
  58. 58.Wang, Y., Elhag, A. A., Jaitly, N., Susskind, J. M., and Bautista, M. A. Swallowing the bitter pill: Simplified scalable conformer generation. In International conference on machine learning, 2024.
  59. 59.Watson, J. L., Juergens, D., Bennett, N. R., Trippe, B. L., Yim, J., Eisenach, H. E., Ahern, W., Borst, A. J., Ragotte, R. J., Milles, L. F., et al. De novo design of protein structure and function with rfdiffusion. Nature, 2023.
  60. 60.Wu, Z., Ramsundar, B., Feinberg, E. N., Gomes, J., Geniesse, C., Pappu, A. S., Leswing, K., and Pande, V. Moleculenet: a benchmark for molecular machine learning. Chemical science, 2018.
  61. 61.Xie, T., Fu, X., Ganea, O.-E., Barzilay, R., and Jaakkola, T. S. Crystal diffusion variational autoencoder for periodic material generation. In International Conference on Learning Representations, 2022.
  62. 62.Xu, M., Powers, A. S., Dror, R. O., Ermon, S., and Leskovec, J. Geometric latent diffusion models for 3d molecule generation. In International Conference on Machine Learning, 2023.
  63. 63.Yang, S., Cho, K., Merchant, A., Abbeel, P., Schuurmans, D., Mordatch, I., and Cubuk, E. D. Scalable diffusion for materials generation. In The Twelfth International Conference on Learning Representations, 2024.
  64. 64.Yim, J., Campbell, A., Foong, A. Y., Gastegger, M., Jimenez-Luna, J., Lewis, S., Satorras, V. G., Veeling, B. S., Barzilay, R., Jaakkola, T., et al. Fast protein backbone generation with se (3) flow matching. arXiv preprint arXiv:2310.05297, 2023a.
  65. 65.Yim, J., Trippe, B. L., De Bortoli, V., Mathieu, E., Doucet, A., Barzilay, R., and Jaakkola, T. Se (3) diffusion model with application to protein backbone generation. In International Conference on Machine Learning, pp. 40001–40039. PMLR, 2023b.
  66. 66.Zeni, C., Pinsler, R., Zugner, D., Fowler, A., Horton, M., Fu, X., et al. Mattergen: a generative model for inorganic materials design. Nature, 2025.
  67. 67.Zhang, L., Rao, A., and Agrawala, M. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023.

Citation

MLA
Joshi, C. K., et al. “All-atom Diffusion Transformers: Unified Generative Modelling of Molecules and Materials”. arXiv, 2025, https://doi.org/10.48550/arxiv.2503.03965.
APA
Joshi, C. K., Fu, X., Liao, Y.-L., Gharakhanyan, V., Miller, B. K., Sriram, A., & Ulissi, Z. W. (2025). All-atom Diffusion Transformers: Unified generative modelling of molecules and materials. arXiv. https://doi.org/10.48550/arxiv.2503.03965
Chicago
Joshi, C. K., X. Fu, Y.-L. Liao, et al. 2025. “All-atom Diffusion Transformers: Unified Generative Modelling of Molecules and Materials”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2503.03965.
Harvard
Joshi, C.K. et al. (2025) “All-atom Diffusion Transformers: Unified generative modelling of molecules and materials”. arXiv. Available at: https://doi.org/10.48550/arxiv.2503.03965.
Vancouver
1. Joshi CK, Fu X, Liao Y-L, Gharakhanyan V, Miller BK, Sriram A, Ulissi ZW (2025) All-atom Diffusion Transformers: Unified generative modelling of molecules and materials. https://doi.org/10.48550/arxiv.2503.03965

BibTeX

@misc{https://doi.org/10.48550/arxiv.2503.03965,
  doi = {10.48550/ARXIV.2503.03965},
  url = {https://arxiv.org/abs/2503.03965},
  author = {Joshi, Chaitanya K. and Fu, Xiang and Liao, Yi-Lun and Gharakhanyan, Vahe and Miller, Benjamin Kurt and Sriram, Anuroop and Ulissi, Zachary W.},
  keywords = {Machine Learning (cs.LG), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {All-atom Diffusion Transformers: Unified generative modelling of molecules and materials},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/