Generative Flow Networks for Discrete Probabilistic Modeling
Dinghuai ZhangNikolay MalkinZhen LiuAlexandra VolokhovaAaron C. CourvilleYoshua Bengio
Proposes energy-based generative flow networks to overcome the slow mixing of traditional MCMC methods in high-dimensional discrete spaces by jointly training an energy function alongside a generative policy that amortizes mode-hopping exploration.
High-dimensional discrete data—such as binary images, text tokens, and symbolic graphs—present major challenges for standard probabilistic generative modeling. Classical energy-based approaches rely heavily on Markov Chain Monte Carlo (MCMC) methods to generate negative samples during training. However, in discrete spaces with complex, separated probability clusters (modes), traditional MCMC mixes slowly and often fails to traverse low-probability barriers. This leads to the discovery of spurious modes and inaccurate data representation, hindering effective deployment in complex real-world domains.
The article introduces and evaluates Energy-Based Generative Flow Networks (EB-GFN), a joint framework that trains an energy function alongside a generative flow network (GFlowNet) sampler directly from discrete datasets. The primary objective is to demonstrate that amortizing the sampling process into a learned stochastic policy enables efficient transitions across separated modes and improves discrete probabilistic modeling without needing predefined structural priors.
The authors construct a non-autoregressive sequential generation policy where discrete data vectors are constructed step-by-step through a directed acyclic graph. Training alternates between updating the GFlowNet via a trajectory balance objective using the current energy model as a reward signal, and updating the energy function using approximate maximum likelihood estimation driven by negative samples from the GFlowNet. The framework also implements a back-and-forth proposal mechanism combining backward erasure and forward construction to approximate high-dimensional block Gibbs sampling. The methodology was evaluated across synthetic 2D datasets remapped to binary strings via Gray codes, physics-based Ising models, and high-dimensional discrete image benchmarks including MNIST, Omniglot, and Caltech Silhouettes.
The experiments show that EB-GFN consistently outperforms or matches established baselines. In Ising model structure recovery, EB-GFN accurately inferred full interaction matrices from discrete samples, showing clear advantages over standard Gibbs and gradient-guided Gibbs sampling on complex, multi-modal negative coupling tasks. On 2D synthetic benchmarks, EB-GFN achieved lower test negative log-likelihood across all datasets and delivered better sample quality than comparable baselines without requiring oversized networks. On high-dimensional discrete image tasks, EB-GFN outperformed existing state-of-the-art methods in test likelihood on three of four benchmarks, including Omniglot and Caltech Silhouettes, and generated sharper visual reconstructions.
These findings indicate that discrete generative modeling can be significantly accelerated and stabilized by replacing iterative MCMC chains with trained sequential policies. By discovering and exploiting structural regularities in data distributions, the framework eliminates the computational bottlenecks and sample quality degradation typical of local MCMC exploration. This reduces training instability and lowers downstream sampling costs, making energy-based discrete modeling more viable for practical applications.
Organizations working with high-dimensional discrete representations should consider adopting GFlowNet-based samplers as an alternative to standard MCMC in generative pipelines. For implementation, practitioners should incorporate gradual proposal scheduling—starting from small local updates and expanding to full-dimension generation—and apply architectural enhancements such as layer normalization, which substantially boosted image modeling likelihoods in the study. Further research is recommended to explore iterating trained GFlowNet proposals as persistent exploration kernels and extending the framework to broader discrete domains such as biological sequences and program synthesis.
The main limitations include the requirement to train two interacting neural networks simultaneously, which increases optimization complexity compared to standalone models. In addition, ablation studies indicate that on challenging benchmarks like discrete image modeling, convergence relies heavily on both the back-and-forth proposal schedule and reverse trajectory sampling. Nevertheless, given the consistent performance gains across synthetic and benchmark tasks, confidence in the core findings remains high for binary discrete spaces.
- Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). Provides fundamental techniques for probabilistic generative modeling in discrete state spaces, establishing key background on discrete data generation that EB-GFN addresses via flow networks.
- Paper: Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions, Emiel Hoogeboom et al. (2021). Introduces foundational methods for learning categorical and discrete probability distributions using generative models, motivating the discrete probabilistic modeling challenge solved in this paper.
- Paper: The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables, Chris J. Maddison et al. (2016). Explains continuous relaxations and gradient estimation strategies for discrete random variables, providing useful conceptual groundwork for discrete probabilistic modeling.
- Paper: GFlowNet Foundations, Yoshua Bengio et al. (2023). Builds an extensive and rigorous mathematical framework for the theory, objectives, and flow properties of Generative Flow Networks established across discrete probabilistic applications.
- Paper: Joint Bayesian Inference of Graphical Structure and Parameters with a Single Generative Flow Network, Tristan Deleu et al. (2023). Extends GFlowNets to complex Bayesian inference tasks over discrete and continuous graphical structures and parameters, continuing the amortized generation principles explored here.
- Paper: Local Search GFlowNets, Minsu Kim et al. (2024). Enhances discrete GFlowNet sampling and exploration efficiency by combining amortized flow generation with local search algorithms.
