Biological Sequence Design with GFlowNets
Moksh JainEmmanuel BengioAlex Hernández-GarcíaJarrid Rector-BrooksBonaventure F. P. DossouChanakya Ajit EkboteJie FuTianyu ZhangMichael KilgourDinghuai Zhang
Proposes an active learning framework powered by GFlowNets and epistemic uncertainty estimation to generate highly diverse, novel, and high-scoring biological sequences across iterative rounds of expensive experimental design.
Designing novel biological sequences such as proteins and DNA is critical for addressing urgent challenges like antimicrobial resistance and targeted therapeutics. However, navigating the astronomically large search space of molecular combinations is slow and costly. Traditional experimental pipelines rely on multi-stage filtering where early, approximate screenings frequently discard viable candidates. Standard machine learning search methods, including reinforcement learning, tend to focus narrowly on a single optimal sequence rather than providing a broad set of alternatives. If that single candidate fails in downstream testing, the entire cycle is wasted. Consequently, drug discovery programs require generative methods that reliably propose batches of candidates that are both high-performing and functionally diverse.
The article introduces and evaluates an active learning framework called GFlowNet-AL for de novo biological sequence design. The main objective is to demonstrate that using Generative Flow Networks (GFlowNets) as candidate generators—combined with uncertainty estimation and historical laboratory data—produces batches of biological sequences that achieve superior performance, diversity, and novelty compared to standard optimization methods.
To evaluate this approach, the researchers constructed an iterative active learning workflow. In this setup, a surrogate model estimates candidate quality along with epistemic uncertainty (the model's lack of confidence in unexplored regions). A GFlowNet then generates candidates proportionally to this reward, ensuring broad exploration across diverse solution modes rather than collapsing to a single local maximum. The framework was tested across three distinct biological design benchmarks: generating antimicrobial peptides (evaluated over 10 active learning rounds with batch sizes of 1,000), optimizing short DNA sequences for human transcription factor binding (a single-round batch of 128 candidates), and designing fluorescent proteins (a single-round batch of 128 candidates across sequences of length 237). The method was compared against established baselines, including reinforcement learning, Bayesian optimization, and model-based generative baselines.
The findings show that GFlowNet-AL consistently outperformed or matched baseline methods across key metrics while generating significantly more diverse and novel sequences. On the antimicrobial peptide task, GFlowNet-AL achieved a high performance score of 0.932 while roughly doubling sequence diversity (22.34 versus 12.12 for reinforcement learning) and tripling novelty relative to the starting library (28.44 versus 9.31). The generated peptides also demonstrated biologically realistic properties, closely matching natural amino acid distributions and maintaining low instability scores. On the protein fluorescence and DNA binding tasks, the framework achieved the highest overall scores while maintaining broad candidate diversity. Ablation studies revealed two key operational drivers: mixing 50% historical offline data with online generation yielded the fastest training convergence, and incorporating robust uncertainty estimates via neural network ensembles significantly improved exploration over single-model proxies.
These results demonstrate that GFlowNets provide a practical mechanism to de-risk biological discovery pipelines. By generating batches of distinct, high-quality candidates rather than near-identical variations, the framework mitigates the failure rate of downstream laboratory testing. Organizations adopting this approach can accelerate experimental design cycles, reduce the cost of wasted synthesis rounds, and improve the odds of identifying viable lead molecules. While traditional methods either struggle with computational scalability or converge prematurely, GFlowNet-AL amortizes search costs and scales efficiently to large batch generation.
Organizations developing sequence design workflows should consider adopting generative flow network architectures coupled with ensemble-based uncertainty modeling. Implementations should leverage existing historical laboratory datasets to seed training batches, targeting an even balance between offline data and on-policy exploration. For further technical development, future efforts should focus on optimizing proxy model retraining across rounds, developing non-autoregressive generation methods, and extending the framework to handle multi-fidelity feedback from different stages of wet-lab evaluations.
Confidence in these findings is supported by consistent performance across three diverse sequence modalities (short DNA, short peptides, and large proteins) and robust comparisons against established baselines. However, stakeholders should note that the evaluations relied on simulated wet-lab oracles and computational proxies rather than physical in-vivo synthesis. In addition, managing two separate learning models—the proxy evaluator and the generative policy—introduces architectural complexity and hyperparameter sensitivity that require careful operational tuning in production environments.
- Paper: GFlowNet Foundations, Yoshua Bengio et al. (2023). It formalizes the core theoretical foundations and training objectives of Generative Flow Networks upon which this biological sequence design framework directly builds.
- Paper: A Survey of Deep Active Learning, Pengzhen Ren et al. (2020). It provides essential background on deep active learning principles and batch uncertainty sampling strategies utilized in the active design loop.
- Paper: Reinforcement Learning with Deep Energy-Based Policies, Tuomas Haarnoja et al. (2017). It introduces energy-based reinforcement learning and maximum entropy policy formulation for diverse exploration that conceptually motivates GFlowNet-based generation.
- Paper: A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning, Eric Brochu et al. (2010). It offers foundational concepts on Bayesian optimization, active acquisition functions, and surrogate modeling for expensive black-box objective functions.
- Paper: Local Search GFlowNets, Minsu Kim et al. (2024). It enhances GFlowNet exploration and high-reward discovery in biological sequence and molecular benchmarks by integrating local search refinement into the training loop.
- Paper: Better Training of GFlowNets with Local Credit and Incomplete Trajectories, Ling Pan et al. (2023). It improves GFlowNet credit assignment on sequence generation tasks by incorporating intermediate energy feedback and learning from incomplete trajectories.
- Paper: Generative Flow Networks for Discrete Probabilistic Modeling, Dinghuai Zhang et al. (2022). It extends GFlowNet methodology to general discrete probabilistic modeling and energy-based training on discrete datasets.
- Paper: Joint Bayesian Inference of Graphical Structure and Parameters with a Single Generative Flow Network, Tristan Deleu et al. (2023). It generalizes GFlowNet sampling to perform joint Bayesian posterior inference over complex discrete graph structures and continuous parameters.
- Paper: Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization, Wenhao Gao et al. (2022). It establishes a standardized benchmark to rigorously evaluate the sample efficiency and optimization performance of generative molecular and sequence design algorithms.
- Paper: ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design, Pascal Notin et al. (2023). It presents large-scale benchmark suites and evaluations for protein fitness prediction and design models across diverse biological assays.
