Biological Sequence Design with GFlowNets

Moksh JainEmmanuel BengioAlex Hernández-GarcíaJarrid Rector-BrooksBonaventure F. P. DossouChanakya Ajit EkboteJie FuTianyu ZhangMichael KilgourDinghuai Zhang

article2022ICML245 citations

Proposes an active learning framework powered by GFlowNets and epistemic uncertainty estimation to generate highly diverse, novel, and high-scoring biological sequences across iterative rounds of expensive experimental design.

Listen

Designing novel biological sequences such as proteins and DNA is critical for addressing urgent challenges like antimicrobial resistance and targeted therapeutics. However, navigating the astronomically large search space of molecular combinations is slow and costly. Traditional experimental pipelines rely on multi-stage filtering where early, approximate screenings frequently discard viable candidates. Standard machine learning search methods, including reinforcement learning, tend to focus narrowly on a single optimal sequence rather than providing a broad set of alternatives. If that single candidate fails in downstream testing, the entire cycle is wasted. Consequently, drug discovery programs require generative methods that reliably propose batches of candidates that are both high-performing and functionally diverse.

The article introduces and evaluates an active learning framework called GFlowNet-AL for de novo biological sequence design. The main objective is to demonstrate that using Generative Flow Networks (GFlowNets) as candidate generators—combined with uncertainty estimation and historical laboratory data—produces batches of biological sequences that achieve superior performance, diversity, and novelty compared to standard optimization methods.

To evaluate this approach, the researchers constructed an iterative active learning workflow. In this setup, a surrogate model estimates candidate quality along with epistemic uncertainty (the model's lack of confidence in unexplored regions). A GFlowNet then generates candidates proportionally to this reward, ensuring broad exploration across diverse solution modes rather than collapsing to a single local maximum. The framework was tested across three distinct biological design benchmarks: generating antimicrobial peptides (evaluated over 10 active learning rounds with batch sizes of 1,000), optimizing short DNA sequences for human transcription factor binding (a single-round batch of 128 candidates), and designing fluorescent proteins (a single-round batch of 128 candidates across sequences of length 237). The method was compared against established baselines, including reinforcement learning, Bayesian optimization, and model-based generative baselines.

The findings show that GFlowNet-AL consistently outperformed or matched baseline methods across key metrics while generating significantly more diverse and novel sequences. On the antimicrobial peptide task, GFlowNet-AL achieved a high performance score of 0.932 while roughly doubling sequence diversity (22.34 versus 12.12 for reinforcement learning) and tripling novelty relative to the starting library (28.44 versus 9.31). The generated peptides also demonstrated biologically realistic properties, closely matching natural amino acid distributions and maintaining low instability scores. On the protein fluorescence and DNA binding tasks, the framework achieved the highest overall scores while maintaining broad candidate diversity. Ablation studies revealed two key operational drivers: mixing 50% historical offline data with online generation yielded the fastest training convergence, and incorporating robust uncertainty estimates via neural network ensembles significantly improved exploration over single-model proxies.

These results demonstrate that GFlowNets provide a practical mechanism to de-risk biological discovery pipelines. By generating batches of distinct, high-quality candidates rather than near-identical variations, the framework mitigates the failure rate of downstream laboratory testing. Organizations adopting this approach can accelerate experimental design cycles, reduce the cost of wasted synthesis rounds, and improve the odds of identifying viable lead molecules. While traditional methods either struggle with computational scalability or converge prematurely, GFlowNet-AL amortizes search costs and scales efficiently to large batch generation.

Organizations developing sequence design workflows should consider adopting generative flow network architectures coupled with ensemble-based uncertainty modeling. Implementations should leverage existing historical laboratory datasets to seed training batches, targeting an even balance between offline data and on-policy exploration. For further technical development, future efforts should focus on optimizing proxy model retraining across rounds, developing non-autoregressive generation methods, and extending the framework to handle multi-fidelity feedback from different stages of wet-lab evaluations.

Confidence in these findings is supported by consistent performance across three diverse sequence modalities (short DNA, short peptides, and large proteins) and robust comparisons against established baselines. However, stakeholders should note that the evaluations relied on simulated wet-lab oracles and computational proxies rather than physical in-vivo synthesis. In addition, managing two separate learning models—the proxy evaluator and the generative policy—introduces architectural complexity and hyperparameter sensitivity that require careful operational tuning in production environments.

Cover for Biological Sequence Design with GFlowNets

Abstract

Design of de novo biological sequences with desired properties, like protein and DNA sequences, often involves an active loop with several rounds of molecule ideation and expensive wet-lab evaluations. These experiments can consist of multiple stages, with increasing levels of precision and cost of evaluation, where candidates are filtered. This makes the diversity of proposed candidates a key consideration in the ideation phase. In this work, we propose an active learning algorithm leveraging epistemic uncertainty estimation and the recently proposed GFlowNets as a generator of diverse candidate solutions, with the objective to obtain a diverse batch of useful (as defined by some utility function, for example, the predicted anti-microbial activity of a peptide) and informative candidates after each round. We also propose a scheme to incorporate existing labeled datasets of candidates, in addition to a reward function, to speed up learning in GFlowNets. We present empirical results on several biological sequence design tasks, and we find that our method generates more diverse and novel batches with high scoring candidates compared to existing approaches.

Table of Contents

  • 1. Introduction
  • 2. Problem Setup
  • 3. GFlowNets For Sequence Design
  • 3.1. Background
  • 3.2. Leveraging Data during Training
  • 3.3. Incorporating Epistemic Uncertainty
  • 4. Related Work
  • 5. Experiments
  • 5.1. Tasks and Evaluation Criteria
  • 5.2. Baselines and Implementation
  • 5.3. Results
  • 5.3.1. ANTI-MICROBIAL PEPTIDE DESIGN
  • 5.3.2. TF-BIND-8
  • 5.3.3. GFP
  • 5.4. Ablations
  • 5.4.1. TRAINING WITH THE ORACLE DATA
  • 5.4.2. EFFECT OF UNCERTAINTY ESTIMATES
  • 6. Conclusion and Future Work
  • Acknowledgements
  • References
  • A. Task Details
  • A.1. Anti-Microbial Peptides
  • A.2. TF-Bind-8
  • A.3. GFP
  • B. Implementation Details
  • B.1. Baselines
  • B.2. GFlowNet
  • C. Additional Results
  • C.1. AMP Generation: Additional Results
  • C.2. TF-Bind-8 and GFP: Additional Results
  • C.3. Effect of Uncertainty: Additional Results

Knowls

  1. Knowl 1 — GFlowNet-AL: Multi-Round Active Learning for Biological Sequence Design

    algorithm

    GFlowNet-AL is an active learning framework for biological sequence design that pairs an epistemic uncertainty-aware surrogate proxy model with a Generative Flow Network (GFlowNet) candidate generator. In each round of active learning, the surrogate proxy fits the accumulated dataset of labeled sequences Di−1\mathcal{D}_{i-1} and provides both a mean prediction μ(x)\mu(x) of sequence fitness and an epistemic uncertainty estimate σ(x)\sigma(x). An acquisition function F(μ,σ)\mathcal{F}(\mu, \sigma), such as the Upper Confidence Bound (UCB, defined as μ(x)+κσ(x)\mu(x) + \kappa \sigma(x) with κ=0.1\kappa = 0.1), defines a non-negative scalar reward function R(x)=(F(μ(x),σ(x)))βR(x) = (\mathcal{F}(\mu(x), \sigma(x)))^\beta (with temperature exponent β=3\beta = 3). The GFlowNet generator policy πθ\pi_\theta is trained to sample candidate sequences proportionally to this reward. A candidate batch is sampled from πθ\pi_\theta, screened with the proxy, evaluated on the black-box oracle O\mathcal{O}, and aggregated into the dataset for subsequent rounds.

    Input:
      O: Oracle function mapping a candidate sequence x to a scalar label y = O(x)
      D_0: Initial dataset of labeled sequence pairs {(x_j, y_j)}
      M: Surrogate proxy model providing predictive mean M.mu(x) and epistemic uncertainty M.sigma(x)
      pi_theta: GFlowNet generative policy parameterized by theta
      F(mu, sigma): Acquisition function combining mean and uncertainty into an unscaled reward
      b: Query batch size evaluated per round
      N: Total number of active learning rounds
      t: Oversampling multiplier to screen candidates with the proxy (set to t = 5)
    Result:
      The augmented dataset D_N containing all generated and evaluated candidates
    Initialize M and pi_theta
    for i = 1 to N do
      Fit proxy model M on current dataset D_{i-1}
      Define reward function R(x) = (F(M.mu(x), M.sigma(x)))^beta
      Train pi_theta using GFlowNet Inner Loop with reward function R and dataset D_{i-1}
      Sample candidate pool S of size t * b from pi_theta
      Select query batch B = {x_1, ..., x_b} subset of S having the highest proxy scores F(M.mu(x), M.sigma(x))
      Evaluate batch B using oracle O: D_hat_i = {(x_1, O(x_1)), ..., (x_b, O(x_b))}
      Update dataset: D_i = D_{i-1} union D_hat_i
    end
    return D_N

    The proxy model MM is parameterized as a Multi-Layer Perceptron (MLP) with 2 hidden layers of 2048 units and ReLU activations, trained with Adam at learning rate 10−410^{-4} using mean squared error loss and early stopping on a 10% validation split. The candidate generator πθ\pi_\theta uses an MLP with 2 hidden layers of 2048 units with action logits corresponding to vocabulary tokens.

  2. Knowl 2 — GFlowNet Inner Loop with Offline Trajectory Mixing

    algorithm

    Because GFlowNet objectives (such as trajectory balance) are off-policy, training can incorporate offline candidate trajectories constructed from known experimental datasets D={(xi,yi)}\mathcal{D} = \{(x_i, y_i)\} alongside online trajectories generated by the active policy. In autoregressive biological sequence generation (where one token is appended per step), the mapping from an initial empty state s0s_0 to any sequence xx is unique, providing an exact backward trajectory. During each training step, a minibatch is formed by mixing a fraction γ∈[0,1)\gamma \in [0, 1) of offline trajectories sampled from D\mathcal{D} and a fraction 1−γ1 - \gamma of on-policy trajectories sampled from an exploratory mixture policy πˉθ=(1−δ)PFθ+δ⋅Uniform\bar{\pi}_\theta = (1 - \delta) P_{F_\theta} + \delta \cdot \text{Uniform}. Mixing offline trajectories speeds up reward convergence and grounds exploration in high-performing regions of the sequence space.

    Input:
      D = {(x_i, y_i)}: Dataset of sequence candidates and known scores
      R(.): Non-negative reward function defined on completed sequences
      gamma: Proportion of offline trajectories in each training minibatch (set to 0.5)
      m: Training minibatch size (set to 32)
      T: Number of parameter update steps
      delta: Uniform action mixture coefficient for exploration (e.g., 0.001 to 0.05)
    Result:
      Learned forward transition policy pi_theta = P_{F_theta} such that pi_theta(x) is proportional to R(x)
    Initialize flow network parameters theta and learnable scalar log Z_theta
    for step = 1 to T do
      Compute m_online = ceil(m * (1 - gamma))
      Sample m_online trajectories from exploration policy pi_bar = (1 - delta) * P_{F_theta} + delta * Uniform
      Sample m - m_online sequences x from dataset D and reconstruct their deterministic trajectories tau = (s_0 -> ... -> x -> s_f)
      Combine sampled trajectories into a minibatch of size m
      Evaluate reward R(x) for the terminal state x of each trajectory in the minibatch
      Compute the Trajectory Balance loss L_TB(tau; theta) for each trajectory in the minibatch
      Update theta and log Z_theta via stochastic gradient descent on the average Trajectory Balance loss
    end
    return pi_theta
  3. Knowl 3 — Trajectory Balance Objective for Autoregressive GFlowNets

    equation

    For an autoregressive discrete sequence generator formulated as a directed acyclic graph G=(S,E)G = (\mathcal{S}, \mathcal{E}), an object x∈Xx \in \mathcal{X} is constructed via a trajectory of states τ=(s0→s1→⋯→sn=x→sf)\tau = (s_0 \to s_1 \to \dots \to s_n = x \to s_f), where s0s_0 is the empty sequence, each transition s→s′s \to s' appends a token from vocabulary A\mathcal{A}, and the transition x→sfx \to s_f is a special termination action. The Trajectory Balance loss LTB(τ;θ)\mathcal{L}_{TB}(\tau; \theta) over a trajectory τ\tau is defined as:

    LTB(τ;θ)=(log⁡Zθ∏s→s′∈τPFθ(s′∣s)R(x))2\mathcal{L}_{TB}(\tau; \theta) = \left( \log \frac{Z_\theta \prod_{s \to s' \in \tau} P_{F_\theta}(s' \mid s)}{R(x)} \right)^2

    where:

    • PFθ(s′∣s)P_{F_\theta}(s' \mid s) is the forward transition probability parameterized by neural network parameters θ\theta.
    • ZθZ_\theta is a learnable scalar parameter representing the total flow (the partition function of the unnormalized measure R(x)R(x)).
    • R(x)∈R+R(x) \in \mathbb{R}^+ is the non-negative target reward assigned to the terminated sequence xx.

    At the global optimum LTB(τ;θ)=0\mathcal{L}_{TB}(\tau; \theta) = 0 for all valid trajectories, the sampling policy πθ(x)=PFθ(x)\pi_\theta(x) = P_{F_\theta}(x) samples complete sequences with probability strictly proportional to their reward:

    πθ(x)=R(x)Z\pi_\theta(x) = \frac{R(x)}{Z}

    In autoregressive sequence generation where each state has a single parent, Trajectory Balance achieves faster credit assignment and greater stability for long sequences and large vocabularies compared to flow-matching objectives.

  4. Knowl 4 — Evaluation Desiderata and Metrics for Generated Biological Sequences

    definition

    In black-box biological sequence design starting from an initial baseline dataset D0\mathcal{D}_0 and using an oracle f:X→R+f: \mathcal{X} \to \mathbb{R}^+, candidate algorithms are evaluated over the top-KK highest-scoring newly generated sequences DBest=TopK(DN∖D0)\mathcal{D}_{\text{Best}} = \text{TopK}(\mathcal{D}_N \setminus \mathcal{D}_0) across three complementary metrics:

    1. Mean Performance: The average oracle fitness score of the top-KK candidates: Mean(D)=1∣D∣∑(xi,yi)∈Dyi\text{Mean}(\mathcal{D}) = \frac{1}{|\mathcal{D}|} \sum_{(x_i, y_i) \in \mathcal{D}} y_i

    2. Diversity: The average pairwise distance between distinct generated candidates in the top-KK set, measuring coverage of distinct modes: Diversity(D)=1∣D∣(∣D∣−1)∑(xi,yi)∈D∑(xj,yj)∈D∖{(xi,yi)}d(xi,xj)\text{Diversity}(\mathcal{D}) = \frac{1}{|\mathcal{D}|(|\mathcal{D}| - 1)} \sum_{(x_i, y_i) \in \mathcal{D}} \sum_{(x_j, y_j) \in \mathcal{D} \setminus \{(x_i, y_i)\}} d(x_i, x_j) where d(xi,xj)d(x_i, x_j) is a sequence distance metric (such as Hamming or Levenshtein edit distance).

    3. Novelty: The average distance from each top-KK generated sequence to its closest neighbor in the starting dataset D0\mathcal{D}_0, quantifying departure from known samples: Novelty(D)=1∣D∣∑(xi,yi)∈Dmin⁡sj∈D0d(xi,sj)\text{Novelty}(\mathcal{D}) = \frac{1}{|\mathcal{D}|} \sum_{(x_i, y_i) \in \mathcal{D}} \min_{s_j \in \mathcal{D}_0} d(x_i, s_j)

  5. Knowl 5 — Empirical Performance on Anti-Microbial Peptide (AMP) Design

    data/table

    Anti-microbial peptide (AMP) design requires generating peptide sequences of length ≤50\le 50 from a vocabulary of 20 amino acids across N=10N = 10 active learning rounds with batch size b=1000b = 1000, starting from an initial dataset D0\mathcal{D}_0 of 3219 positive AMPs and 4611 non-AMPs from DBAASP. The table reports the Mean Performance, Diversity, and Novelty for the top K=100K = 100 generated sequences (DBest \mathcal{D}_{\text{Best}}).

    Method Performance Diversity Novelty
    GFlowNet-AL 0.932±0.0020.932 \pm 0.002 22.34±1.24\mathbf{22.34 \pm 1.24} 28.44±1.32\mathbf{28.44 \pm 1.32}
    DynaPPO 0.938±0.009\mathbf{0.938 \pm 0.009} 12.12±1.7112.12 \pm 1.71 9.31±0.699.31 \pm 0.69
    COMs 0.761±0.0090.761 \pm 0.009 19.38±0.1419.38 \pm 0.14 26.47±1.3026.47 \pm 1.30
    GFlowNet 0.868±0.0150.868 \pm 0.015 11.32±0.6711.32 \pm 0.67 15.72±0.4415.72 \pm 0.44

    GFlowNet-AL achieves near-optimal oracle performance while producing batches that are roughly twice as diverse (22.3422.34 vs 12.1212.12) and three times as novel (28.4428.44 vs 9.319.31) compared to reinforcement learning (DynaPPO). Vanilla GFlowNet (without offline data or proxy uncertainty) suffers from lower performance (0.8680.868) and limited diversity (11.3211.32). Physicochemical analysis of the Top-100 GFlowNet-AL peptides demonstrates biological viability: mean instability index is 26.526.5 (maximum 3636, below the instability threshold of 4040) and the amino acid distribution closely matches natural AMPs.

  6. Knowl 6 — Empirical Performance on TF-Bind-8 and GFP Benchmarks

    data/table

    The proposed method was evaluated in single-round (N=1,b=128N = 1, b = 128) offline Model-Based Optimization (MBO) on two standard sequence benchmarks: TF-Bind-8 (length 8 DNA sequences, vocabulary size 4 nucleobases, starting dataset ∣D0∣=32,898|\mathcal{D}_0| = 32,898 containing only the lower 50% scoring sequences) and GFP (length 237 protein sequences, vocabulary size 20 amino acids, starting dataset ∣D0∣=5,000|\mathcal{D}_0| = 5,000 drawn between the 50th and 60th percentiles). The tables present Performance, Diversity, and Novelty metrics on the top K=128K = 128 candidates.

    TF-Bind-8 (K=128K = 128):

    Method Performance Diversity Novelty
    GFlowNet-AL 0.84±0.05\mathbf{0.84 \pm 0.05} 4.53±0.464.53 \pm 0.46 2.12±0.04\mathbf{2.12 \pm 0.04}
    DynaPPO 0.58±0.020.58 \pm 0.02 5.18±0.045.18 \pm 0.04 0.83±0.030.83 \pm 0.03
    COMs 0.74±0.040.74 \pm 0.04 4.36±0.244.36 \pm 0.24 1.16±0.111.16 \pm 0.11
    BO-qEI 0.44±0.050.44 \pm 0.05 4.78±0.174.78 \pm 0.17 0.62±0.230.62 \pm 0.23
    CbAS 0.45±0.140.45 \pm 0.14 5.35±0.165.35 \pm 0.16 0.46±0.040.46 \pm 0.04
    MINs 0.40±0.140.40 \pm 0.14 5.57±0.15\mathbf{5.57 \pm 0.15} 0.36±0.000.36 \pm 0.00
    CMA-ES 0.47±0.120.47 \pm 0.12 4.89±0.014.89 \pm 0.01 0.64±0.210.64 \pm 0.21
    AmortizedBO 0.62±0.010.62 \pm 0.01 4.97±0.064.97 \pm 0.06 1.00±0.571.00 \pm 0.57
    GFlowNet 0.72±0.030.72 \pm 0.03 4.72±0.134.72 \pm 0.13 1.14±0.301.14 \pm 0.30

    GFP (K=128K = 128):

    Method Performance Diversity Novelty
    GFlowNet-AL 0.853±0.004\mathbf{0.853 \pm 0.004} 211.51±0.73\mathbf{211.51 \pm 0.73} 210.56±0.82\mathbf{210.56 \pm 0.82}
    DynaPPO 0.794±0.0020.794 \pm 0.002 206.19±0.19206.19 \pm 0.19 203.20±0.47203.20 \pm 0.47
    COMs 0.831±0.0030.831 \pm 0.003 204.14±0.14204.14 \pm 0.14 201.64±0.42201.64 \pm 0.42
    BO-qEI 0.045±0.0030.045 \pm 0.003 139.89±0.18139.89 \pm 0.18 203.60±0.06203.60 \pm 0.06
    CbAS 0.817±0.0120.817 \pm 0.012 5.42±0.185.42 \pm 0.18 1.81±0.161.81 \pm 0.16
    MINs 0.761±0.0070.761 \pm 0.007 5.39±0.005.39 \pm 0.00 2.42±0.002.42 \pm 0.00
    CMA-ES 0.063±0.0030.063 \pm 0.003 201.43±0.12201.43 \pm 0.12 203.82±0.09203.82 \pm 0.09
    AmortizedBO 0.051±0.0010.051 \pm 0.001 205.32±0.12205.32 \pm 0.12 202.34±0.25202.34 \pm 0.25
    GFlowNet 0.743±0.0500.743 \pm 0.050 200.72±3.42200.72 \pm 3.42 202.11±1.54202.11 \pm 1.54

    On GFP, generative VAE-based baselines (CbAS, MINs) experience mode collapse onto the training set, yielding negligible novelty (1.811.81 and 2.422.42) and diversity (5.425.42 and 5.395.39), whereas GFlowNet-AL produces candidates with high performance (0.8530.853), diversity (211.51211.51), and novelty (210.56210.56).

  7. Knowl 7 — Optimal Offline Trajectory Proportion in GFlowNet Training

    empirical result

    In GFlowNet-AL, the proportion γ∈[0,1]\gamma \in [0, 1] of offline trajectories sampled from the dataset D\mathcal{D} during each training minibatch controls the balance between offline grounding and on-policy exploration. Evaluating top-KK (K=100K = 100) sequence rewards across training steps on the AMP task under γ∈{0,0.25,0.50,0.75,1.0}\gamma \in \{0, 0.25, 0.50, 0.75, 1.0\} reveals:

    • Pure online training (γ=0\gamma = 0) converges slowly, reaching a top-KK reward of only ≈0.35\approx 0.35 after 10,000 steps.
    • Mixing offline data substantially accelerates learning speed and final reward: γ=0.25\gamma = 0.25 reaches ≈0.55\approx 0.55, while γ=0.50\gamma = 0.50 (50% offline data) achieves the highest top-KK reward of ≈0.65\approx 0.65 within 10,000 steps.
    • Increasing offline data beyond 50% degrades performance: γ=0.75\gamma = 0.75 and γ=1.0\gamma = 1.0 (pure offline) stall at top-KK rewards under 0.450.45 due to insufficient exploration and policy overfitting to the static training dataset.

    Thus, a 1:1 mixture (γ=0.50\gamma = 0.50) of on-policy exploratory trajectories and offline dataset trajectories provides the optimal trade-off between exploration and rapid credit assignment.

  8. Knowl 8 — Effect of Epistemic Uncertainty Estimation Methods on Candidate Diversity

    data/table

    Ablating the epistemic uncertainty estimation component in GFlowNet-AL across active learning rounds on the AMP design task demonstrates that incorporating uncertainty into the reward function directly enhances candidate diversity and novelty, with higher-fidelity uncertainty estimators yielding greater improvements.

    Surrogate Uncertainty Method Performance Diversity Novelty
    Deep Ensembles (5 models) 0.932±0.002\mathbf{0.932 \pm 0.002} 22.34±1.24\mathbf{22.34 \pm 1.24} 28.44±1.32\mathbf{28.44 \pm 1.32}
    MC Dropout (25 samples) 0.921±0.0040.921 \pm 0.004 18.58±1.7818.58 \pm 1.78 19.58±1.1219.58 \pm 1.12
    No Uncertainty (Single model) 0.909±0.0080.909 \pm 0.008 16.42±0.7416.42 \pm 0.74 17.24±1.4417.24 \pm 1.44

    Comparing acquisition functions with Deep Ensembles shows that Upper Confidence Bound (UCB: Performance 0.932±0.0020.932 \pm 0.002, Diversity 22.34±1.2422.34 \pm 1.24, Novelty 28.44±1.3228.44 \pm 1.32) and Expected Improvement (EI: Performance 0.928±0.0020.928 \pm 0.002, Diversity 23.61±1.0523.61 \pm 1.05, Novelty 26.52±1.5626.52 \pm 1.56) perform comparably. The choice of uncertainty estimation quality (Ensembles vs MC Dropout vs None) is the dominant factor determining exploration efficiency and generation diversity.

  9. Knowl 9 — Limitations of GFlowNet-AL in Active Sequence Design

    limitation

    GFlowNet-AL exhibits specific practical and architectural limitations:

    1. Dual-Learner Overhead: The architecture decouples the surrogate proxy (learner 1) and the generative policy (learner 2), each with distinct optimization objectives, training routines, and hyperparameters that must be tuned simultaneously.
    2. Proxy Retraining in Continual Settings: Retraining the surrogate proxy model from scratch across multiple active learning rounds is computationally expensive; continual learning mechanisms to update the proxy incrementally without catastrophic forgetting remain unaddressed.
    3. Single-Fidelity Oracle Assumption: The framework assumes queries are directed to a single oracle at a fixed cost, whereas real-world drug and sequence discovery involves hierarchical multi-fidelity evaluation pipelines (ranging from high-throughput in-silico screening to low-throughput in-vitro and animal trials) that are not natively handled by the outer loop policy.

Coverage note — None was omitted; all key algorithms, theoretical training objectives, evaluation metrics, empirical results across AMP, TF-Bind-8, and GFP tasks, ablations on trajectory mixing and uncertainty, and stated limitations are fully represented.

References

  1. 1.Aggarwal, C. C., Kong, X., Gu, Q., Han, J., and Philip, S. Y. Active learning: A survey. In Data Classification: Algorithms and Applications, pp. 571–605. CRC Press, 2014.
  2. 2.Angermueller, C., Dohan, D., Belanger, D., Deshpande, R., Murphy, K., and Colwell, L. Model-based reinforcement learning for biological sequence design. In International Conference on Learning Representations, 2019.
  3. 3.Audet, C. and Hare, W. Derivative-Free and Blackbox Optimization. Springer Series in Operations Research and Financial Engineering. Springer International Publishing, 2017. ISBN 9783319689135. URL https://books.google.ca/books?id=ejVBDwAAQBAJ.
  4. 4.Barrera, L. A., Vedenko, A., Kurland, J. V., Rogers, J. M., Gisselbrecht, S. S., Rossin, E. J., Woodard, J., Mariani, L., Kock, K. H., Inukai, S., Siggers, T., Shokri, L., Gordan, R., Sahni, N., Cotsapas, C., Hao, T., Yi, S., Kellis, M., Daly, M. J., Vidal, M., Hill, D. E., and Bulyk, M. L. Survey of variation in human transcription factors reveals prevalent dna binding changes. Science, 351(6280):1450–1454, 2016a. doi: 10.1126/science.aad2257. URL https://www.science.org/doi/abs/10.1126/science.aad2257.
  5. 5.Barrera, L. A., Vedenko, A., Kurland, J. V., Rogers, J. M., Gisselbrecht, S. S., Rossin, E. J., Woodard, J., Mariani, L., Kock, K. H., Inukai, S., et al. Survey of variation in human transcription factors reveals prevalent dna binding changes. Science, 351(6280):1450–1454, 2016b.
  6. 6.Belanger, D., Vora, S., Mariet, Z., Deshpande, R., Dohan, D., Angermueller, C., Murphy, K., Chapelle, O., and Colwell, L. Biological sequences design using batched bayesian optimization, 2019.
  7. 7.Bengio, E., Jain, M., Korablyov, M., Precup, D., and Bengio, Y. Flow network based generative models for noniterative diverse candidate generation. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, 2021a. URL https://openreview.net/forum?id=Arn2E4IRjEB.
  8. 8.Bengio, Y., Deleu, T., Hu, E. J., Lahlou, S., Tiwari, M., and Bengio, E. Gflownet foundations, 2021b.
  9. 9.Boitreaud, J., Mallet, V., Oliver, C., and Waldispuhl, J. Optimol: Optimization of binding affinities in chemical space for drug discovery. Journal of Chemical Information and Modeling, 60(12):5658–5666, 2020. doi: 10.1021/acs.jcim.0c00833. URL https://doi.org/10.1021/acs.jcim.0c00833. PMID: 32986426.
  10. 10.Brookes, D., Park, H., and Listgarten, J. Conditioning by adaptive sampling for robust design. In International conference on machine learning, pp. 773–782. PMLR, 2019a.
  11. 11.Brookes, D. H., Park, H., and Listgarten, J. Conditioning by adaptive sampling for robust design. In ICML, 2019b.
  12. 12.Buesing, L., Heess, N., and Weber, T. Approximate inference in discrete distributions with monte carlo tree search and value functions. Artificial Intelligence and Statistics (AISTATS), 2019.
  13. 13.Cock, P. J. A., Antao, T., Chang, J. T., Chapman, B. A., Cox, C. J., Dalke, A., Friedberg, I., Hamelryck, T., Kauff, F., Wilczynski, B., and de Hoon, M. J. L. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics, 25(11):1422–1423, 03 2009. ISSN 1367-4803. doi: 10.1093/bioinformatics/btp163. URL https://doi.org/10.1093/bioinformatics/btp163.
  14. 14.Dadkhahi, H., Rios, J., Shanmugam, K., Das, Payel Melnyk, I., Das, P., Chenthamarakshan, V., and Lozano, A. Fourier representations for black-box optimization over categorical variables. arXiv preprint arXiv:2111.06801, 2021.
  15. 15.Das, P., Sercu, T., Wadhawan, K., Padhi, I., Gehrmann, S., Cipcigan, F., Chenthamarakshan, V., Strobelt, H., Dos Santos, C., Chen, P.-Y., et al. Accelerated antimicrobial discovery via deep generative models and molecular dynamics simulations. Nature Biomedical Engineering, 5(6):613–623, 2021.
  16. 16.Elnaggar, A., Heinzinger, M., Dallago, C., Rihawi, G., Wang, Y., Jones, L., Gibbs, T., Feher, T., Angerer, C., Steinegger, M., et al. Prottrans: towards cracking the language of life's code through self-supervised deep learning and high performance computing. arXiv preprint arXiv:2007.06225, 2020.
  17. 17.Gal, Y. and Ghahramani, Z. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML'16, pp. 1050–1059, 2016.
  18. 18.Garnett, R. Bayesian Optimization. Cambridge University Press, 2022. in preparation.
  19. 19.Guo, H., Tan, B., Liu, Z., Xing, E. P., and Hu, Z. Text generation with efficient (soft) q-learning, 2021.
  20. 20.Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. Reinforcement learning with deep energy-based policies. International Conference on Machine Learning (ICML), 2017.
  21. 21.Hansen, N. The cma evolution strategy: a comparing review. Towards a new evolutionary computation, pp. 75–102, 2006.
  22. 22.Hoffman, S. C., Chenthamarakshan, V., Wadhawan, K., Chen, P.-Y., and Das, P. Optimizing molecules using efficient queries from property evaluations. Nature Machine Intelligence, pp. 1–11, 2021.
  23. 23.Jain, M., Lahlou, S., Nekoei, H., Butoi, V., Bertin, P., Rector-Brooks, J., Korablyov, M., and Bengio, Y. Deup: Direct epistemic uncertainty prediction, 2021.
  24. 24.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization, 2017.
  25. 25.Kumar, A. and Levine, S. Model inversion networks for model-based optimization. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 5126–5137. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/373e4c5d8edfa8b74fd4b6791d0cf6dc-Paper.pdf.
  26. 26.Lakshminarayanan, B., Pritzel, A., and Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17, pp. 6405–6416, Red Hook, NY, USA, 2017. Curran Associates Inc. ISBN 9781510860964.
  27. 27.Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y. Trajectory balance: Improved credit assignment in gflownets, 2022.
  28. 28.Melnyk, I., Das, P., Chenthamarakshan, V., and Lozano, A. Benchmarking deep generative models for diverse antibody sequence design. arXiv preprint arXiv:2111.06801, 2021.
  29. 29.Mockus, J. On bayesian methods for seeking the extremum. In Marchuk, G. I. (ed.), Optimization Techniques IFIP Technical Conference Novosibirsk, July 1–7, 1974, pp. 400–404, Berlin, Heidelberg, 1975. Springer Berlin Heidelberg. ISBN 978-3-540-37497-8.
  30. 30.Moss, H. B., Leslie, D. S., Beck, D., Gonzalez, J., and Rayson, P. Boss: Bayesian optimization over string spaces. In Advances in Neural Information Processing Systems, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/b19aa25ff58940d974234b48391b9549-Abstract.html.
  31. 31.Mullis, M. M., Rambo, I. M., Baker, B. J., and Reese, B. K. Diversity, ecology, and prevalence of antimicrobials in nature. Frontiers in microbiology, 10:2518, 2019.
  32. 32.Murray, C. J., Ikuta, K. S., Sharara, F., Swetschinski, L., Robles Aguilar, G., Gray, A., Han, C., Bisignano, C., Rao, P., Wool, E., Johnson, S. C., Browne, A. J., Chipeta, M. G., Fell, F., Hackett, S., Haines-Woodhouse, G., Kashef Hamadani, B. H., Kumaran, E. A. P., McManigal, B., Agarwal, R., Akech, S., Albertson, S., Amuasi, J., Andrews, J., Aravkin, A., Ashley, E., Bailey, F., Baker, S., Basnyat, B., Bekker, A., Bender, R., Bethou, A., Bielicki, J., Boonkasidecha, S., Bukosia, J., Carvalheiro, C., Castaneda-Orjuela, C., Chansamouth, V., Chaurasia, S., Chiurchiu, S., Chowdhury, F., Cook, A. J., Cooper, B., Cressey, T. R., Criollo-Mora, E., Cunningham, M., Darboe, S., Day, N. P. J., De Luca, M., Dokova, K., Dramowski, A., Dunachie, S. J., Eckmanns, T., Eibach, D., Emami, A., Feasey, N., Fisher-Pearson, N., Forrest, K., Garrett, D., Gastmeier, P., Giref, A. Z., Greer, R. C., Gupta, V., Haller, S., Haselbeck, A., Hay, S. I., Holm, M., Hopkins, S., Iregbu, K. C., Jacobs, J., Jarovsky, D., Javanmardi, F., Khorana, M., Kissoon, N., Kobeissi, E., Kostyanev, T., Krapp, F., Krumkamp, R., Kumar, A., Kyu, H. H., Lim, C., Limmathurotsakul, D., Loftus, M. J., Lunn, M., Ma, J., Mturi, N., Munera-Huertas, T., Musicha, P., Mussi-Pinhata, M. M., Nakamura, T., Nanavati, R., Nangia, S., Newton, P., Ngoun, C., Novotney, A., Nwakanma, D., Obiero, C. W., Olivas-Martinez, A., Olliaro, P., Ooko, E., Ortiz-Brizuela, E., Peleg, A. Y., Perrone, C., Plakkal, N., de Leon, A. P., Raad, M., Ramdin, T., Riddell, A., Roberts, T., Robotham, J. V., Roca, A., Rudd, K. E., Russell, N., Schnall, J., Scott, J. A. G., Shivamallappa, M., Sifuentes-Osornio, J., Steenkeste, N., Stewardson, A. J., Stoeva, T., Tasak, N., Thaiprakong, A., Thwaites, G., Turner, C., Turner, P., van Doorn, H. R., Velaphi, S., Vongpradith, A., Vu, H., Walsh, T., Waner, S., Wangrangsimakul, T., Wozniak, T., Zheng, P., Sartorius, B., Lopez, A. D., Stergachis, A., Moore, C., Dolecek, C., and Naghavi, M. Global burden of bacterial antimicrobial resistance in 2019: a systematic analysis. The Lancet, 2022. ISSN 0140-6736. doi: https://doi.org/10.1016/S0140-6736(21)02724-0. URL https://www.sciencedirect.com/science/article/pii/S0140673621027240.
  33. 33.Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. Bridging the gap between value and policy based reinforcement learning. Advances in Neural Information Processing Systems, 30:2775–2785, 2017.
  34. 34.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32, pp. 8024–8035. Curran Associates, Inc., 2019.
  35. 35.Pirtskhalava, M., Amstrong, A. A., Grigolava, M., Chubinidze, M., Alimbarashvili, E., Vishnepolsky, B., Gabrielian, A., Rosenthal, A., Hurt, D. E., and Tartakovsky, M. Dbaasp v3: Database of antimicrobial/cytotoxic activity and structure of peptides as a resource for development of new therapeutics. Nucleic Acids Research, 49(D1):D288–D297, 2021.
  36. 36.Pyzer-Knapp, E. O. Bayesian optimization for accelerated drug discovery. IBM Journal of Research and Development, 62(6):2:1–2:7, 2018. doi: 10.1147/JRD.2018.2881731.
  37. 37.Rao, R., Bhattacharya, N., Thomas, N., Duan, Y., Chen, X., Canny, J., Abbeel, P., and Song, Y. S. Evaluating protein transfer learning with tape. In Advances in Neural Information Processing Systems, 2019a.
  38. 38.Rao, R., Bhattacharya, N., Thomas, N., Duan, Y., Chen, X., Canny, J. F., Abbeel, P., and Song, Y. S. Evaluating protein transfer learning with tape. bioRxiv, 2019b.
  39. 39.Rasmussen, C. and Williams, C. Gaussian Processes for Machine Learning. Adaptive Computation and Machine Learning series. MIT Press, 2005. ISBN 9780262182539. URL https://books.google.ca/books?id=GhoSngEACAAJ.
  40. 40.Sarkisyan, K. S., Bolotin, D. A., Meer, M. V., Usmanova, D. R., Mishin, A. S., Sharonov, G. V., Ivankov, D. N., Bozhanova, N. G., Baranov, M. S., Soylemez, O., et al. Local fitness landscape of the green fluorescent protein. Nature, 533(7603):397–401, 2016.
  41. 41.Sinai, S., Wang, R., Whatley, A., Slocum, S., Locane, E., and Kelsic, E. Adalead: A simple and robust adaptive greedy search algorithm for sequence design. arXiv preprint, 2020.
  42. 42.Srinivas, N., Krause, A., Kakade, S., and Seeger, M. Gaussian process optimization in the bandit setting: No regret and experimental design. In Proceedings of the 27th International Conference on International Conference on Machine Learning, ICML'10, pp. 1015–1022, Madison, WI, USA, 2010. Omnipress. ISBN 9781605589077.
  43. 43.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. MIT press, 2018.
  44. 44.Swersky, K., Rubanova, Y., Dohan, D., and Murphy, K. Amortized bayesian optimization over discrete spaces. In Peters, J. and Sontag, D. (eds.), Proceedings of the 36th Conference on Uncertainty in Artificial Intelligence (UAI), volume 124 of Proceedings of Machine Learning Research, pp. 769–778. PMLR, 03–06 Aug 2020. URL https://proceedings.mlr.press/v124/swersky20a.html.
  45. 45.Terayama, K., Sumita, M., Tamura, R., and Tsuda, K. Black-box optimization for automated discovery. Accounts of Chemical Research, XXXX, 02 2021. doi: 10.1021/acs.accounts.0c00713.
  46. 46.Trabucco, B., Kumar, A., Geng, X., and Levine, S. Conservative objective models for effective offline model-based optimization. In International Conference on Machine Learning, pp. 10358–10368. PMLR, 2021a.
  47. 47.Trabucco, B., Kumar, A., Geng, X., and Levine, S. Design-bench: Benchmarks for data-driven offline model-based optimization, 2021b. URL https://openreview.net/forum?id=cQzf26aA3vM.
  48. 48.Wilson, J. T., Moriconi, R., Hutter, F., and Deisenroth, M. P. The reparameterization trick for acquisition functions, 2017.
  49. 49.Zhang, D., Fu, J., Bengio, Y., and Courville, A. C. Unifying likelihood-free inference with black-box sequence design and beyond. ArXiv, abs/2110.03372, 2021.
  50. 50.Zhang, D., Malkin, N., Liu, Z., Volokhova, A., Courville, A. C., and Bengio, Y. Generative flow networks for discrete probabilistic modeling. ArXiv, abs/2202.01361, 2022.

Citation

MLA
Jain, M., et al. “Biological Sequence Design with GFlowNets”. International Conference on Machine Learning, vol. 162, 2022, pp. 9786–801, https://proceedings.mlr.press/v162/jain22a.html.
APA
Jain, M., Bengio, E., Hernandez-Garcia, A., Rector-Brooks, J., Dossou, B. F. P., Ekbote, C. A., Fu, J., Zhang, T., Kilgour, M., Zhang, D., Simine, L., Das, P., & Bengio, Y. (2022). Biological Sequence Design with GFlowNets. International Conference on Machine Learning, 162, 9786–9801. https://proceedings.mlr.press/v162/jain22a.html
Chicago
Jain, M., E. Bengio, A. Hernandez-Garcia, et al. 2022. “Biological Sequence Design with GFlowNets”. International Conference on Machine Learning 162: 9786–9801. https://proceedings.mlr.press/v162/jain22a.html.
Harvard
Jain, M. et al. (2022) “Biological Sequence Design with GFlowNets”, International Conference on Machine Learning. PMLR, pp. 9786–9801. Available at: https://proceedings.mlr.press/v162/jain22a.html.
Vancouver
1. Jain M, Bengio E, Hernandez-Garcia A, et al (2022) Biological Sequence Design with GFlowNets. In: International Conference on Machine Learning. PMLR, pp 9786–9801

BibTeX

@InProceedings{pmlr-v162-jain22a,
  title = 	 {Biological Sequence Design with {GF}low{N}ets},
  author =       {Jain, Moksh and Bengio, Emmanuel and Hernandez-Garcia, Alex and Rector-Brooks, Jarrid and Dossou, Bonaventure F. P. and Ekbote, Chanakya Ajit and Fu, Jie and Zhang, Tianyu and Kilgour, Michael and Zhang, Dinghuai and Simine, Lena and Das, Payel and Bengio, Yoshua},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {9786--9801},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/jain22a/jain22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/jain22a.html},
  abstract = 	 {Design of de novo biological sequences with desired properties, like protein and DNA sequences, often involves an active loop with several rounds of molecule ideation and expensive wet-lab evaluations. These experiments can consist of multiple stages, with increasing levels of precision and cost of evaluation, where candidates are filtered. This makes the diversity of proposed candidates a key consideration in the ideation phase. In this work, we propose an active learning algorithm leveraging epistemic uncertainty estimation and the recently proposed GFlowNets as a generator of diverse candidate solutions, with the objective to obtain a diverse batch of useful (as defined by some utility function, for example, the predicted anti-microbial activity of a peptide) and informative candidates after each round. We also propose a scheme to incorporate existing labeled datasets of candidates, in addition to a reward function, to speed up learning in GFlowNets. We present empirical results on several biological sequence design tasks, and we find that our method generates more diverse and novel batches with high scoring candidates compared to existing approaches.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/