RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design

Meghana KshirsagarChing-An ChengAllen NieFanglei XueRahul DodhiaJuan M. Lavista FerresKevin Kaichuang YangF. Dimaio

article2026arXiv0 citations

Introduces RosettaSearch, an inference-time optimization framework that uses large language models as generative search agents to rescue failed ProteinMPNN and LigandMPNN candidates, boosting protein sequence design success rates by 2.5 times without requiring model retraining.

Listen

Protein sequence design—the task of generating amino acid sequences that reliably fold into specific three-dimensional target structures—is essential for advancing drug discovery, enzyme engineering, and therapeutic development. However, state-of-the-art computational design tools frequently produce candidate sequences with low structural fidelity. Because these tools rely on single-pass autoregressive generation, they lack self-correction mechanisms to detect or fix structural errors. Consequently, poorly folded designs often proceed to expensive wet-lab testing where they fail, creating a major cost and timeline bottleneck. The article presents RosettaSearch, an inference-time optimization framework that uses large language models as generative optimizers to iteratively refine protein sequences based on direct structural feedback, eliminating the need for expensive model retraining or task-specific fine-tuning.

The main objective of the article is to demonstrate that frontier language models can serve as generative optimizers within a structured, priority-based search algorithm to systematically improve the structural fidelity of protein designs under strict computational budgets. To achieve this, the authors evaluated RosettaSearch across approximately 400 monomeric proteins from the Protein Data Bank (PDB) with low baseline fidelity, 275 de novo computational backbones from the Dayhoff atlas, and 50 computationally optimized binders from BindCraft. The framework couples candidate sequence generation with structural predictions from RosettaFold3, which provides scalar rewards and residue-level text annotations identifying low-confidence and misaligned regions. The optimization was evaluated under tight evaluation budgets (up to 75 structure calls per target) and independently validated using a separate structure prediction model, Chai-1, to guard against model-specific evaluation bias.

The key findings show that RosettaSearch substantially improves protein design quality across diverse benchmarks. First, when applied to suboptimal designs from LigandMPNN, RosettaSearch improved structural similarity metrics by 18% to 68%, driving a 2.5-fold increase in the design success rate (from 7.9% to 20.5% under multi-objective priority search). Under a more stringent multi-metric threshold, the success rate more than tripled from 2.5% to 8.9%. Second, these gains proved robust when evaluated with the independent Chai-1 predictor, confirming that performance improvements reflect genuine structural enhancement rather than overfitting to the primary prediction tool. In contrast, an information-matched baseline using random mutations failed to improve starting designs, proving that the language model's reasoning is essential for proposing effective, chemically informed edits. Third, RosettaSearch demonstrated strong generalization across other problem settings: it raised the success rate of de novo Dayhoff backbone designs from 72.4% to 89.5% without any reference sequence context, and consistently improved structural fidelity across 48 complex protein binders from BindCraft.

These findings indicate that generative search can serve as a highly flexible, plug-and-play optimization layer on top of existing sequence generation pipelines. Because RosettaSearch operates entirely at inference time and accommodates modular reward functions, engineering teams can optimize for multiple structural and functional constraints without undergoing costly model retraining or dataset curation. In addition, comparative evaluations across different model families (o4-mini, o3-mini, and Gemini-3) showed that optimization performance scales directly with underlying reasoning capability, highlighting that general-purpose reasoning models can be effectively harnessed for complex biological design.

Organizations developing computational protein pipelines should consider integrating inference-time generative search to refine candidate pools before committing to wet-lab synthesis. When implementing this approach, teams should utilize parallel priority search over greedy sequential revision, as parallel exploration significantly reduces premature convergence. Furthermore, systems should combine global numerical scores with residue-level textual feedback to provide actionable spatial context, while soft textual constraints should be enforced to prevent common optimization failure modes such as repetitive sequence motifs or length drift.

Despite these strong results, several limitations should be noted. RosettaSearch focuses on optimizing a single high-fidelity sequence per backbone rather than generating large, diverse candidate libraries, and the optimization process remains computationally bounded by the speed of structure prediction models (averaging approximately 30.6 minutes per protein). Furthermore, expert analysis revealed that while the language models propose valid design intents, they occasionally exhibit reasoning-action inconsistencies, such as violating target mutation budgets. Overall, confidence in the demonstrated fidelity gains is high due to cross-validation with an independent structural oracle, though physical wet-lab synthesis remains the necessary final validation step for generated candidates.

arXiv: 2604.17175

No sufficiently relevant recommendations were found.

Cover for RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design

Abstract

We introduce RosettaSearch, an inference-time multi-objective optimization approach for backbone conditioned protein sequence design. We use large language models (LLMs) as a generative optimizer within a search algorithm capable of controlled exploration and exploitation, using rewards computed from RosettaFold3, a structure prediction model, under a strict computational budget. In a large-scale evaluation, we apply RosettaSearch to 400 suboptimal sequences generated by LigandMPNN (a state-of-the-art model trained for protein sequence design), recovering high-fidelity designs that LigandMPNN's single-pass decoding fails to produce. RosettaSearch's designs show improvements in structural fidelity metrics ranging between 18% to 68%, translating to a 2.5x improvement in design success rate. We observe that these gains in success rate are robust when RosettaSearch-designed sequences are evaluated with an independent structure prediction oracle (Chai-1) and generalize across two distinct LLM families (o4-mini and Gemini-3), with performance scaling consistently with reasoning capability.

We further demonstrate that RosettaSearch improves the sequence fidelity of ProteinMPNN designs for de novo backbones from the Dayhoff atlas, showing that the approach generalizes beyond native protein structures to computationally generated backbones. We also demonstrate a multi-modal extension of RosettaSearch with vision-language models, where images of predicted protein structures are used as feedback to incorporate structural context to guide protein sequence generation. To our knowledge, this is the first large-scale demonstration that LLMs can serve as effective generative optimizers for backbone-conditioned protein sequence design, yielding systematic gains without any model retraining.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Design of RosettaSearch
  • 3.1 Metrics used to quantify structural fidelity of designed protein sequences
  • 3.2 Formulating effective reward and feedback
  • 3.2.1 Reward and Feedback
  • 3.2.2 Feedback to Prevent Reward Hacking
  • 3.3 Generative optimization with LLMs
  • 3.3.1 Iteration Schemes
  • 3.3.2 Multiobjective Optimization
  • 3.3.3 Meta Design
  • 3.4 Final design choices
  • 4 Datasets
  • 5 Experimental Results
  • 5.1 Evaluation metrics
  • 5.2 Random mutations baseline
  • 5.3 Cumulative results
  • 5.4 Priority search generally outperforms sequential revision
  • 5.5 RosettaSearch improves sequence fidelity
  • 5.6 Fidelity metrics improve over search steps
  • 5.7 Multi-modal extension with vision-language models
  • 5.8 RosettaSearch improves BindCraft binders
  • 5.9 Human experts’ thoughts on the LLM reasoning
  • 5.10 LLM experts’ opinion of reasoning
  • 5.11 Evaluation on the Dayhoff dataset
  • 5.12 Results from ablation of the contextual signals
  • 6 Further results
  • 7 Conclusion
  • 8 Acknowledgements
  • References
  • A Appendix
  • A.1 Calculation of C​αC\alpha-RMSD
  • A.2 Multi-objective optimization details
  • A.3 Where are the most updates made?
  • A.4 Characterizing the amino-acid substitution patterns
  • A.5 Does the reasoning indicate hallucination?

Knowls

  1. Knowl 1 — RosettaSearch Framework for Backbone-Conditioned Protein Sequence Design

    model/method

    RosettaSearch is an inference-time multi-objective generative optimization framework for backbone-conditioned protein sequence design. Rather than generating sequences in a single-pass feedforward manner or retraining deep generative models, RosettaSearch employs a large language model (LLM) or vision-language model (VLM) as a generative optimizer within an iterative search loop.

    Given a target backbone 3D atomic structure x∗x^* and an initial candidate sequence s0∈ALs_0 \in \mathcal{A}^L (over the 20 standard amino acid alphabet A\mathcal{A} of length LL), the system iteratively:

    1. Evaluates candidate sequences using an external biophysical structure prediction oracle ff, specifically RosettaFold3, in single-sequence mode (without multiple sequence alignment conditioning) to yield predicted 3D coordinates x^=f(s)\hat{x} = f(s).
    2. Computes global scalar rewards and residue-level local textual feedback identifying problematic structural regions (such as contiguous stretches of low confidence or high geometric deviation).
    3. Prompts the generative optimizer with the candidate sequence, structural feedback, task instructions, and soft biological constraints (e.g., sequence length, low k-mer redundancy, and preservation of catalytic/ligand-binding residues).
    4. Obtains proposed sequence mutations from the optimizer and updates a priority search queue to explore multiple refinement trajectories in parallel.

    The framework operates purely at inference time without requiring any gradient updates, policy retraining, or model fine-tuning when adapting to different reward criteria or backbones.

  2. Knowl 2 — Priority Search Algorithm for Sequence Optimization

    algorithm

    The Priority Search algorithm coordinates candidate evaluation and generative LLM proposals in a priority-based parallel search to avoid premature convergence to local optima.

    Input: Starting protein sequence ϕ0\phi_0, LLM optimizer O\mathcal{O}, structural evaluation guide G\mathcal{G}, number of exploration candidates KK, number of proposals per candidate NN, priority scoring function S(⋅)S(\cdot)
    Output: Best candidate designs {ϕt∗}\{\phi^*_t\}
    Initialize priority queue M←∅\mathcal{M} \leftarrow \emptyset
    Insert KK copies of (ϕ0,−∞)(\phi_0, -\infty) into M\mathcal{M} with maximal priority
    for iteration t=1,2,…t = 1, 2, \dots do
        Select best candidate ϕt∗\phi^*_t from c∗←arg⁡max⁡c∈MS(c)c^* \leftarrow \arg\max_{c \in \mathcal{M}} S(c)
        Select top-KK candidates from M\mathcal{M} as set C\mathcal{C}
        for each candidate c=(ϕ,score)∈Cc = (\phi, \text{score}) \in \mathcal{C} do
            Run ϕ\phi using guide G\mathcal{G} to extract updated score and feedback rr
            Collect rollout Rc←{(ϕ,score,r)}R_c \leftarrow \{(\phi, \text{score}, r)\}
        end for
        Initialize proposal set P←∅\mathcal{P} \leftarrow \emptyset
        for each candidate c∈Cc \in \mathcal{C} do
            Extract aggregated feedback f←⋃{r∣r∈Rc}f \leftarrow \bigcup \{r \mid r \in R_c\}
            for proposal index p=1p = 1 to NN do
                Generate sequence mutation ϕ′←O.step(ϕ,f)\phi' \leftarrow \mathcal{O}.\text{step}(\phi, f)
                Evaluate ϕ′\phi' with guide G\mathcal{G} to obtain score′\text{score}'
                Add candidate c′=(ϕ′,score′)c' = (\phi', \text{score}') to P\mathcal{P}
            end for
        end for
        for each candidate c∈C∪Pc \in \mathcal{C} \cup \mathcal{P} do
            Compute priority S(c)S(c)
            Insert cc into priority queue M\mathcal{M}
        end for
    end for
    return best designs {ϕt∗}\{\phi^*_t\}

    In standard execution, K=3K = 3 top candidate sequences are selected per round and N=2N = 2 mutant proposals are generated per candidate by the LLM optimizer, evaluated by RosettaFold3, and returned to the priority queue.

  3. Knowl 3 — Multi-Objective Reward and Structured Text Feedback Formulation

    equation

    The scalar reward R(s)R(s) used to guide sequence optimization is formulated as a weighted sum of three structural fidelity metrics:

    R(s)=wtm⋅TM(x^,x∗)+wp⋅pLDDT(x^)+wr⋅Cα-RMSD(x^,x∗)R(s) = w_{\text{tm}} \cdot \text{TM}(\hat{x}, x^*) + w_p \cdot \text{pLDDT}(\hat{x}) + w_r \cdot C_\alpha\text{-RMSD}(\hat{x}, x^*)

    where x^=f(s)\hat{x} = f(s) is the 3D structure predicted from the sequence ss by RosettaFold3 in single-sequence mode, x∗x^* is the target backbone atomic coordinates, pLDDT(x^)\text{pLDDT}(\hat{x}) is the average per-residue predicted Local Distance Difference Test scaled to [0,1][0, 1], TM(x^,x∗)\text{TM}(\hat{x}, x^*) is the length-normalized template-matching score, and Cα-RMSD(x^,x∗)C_\alpha\text{-RMSD}(\hat{x}, x^*) is the root mean square deviation of CαC_\alpha atoms scaled to [0,1][0, 1] after structural superposition via a remote-homology structural alignment algorithm. Values of Cα-RMSDC_\alpha\text{-RMSD} that exceed a predefined penalty threshold receive an explicit negative correction. The optimal weights are:

    wp=111,wtm=511,wr=511w_p = \frac{1}{11}, \quad w_{\text{tm}} = \frac{5}{11}, \quad w_r = \frac{5}{11}

    Alongside R(s)R(s), the LLM is provided with structured local textual feedback:

    • pLDDT Feedback: Contiguous residue index intervals where per-residue pLDDT is below adaptive thresholds (e.g., <35< 35).
    • Geometric Deviation Feedback: Contiguous residue intervals where local structural deviation from x∗x^* exceeds thresholds (e.g., >5.0 A˚> 5.0\text{ \AA} deviation or binned CαC_\alpha distances).
    • Anti-Reward-Hacking Constraints: Textual notifications flagging if a proposed sequence exhibits >5%> 5\% repeated 6-mers, deviates from the required sequence length LL, or mutates designated ligand-binding residues.
  4. Knowl 4 — Textual Step-Size Control and Best-of-N Search Strategy

    model/method

    To control the exploration-versus-exploitation trade-off during LLM-driven optimization, RosettaSearch constrains the number of amino acid residue substitutions allowed per optimization step by explicitly specifying a textual step-size regime in the LLM prompt:

    • Unconstrained: Arbitrary number of residue edits permitted.
    • Conservative: 11 to 22 residue changes per step.
    • Moderate: Up to 55 residue changes per step.
    • Aggressive: Up to 1010 residue changes per step.
    • Drastic: Up to 2020 residue changes per step.

    To ensure robustness against local optima and search non-monotonicity under a strict computational budget, a Best-of-NN strategy is deployed. For each target protein, RosettaSearch executes three independent Priority Search runs corresponding to three optimal step-size regimes: 'aggressive' (1010 edits), 'moderate' (55 edits), and 'unconstrained' (any edits). Each regime is allocated a budget of 2525 RosettaFold3 structure evaluations, yielding a total evaluation budget of N=3×25=75N = 3 \times 25 = 75 candidate structures per target backbone. The candidate achieving the highest composite reward across all NN evaluated sequences is selected as the final design.

  5. Knowl 5 — Evaluation Metrics for Protein Structural Fidelity and Success Rate

    definition

    Structural fidelity of a designed protein sequence ss of length LL folding into a target backbone structure x∗x^* is evaluated by predicting its single-sequence 3D coordinates x^=f(s)\hat{x} = f(s) and measuring:

    • ss-pLDDT (Single-Sequence pLDDT): Mean predicted Local Distance Difference Test across all residues without multiple sequence alignment (MSA) inputs, assessing sequence-intrinsic folding confidence in [0,1][0, 1].
    • TM-score: Template-matching score between x^\hat{x} and x∗x^*, assessing global fold topology similarity in [0,1][0, 1].
    • CαC_\alpha-RMSD: Root-mean-square deviation of corresponding CαC_\alpha atoms between x^\hat{x} and x∗x^* (in A˚\text{\AA}) computed via structural superposition.

    Design quality across a dataset is quantified by two composite binary success thresholds:

    • Success Rate-1: The fraction of designed sequences whose predicted structures simultaneously satisfy: pLDDT≥0.8andTM-score≥0.8\text{pLDDT} \ge 0.8 \quad \text{and} \quad \text{TM-score} \ge 0.8
    • Success Rate-2: A more stringent criterion requiring simultaneous high confidence, high global topology alignment, and low atomic deviation: pLDDT≥0.8,TM-score≥0.8,andCα-RMSD≤1.5 A˚\text{pLDDT} \ge 0.8, \quad \text{TM-score} \ge 0.8, \quad \text{and} \quad C_\alpha\text{-RMSD} \le 1.5\text{ \AA}
  6. Knowl 6 — Performance Comparison Across Initializations and Structural Oracles

    data/table

    On a curated dataset of approximately 400 monomeric PDB protein backbones where initial designs had low structural fidelity, RosettaSearch was evaluated across different initialization strategies and optimization objectives. Final designs generated using RosettaFold3 (RF3) feedback were re-evaluated with an independent structure prediction model, Chai-1, to verify out-of-distribution generalization.

    Initialization Objective / Method Success Rate (RF3) Success Rate (Chai-1)
    Baselines
    (a) Native monomer sequences Starting baseline 7.0% 2.4%
    (b) LigandMPNN designs Starting baseline 7.9% 3.1%
    (c) Random protein sequences Starting baseline 0.0% 0.0%
    LigandMPNN initialization Random mutations baseline 7.9% 3.1%
    RosettaSearch (o4-mini optimizer)
    (a) Native sequences Single obj. (ss-pLDDT) 7.5% 4.2%
    Multi-obj. (ss-pLDDT, TM-score) 9.1% 4.7%
    Multi-obj. (ss-pLDDT, TM, CαC_\alpha-RMSD) 8.4% 4.3%
    Cumulative 10.1% —
    (b) LigandMPNN designs Single obj. (ss-pLDDT) 18.4% 5.0%
    Multi-obj. (ss-pLDDT, TM-score) 17.1% 6.7%
    Multi-obj. (ss-pLDDT, TM, CαC_\alpha-RMSD) 20.5% 5.3%
    Cumulative 21.6% —

    When evaluated under the more stringent Success Rate-2 (pLDDT ≥0.8\ge 0.8, TM-score ≥0.8\ge 0.8, CαC_\alpha-RMSD ≤1.5 A˚\le 1.5\text{ \AA}), LigandMPNN baseline achieves 2.5%2.5\% (10/39410/394) while RosettaSearch multi-objective achieves 8.9%8.9\% (35/39435/394). Evaluating alternative LLMs on the LigandMPNN multi-objective task yields a 19.4%19.4\% RF3 success rate for Gemini-3 and 12.3%12.3\% for o3-mini (compared to 20.5%20.5\% for o4-mini).

  7. Knowl 7 — Priority Search vs. Sequential Revision Optimization

    data/table

    A comparison between Sequential Revision (a greedy ReAct-style iterative update loop) and Priority Search (parallel beam-like queue search) was conducted on the 400 PDB monomer dataset using o4-mini under identical evaluation budgets.

    Setting / Algorithm ss-pLDDT TM-score CαC_\alpha-RMSD ()Ä
    Single-Objective (ss-pLDDT)
    Baseline (Starting) 0.59 0.44 19.2
    Sequential Revision 0.670.67 (+17.0%) 0.540.54 (+35.5%) 14.614.6 (-20.0%)
    Priority Search 0.690.69 (+18.1 ±\pm 0.85%) 0.570.57 (+56.9 ±\pm 4.92%) 13.013.0 (-31.2 ±\pm 1.28%)
    (p=1.61×10−77p=1.61 \times 10^{-77}) (p=3.42×10−57p=3.42 \times 10^{-57}) (p=8.71×10−40p=8.71 \times 10^{-40})
    Multi-Objective (ss-pLDDT, TM-score, CαC_\alpha-RMSD)
    Baseline (Starting) 0.59 0.44 19.2
    Sequential Revision 0.710.71 (+24.2 ±\pm 1.1%) 0.610.61 (+65.1 ±\pm 5.2%) 11.711.7 (-32.6 ±\pm 1.3%)
    Priority Search 0.680.68 (+18.0 ±\pm 0.82%) 0.620.62 (+67.7 ±\pm 5.82%) 11.211.2 (-37.7 ±\pm 1.24%)
    (p=4.9×10−81p=4.9 \times 10^{-81}) (p=6.1×10−68p=6.1 \times 10^{-68}) (p=1.1×10−45p=1.1 \times 10^{-45})
    Cumulative Best-of-All +38.5% +86.4% -44.4%

    Priority Search achieves greater geometric accuracy (larger CαC_\alpha-RMSD reductions and higher TM-score gains) by maintaining a diversity of candidate trajectories and deferring commitment until downstream oracle scores are observed.

  8. Knowl 8 — Context and Feedback Ablation Against Random Mutation Baselines

    data/table

    To analyze what sources of guidance drive sequence design improvements, an ablation study was conducted starting from LigandMPNN candidate sequences (baseline success rate 8.5%8.5\% under sequential revision evaluation):

    Optimization Context / Configuration Success Rate (RF3)
    LigandMPNN initial sequences 8.5%
    + Reward-only sequential revision (scalar reward, no text feedback) 9.3%
    + Detailed textual feedback (residue ranges of low pLDDT / TM) 11.4%
    + Native sequence provided in prompt context 17.2%
    + Protein functional description included in prompt 17.4%
    + Retrieval-Augmented Generation (RAG from external database) 16.1%

    In an information-, compute-, and initialization-matched random mutations baseline—where identical residue-level feedback flags were provided but random amino acid substitutions were applied instead of LLM reasoning—the success rate remained at 7.9%7.9\% under RF3 (3.1%3.1\% under Chai-1), showing 0.0%0.0\% net improvement over the unoptimized LigandMPNN designs. This demonstrates that structural feedback alone is insufficient; the LLM's chemically targeted substitutions are essential.

  9. Knowl 9 — Fidelity Improvement on De Novo Backbones from the Dayhoff Atlas

    data/table

    RosettaSearch was evaluated on 275 de novo protein backbones generated unconditionally by RFDiffusion from the Dayhoff atlas, with initial sequences generated by ProteinMPNN. In this benchmark, no native sequence exists, so the LLM optimizer is provided zero reference sequence context and must rely strictly on RosettaFold3 structural feedback.

    Metric ProteinMPNN (Baseline) RosettaSearch (Optimized)
    ss-pLDDT (average) ↑\uparrow 0.81 0.85
    TM-score (average) ↑\uparrow 0.90 0.96
    CαC_\alpha-RMSD (average) ↓\downarrow 3.32 A˚3.32\text{ \AA} 1.78 A˚1.78\text{ \AA}
    Success Rate-1 (pLDDT ≥0.8\ge 0.8, TM ≥0.8\ge 0.8) 72.4% (199/275) 89.5% (246/275)
    Success Rate-2 (pLDDT ≥0.8\ge 0.8, TM ≥0.8\ge 0.8, RMSD ≤1.5 A˚\le 1.5\text{ \AA}) 45.5% (125/275) 67.3% (185/275)

    These results confirm that RosettaSearch generalizes to computationally generated, de novo protein backbones without requiring reference sequence templates.

  10. Knowl 10 — Refinement of High-Affinity Binders Generated by BindCraft

    data/table

    RosettaSearch was tested on 50 complex protein binders randomly sampled from the BindCraft dataset, which were pre-designed through AlphaFold2 fine-tuning pipelines to bind target enzymes. RosettaSearch optimized these candidates while retaining the annotated binding pocket residues interacting with the target.

    Metric BindCraft (Baseline) RosettaSearch (Optimized)
    ss-pLDDT ↑\uparrow 0.85 0.90
    TM-score ↑\uparrow 0.93 0.95
    CαC_\alpha-RMSD ↓\downarrow 1.37 A˚1.37\text{ \AA} 0.99 A˚0.99\text{ \AA}

    RosettaSearch improved structural fidelity metrics across all 48 evaluated binders while preserving the target enzyme binding interface.

  11. Knowl 11 — Reasoning Quality and Failure Modes of LLM Optimizers

    empirical result

    Evaluation of chain-of-thought reasoning produced by the o4-mini generative optimizer was conducted using both domain human experts and Gemini-3 as an automated independent judge across 50 proposal steps.

    • Logical Validity: Gemini-3 experts unanimously rated 56%56\% (28/50) of reasoning instances as logically sound and valid, and 44%44\% (22/50) as invalid.
    • Dominant Failure Modes:
      1. Mutation Budget Violation: The model explicitly planned a constrained set of edits (e.g., at most 10 substitutions) in its reasoning trace but emitted an output sequence with 15–30+ amino acid mutations.
      2. Hallucinated Residue Indexing: The model misidentified residue positions, made false assumptions about native amino acids, or reasoned about regions unflagged by the feedback.
      3. No-Op Generations: Emitting the identical starting sequence despite describing detailed mutation plans.
      4. Context Insensitivity: Proposing chemically sensible substitutions in isolation (e.g., mutating buried charged residues to nonpolar) where the residue was not actually buried in 3D space.

    Positional analysis across 3,308 successful trajectories revealed that edits were distributed across the protein with a slight preference for the C-terminal region (30.4%30.4\% in the 75–100% position bin vs. 22.3%22.3\%, 22.4%22.4\%, and 24.9%24.9\% in the other quartiles).

Coverage note — None was omitted; all contributed algorithms, mathematical formulations, core experimental benchmarks across native/de novo/binder datasets, ablations, and reasoning analyses are covered.

References

  1. 1.Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., Ronneberger, O., Willmore, L., Ballard, A. J., Bambrick, J., Bodenstein, S. W., Evans, D. A., Hung, C.-C., O’Neill, M., Reiman, D., Tunyasuvunakool, K., Wu, Z., Zemgulyt ˇ e, A., Arvaniti, E., Beattie, C., Bertolli, O., Bridgland, ˙ A., Cherepanov, A., Congreve, M., Cowen-Rivers, A. I., Cowie, A., Figurnov, M., Fuchs, F. B., Gladman, H., Jain, R., Khan, Y. A., Low, C. M. R., Perlin, K., Potapenko, A., Savy, P., Singh, S., Stecula, A., Thillaisundaram, A., Tong, C., Yakneen, S., Zhong, E. D., Zielinski, M., Zˇ´ıdek, A., Bapst, V., Kohli, P., Jaderberg, M., Hassabis, D., and Jumper, J. M. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature, pp. 1–3, May 2024. ISSN 1476-4687. doi: 10.1038/s41586-024-07487-w.
  2. 2.Agrawal, L. A., Tan, S., Soylu, D., Ziems, N., Khare, R., Opsahl-Ong, K., Singhvi, A., Shandilya, H., Ryan, M. J., Jiang, M., et al. Gepa: Reflective prompt evolution can outperform reinforcement learning. arXiv preprint arXiv:2507.19457, 2025.
  3. 3.Anishchenko, I., Pellock, S. J., Chidyausiku, T. M., Ramelot, T. A., Ovchinnikov, S., Hao, J., Bafna, K., Norn, C., Kang, A., Bera, A. K., et al. De novo protein design by deep network hallucination. Nature, 600(7889):547–552, 2021.
  4. 4.Chen, Y., Arkin, J., Hao, Y., Zhang, Y., Roy, N., and Fan, C. Prompt optimization in multi-step tasks (promst): Integrating human feedback and preference alignment. CoRR, 2024.
  5. 5.Cheng, C.-A., Nie, A., and Swaminathan, A. Trace is the next autodiff: Generative optimization with rich feedback, execution traces, and llms. Advances in Neural Information Processing Systems, 37: 71596–71642, 2024.
  6. 6.Corley, N., Mathis, S., Krishna, R., Bauer, M. S., Thompson, T. R., Ahern, W., Kazman, M. W., Brent, R. I., Didi, K., Kubaney, A., et al. Accelerating biomolecular modeling with atomworks and rf3. bioRxiv, 2025.
  7. 7.Dauparas, J., Anishchenko, I., Bennett, N., Bai, H., Ragotte, R. J., Milles, L. F., Wicky, B. I. M., Courbet, A., de Haas, R. J., Bethel, N., Leung, P. J. Y., Huddy, T. F., Pellock, S., Tischer, D., Chan, F., Koepnick, B., Nguyen, H., Kang, A., Sankaran, B., Bera, A. K., King, N. P., and Baker, D. Robust deep learning–based protein sequence design using ProteinMPNN. Science, 378(6615): 49–56, October 2022. doi: 10.1126/science.add2187.
  8. 8.Dauparas, J., Lee, G. R., Pecoraro, R., An, L., Anishchenko, I., Glasscock, C., and Baker, D. Atomic context-conditioned protein sequence design using LigandMPNN. Nature Methods, 22(4):717–723, April 2025. ISSN 1548-7105. doi: 10.1038/s41592-025-02626-1.
  9. 9.Discovery, C., Boitreaud, J., Dent, J., McPartlon, M., Meier, J., Reis, V., Rogozhnikov, A., and Wu, K. Chai-1: Decoding the molecular interactions of life, October 2024.
  10. 10.Ferruz, N., Schmidt, S., and Hocker, B. Protgpt2 is a deep unsupervised language model for protein ¨ design. Nature communications, 13(1):4348, 2022.
  11. 11.Ghareeb, A. E., Chang, B., Mitchener, L., Yiu, A., Szostkiewicz, C. J., Laurent, J. M., Razzak, M. T., White, A. D., Hinks, M. M., and Rodriques, S. G. Robin: A multi-agent system for automating scientific discovery. arXiv preprint arXiv:2505.13400, 2025.
  12. 12.Gottweis, J., Weng, W.-H., Daryin, A., Tu, T., Palepu, A., Sirkovic, P., Myaskovsky, A., Weissenberger, F., Rong, K., Tanno, R., et al. Towards an ai co-scientist. arXiv preprint arXiv:2502.18864, 2025.
  13. 13.Greisen, A., Cong, L., Ovchinnikov, S., et al. Reasoning models outperform standard language models in de novo protein design. In Open Conference of AI Agents for Science 2025.
  14. 14.Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., Vardhamanan, S., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., et al. Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714, 2023.
  15. 15.Korbeld, K. T., Viliuga, V., and Furst, M. Limitations of the refolding pipeline for de novo protein ¨ design. bioRxiv, pp. 2025–12, 2025.
  16. 16.Lutz, I. D., Wang, S., Norn, C., Courbet, A., Borst, A. J., Zhao, Y. T., Dosey, A., Cao, L., Xu, J., Leaf, E. M., et al. Top-down design of protein architectures with reinforcement learning. Science, 380 (6642):266–273, 2023.
  17. 17.Narayanan, S. M., Braza, J. D., Griffiths, R.-R., Bou, A., Wellawatte, G., Ramos, M. C., Mitchener, L., Rodriques, S. G., and White, A. D. Training a scientific reasoning model for chemistry. arXiv preprint arXiv:2506.17238, 2025.
  18. 18.Nie, A., Cheng, C.-A., Kolobov, A., and Swaminathan, A. The importance of directional feedback for llm-based optimizers. arXiv preprint arXiv:2405.16434, 2024.
  19. 19.Pacesa, M., Nickel, L., Schellhaas, C., Schmidt, J., Pyatova, E., Kissling, L., Barendse, P., Choudhury, J., Kapoor, S., Alcaraz-Serna, A., Cho, Y., Ghamary, K. H., Vinue, L., Yachnin, B. J., Wollacott, ´ A. M., Buckley, S., Westphal, A. H., Lindhoud, S., Georgeon, S., Goverde, C. A., Hatzopoulos, G. N., Gonczy, P., Muller, Y. D., Schwank, G., Swarts, D. C., Vecchio, A. J., Schneider, B. L., ¨ Ovchinnikov, S., and Correia, B. E. One-shot design of functional protein binders with BindCraft. Nature, 646(8084):483–492, October 2025. ISSN 1476-4687. doi: 10.1038/s41586-025-09429-6.
  20. 20.Pryzant, R., Iter, D., Li, J., Lee, Y. T., Zhu, C., and Zeng, M. Automatic prompt optimization with” gradient descent” and beam search. arXiv preprint arXiv:2305.03495, 2023.
  21. 21.Riesselman, A. J., Ingraham, J. B., and Marks, D. S. Deep generative models of genetic variation capture the effects of mutations. Nature methods, 15(10):816–822, 2018.
  22. 22.Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J., Ellenberg, J. S., Wang, P., Fawzi, O., et al. Mathematical discoveries from program search with large language models. Nature, 625(7995):468–475, 2024.
  23. 23.Skalse, J., Howe, N., Krasheninnikov, D., and Krueger, D. Defining and characterizing reward gaming. Advances in Neural Information Processing Systems, 35:9460–9471, 2022.
  24. 24.Stracquadanio, G. and Nicosia, G. Computational energy-based redesign of robust proteins. Computers & chemical engineering, 35(3):464–473, 2011.
  25. 25.Sun, H., Liu, X., Gong, Y., Zhang, Y., Jiang, D., Yang, L., and Duan, N. Allies: Prompting large language model with beam search. arXiv preprint arXiv:2305.14766, 2023.
  26. 26.Swanson, K., Wu, W., Nash, L. B., Pak, J. E., and Zou, J. The virtual lab: Ai agents design new sars-cov-2 nanobodies with experimental validation. biorxiv. 2024.
  27. 27.Tang, S., Zhang, Y., and Chatterjee, P. Peptune: De novo generation of therapeutic peptides with multi-objective-guided discrete diffusion. ArXiv, pp. arXiv–2412, 2025.
  28. 28.Tunyasuvunakool, K., Adler, J., Wu, Z., Green, T., Zielinski, M., Zˇ´ıdek, A., Bridgland, A., Cowie, A., Meyer, C., Laydon, A., et al. Highly accurate protein structure prediction for the human proteome. Nature, 596(7873):590–596, 2021.
  29. 29.Wang, X., Li, C., Wang, Z., Bai, F., Luo, H., Zhang, J., Jojic, N., Xing, E. P., and Hu, Z. Promptagent: Strategic planning with language models enables expert-level prompt optimization. arXiv preprint arXiv:2310.16427, 2023.
  30. 30.Watson, J. L., Juergens, D., Bennett, N. R., Trippe, B. L., Yim, J., Eisenach, H. E., Ahern, W., Borst, A. J., Ragotte, R. J., Milles, L. F., Wicky, B. I. M., Hanikel, N., Pellock, S. J., Courbet, A., Sheffler, W., Wang, J., Venkatesh, P., Sappington, I., Torres, S. V., Lauko, A., De Bortoli, V., Mathieu, E., Ovchinnikov, S., Barzilay, R., Jaakkola, T. S., DiMaio, F., Baek, M., and Baker, D. De novo design of protein structure and function with RFdiffusion. Nature, pp. 1–3, July 2023. ISSN 1476-4687. doi: 10.1038/s41586-023-06415-8.
  31. 31.Wei, A., Nie, A., Teixeira, T. S., Yadav, R., Lee, W., Wang, K., and Aiken, A. Improving parallel program performance with llm optimizers via agent-system interfaces. arXiv preprint arXiv:2410.15625, 2024.
  32. 32.Xia, Y., Jin, P., Xie, S., He, L., Cao, C., Luo, R., Liu, G., Wang, Y., Liu, Z., Chen, Y.-J., et al. Nature language model: Deciphering the language of nature for scientific discovery. arXiv preprint arXiv:2502.07527, 2025.
  33. 33.Xu, W., Banburski-Fahey, A., and Jojic, N. Reprompting: Automated chain-of-thought prompt inference through gibbs sampling. arXiv preprint arXiv:2305.09993, 2023.
  34. 34.Xue, F., Kubaney, A., Guo, Z., Min, J. K., Liu, G., Yang, Y., and Baker, D. Improving Protein Sequence Design through Designability Preference Optimization, May 2025.
  35. 35.Yang, K. K., Wu, Z., and Arnold, F. H. Machine-learning-guided directed evolution for protein engineering. Nature methods, 16(8):687–694, 2019.
  36. 36.Yang, K. K., Alamdari, S., Lee, A. J., Kaymak-Loveless, K., Char, S., Brixi, G., Domingo-Enrich, C., Wang, C., Lyu, S., Fusi, N., et al. The dayhoff atlas: scaling sequence diversity for improved protein generation. bioRxiv, pp. 2025–07, 2025.
  37. 37.Yin, R., Feng, B. Y., Varshney, A., and Pierce, B. G. Benchmarking alphafold for protein complex modeling reveals accuracy determinants. Protein Science, 31(8):e4379, 2022.
  38. 38.Yuksekgonul, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., Guestrin, C., and Zou, J. Textgrad: Automatic” differentiation” via text. arXiv preprint arXiv:2406.07496, 2024.
  39. 39.Zhang, Y. and Skolnick, J. Tm-align: a protein structure alignment algorithm based on the tm-score. Nucleic acids research, 33(7):2302–2309, 2005.
  40. 40.Zhou, X., Xue, D., Chen, R., Zheng, Z., Wang, L., and Gu, Q. Antigen-specific antibody design via direct energy-based preference optimization. Advances in Neural Information Processing Systems, 37:120861–120891, 2024.

Citation

MLA
Kshirsagar, M., et al. “RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design”. arXiv, 2026, http://arxiv.org/abs/2604.17175v2.
APA
Kshirsagar, M., Nie, A., Cheng, C.-A., Xue, F., Dodhia, R., Ferres, J. L., Yang, K. K., & DiMaio, F. (2026). RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design. arXiv. http://arxiv.org/abs/2604.17175v2
Chicago
Kshirsagar, M., A. Nie, C.-A. Cheng, et al. 2026. “RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design”. arXiv. http://arxiv.org/abs/2604.17175v2.
Harvard
Kshirsagar, M. et al. (2026) “RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.17175v2.
Vancouver
1. Kshirsagar M, Nie A, Cheng C-A, Xue F, Dodhia R, Ferres JL, Yang KK, DiMaio F (2026) RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design. arXiv

BibTeX

@article{kshirsagar2026rosettasearch,
  title = {RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design},
  author = {Kshirsagar, Meghana and Nie, Allen and Cheng, Ching-An and Xue, Fanglei and Dodhia, Rahul and Ferres, Juan Lavista and Yang, Kevin K. and DiMaio, Frank},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.17175v2},
  eprint = {2604.17175}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/