RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design
Meghana KshirsagarChing-An ChengAllen NieFanglei XueRahul DodhiaJuan M. Lavista FerresKevin Kaichuang YangF. Dimaio
Introduces RosettaSearch, an inference-time optimization framework that uses large language models as generative search agents to rescue failed ProteinMPNN and LigandMPNN candidates, boosting protein sequence design success rates by 2.5 times without requiring model retraining.
Protein sequence design—the task of generating amino acid sequences that reliably fold into specific three-dimensional target structures—is essential for advancing drug discovery, enzyme engineering, and therapeutic development. However, state-of-the-art computational design tools frequently produce candidate sequences with low structural fidelity. Because these tools rely on single-pass autoregressive generation, they lack self-correction mechanisms to detect or fix structural errors. Consequently, poorly folded designs often proceed to expensive wet-lab testing where they fail, creating a major cost and timeline bottleneck. The article presents RosettaSearch, an inference-time optimization framework that uses large language models as generative optimizers to iteratively refine protein sequences based on direct structural feedback, eliminating the need for expensive model retraining or task-specific fine-tuning.
The main objective of the article is to demonstrate that frontier language models can serve as generative optimizers within a structured, priority-based search algorithm to systematically improve the structural fidelity of protein designs under strict computational budgets. To achieve this, the authors evaluated RosettaSearch across approximately 400 monomeric proteins from the Protein Data Bank (PDB) with low baseline fidelity, 275 de novo computational backbones from the Dayhoff atlas, and 50 computationally optimized binders from BindCraft. The framework couples candidate sequence generation with structural predictions from RosettaFold3, which provides scalar rewards and residue-level text annotations identifying low-confidence and misaligned regions. The optimization was evaluated under tight evaluation budgets (up to 75 structure calls per target) and independently validated using a separate structure prediction model, Chai-1, to guard against model-specific evaluation bias.
The key findings show that RosettaSearch substantially improves protein design quality across diverse benchmarks. First, when applied to suboptimal designs from LigandMPNN, RosettaSearch improved structural similarity metrics by 18% to 68%, driving a 2.5-fold increase in the design success rate (from 7.9% to 20.5% under multi-objective priority search). Under a more stringent multi-metric threshold, the success rate more than tripled from 2.5% to 8.9%. Second, these gains proved robust when evaluated with the independent Chai-1 predictor, confirming that performance improvements reflect genuine structural enhancement rather than overfitting to the primary prediction tool. In contrast, an information-matched baseline using random mutations failed to improve starting designs, proving that the language model's reasoning is essential for proposing effective, chemically informed edits. Third, RosettaSearch demonstrated strong generalization across other problem settings: it raised the success rate of de novo Dayhoff backbone designs from 72.4% to 89.5% without any reference sequence context, and consistently improved structural fidelity across 48 complex protein binders from BindCraft.
These findings indicate that generative search can serve as a highly flexible, plug-and-play optimization layer on top of existing sequence generation pipelines. Because RosettaSearch operates entirely at inference time and accommodates modular reward functions, engineering teams can optimize for multiple structural and functional constraints without undergoing costly model retraining or dataset curation. In addition, comparative evaluations across different model families (o4-mini, o3-mini, and Gemini-3) showed that optimization performance scales directly with underlying reasoning capability, highlighting that general-purpose reasoning models can be effectively harnessed for complex biological design.
Organizations developing computational protein pipelines should consider integrating inference-time generative search to refine candidate pools before committing to wet-lab synthesis. When implementing this approach, teams should utilize parallel priority search over greedy sequential revision, as parallel exploration significantly reduces premature convergence. Furthermore, systems should combine global numerical scores with residue-level textual feedback to provide actionable spatial context, while soft textual constraints should be enforced to prevent common optimization failure modes such as repetitive sequence motifs or length drift.
Despite these strong results, several limitations should be noted. RosettaSearch focuses on optimizing a single high-fidelity sequence per backbone rather than generating large, diverse candidate libraries, and the optimization process remains computationally bounded by the speed of structure prediction models (averaging approximately 30.6 minutes per protein). Furthermore, expert analysis revealed that while the language models propose valid design intents, they occasionally exhibit reasoning-action inconsistencies, such as violating target mutation budgets. Overall, confidence in the demonstrated fidelity gains is high due to cross-validation with an independent structural oracle, though physical wet-lab synthesis remains the necessary final validation step for generated candidates.
- Paper: Learning inverse folding from millions of predicted structures, Chloe Hsu et al. (2022). Establishes modern deep learning paradigms for inverse folding and backbone-conditioned protein sequence design that RosettaSearch aims to optimize at inference time.
- Paper: Highly accurate protein structure prediction with AlphaFold, John Jumper et al. (2021). Provides the foundational breakthrough in machine learning-based protein structure prediction that underpins the structural reward models and oracles used to guide sequence optimization.
- Paper: Biological Sequence Design with GFlowNets, Moksh Jain et al. (2022). Introduces generative search and active learning frameworks for biological sequence design, establishing the core problem formulation of navigating complex sequence spaces with reward feedback.
- Paper: ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design, Pascal Notin et al. (2023). Defines standardized benchmarks and validation methodologies for evaluating protein fitness, zero-shot mutation scoring, and computational design fidelity.
- Paper: Symbolic Regression with a Learned Concept Library, Arya Grayeli et al. (2024). Demonstrates how large language models can act as generative operators to guide search and optimization algorithms across complex discrete hypothesis spaces.
No sufficiently relevant recommendations were found.
