Built independently by an author, for readers. Read the story and support ChapterPal

keyword

protein structure prediction

Protein structure prediction is the computational process of determining the three-dimensional spatial conformation of a protein from its one-dimensional amino acid sequence. Because a protein's biological function is largely dictated by its folded geometric shape, predicting its native structure helps elucidate molecular mechanisms while addressing the time and resource constraints of experimental techniques such as X-ray crystallography and cryo-electron microscopy. The process encompasses modeling various levels of organization, including secondary structural motifs as well as complex tertiary and quaternary assemblies. Modern approaches utilize template-based homology modeling, physics-based energy simulations, and deep learning architectures trained on evolutionary sequence alignments and experimental structural databases, making it an essential foundation for structural bioinformatics, drug discovery, and de novo protein design.

4 items

RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design

RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design

Meghana Kshirsagar, Ching-An Cheng, Allen Nie, Fanglei Xue, Rahul Dodhia, Juan M. Lavista Ferres, Kevin Kaichuang Yang, F. Dimaio

OrganizationsDepartment of BiochemistryGoogleInstitute for Protein DesignMicrosoftUniversity of Washington

Why you should read this

Introduces RosettaSearch, an inference-time optimization framework that uses large language models as generative search agents to rescue failed ProteinMPNN and LigandMPNN candidates, boosting protein sequence design success rates by 2.5 times without requiring model retraining.

We introduce RosettaSearch, an inference-time multi-objective optimization approach for backbone conditioned protein sequence design. We use large language models (LLMs) as a generative optimizer within a search algorithm capable of controlled exploration and exploitation, using rewards computed from RosettaFold3, a structure prediction model, under a strict computational budget. In a large-scale evaluation, we apply RosettaSearch to 400 suboptimal sequences generated by LigandMPNN (a state-of-the-art model trained for protein sequence design), recovering high-fidelity designs that LigandMPNN's single-pass decoding fails to produce. RosettaSearch's designs show improvements in structural fidelity metrics ranging between 18% to 68%, translating to a 2.5x improvement in design success rate. We observe that these gains in success rate are robust when RosettaSearch-designed sequences are evaluated with an independent structure prediction oracle (Chai-1) and generalize across two distinct LLM families (o4-mini and Gemini-3), with performance scaling consistently with reasoning capability. We further demonstrate that RosettaSearch improves the sequence fidelity of ProteinMPNN designs for de novo backbones from the Dayhoff atlas, showing that the approach generalizes beyond native protein structures to computationally generated backbones. We also demonstrate a multi-modal extension of RosettaSearch with vision-language models, where images of predicted protein structures are used as feedback to incorporate structural context to guide protein sequence generation. To our knowledge, this is the first large-scale demonstration that LLMs can serve as effective generative optimizers for backbone-conditioned protein sequence design, yielding systematic gains without any model retraining.

Added

2026-09-29

Protein Representation Learning by Geometric Structure Pretraining

Protein Representation Learning by Geometric Structure Pretraining

Zuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan, Aurélie C. Lozano, Payel Das, Jian Tang

OrganizationsCIFARHEC MontréalIBMMila – Québec Artificial Intelligence InstituteUniversité de MontréalUniversity of Cambridge

Why you should read this

Develops a geometric pretraining framework for protein graphs using multiview contrastive learning and self-prediction tasks, matching or exceeding sequence-based language models on function and fold classification while requiring far less training data.

Learning effective protein representations is critical in a variety of tasks in biology such as predicting protein function or structure. Existing approaches usually pretrain protein language models on a large number of unlabeled amino acid sequences and then finetune the models with some labeled data in downstream tasks. Despite the effectiveness of sequence-based approaches, the power of pretraining on known protein structures, which are available in smaller numbers only, has not been explored for protein property prediction, though protein structures are known to be determinants of protein function. In this paper, we propose to pretrain protein representations according to their 3D structures. We first present a simple yet effective encoder to learn the geometric features of a protein. We pretrain the protein graph encoder by leveraging multiview contrastive learning and different self-prediction tasks. Experimental results on both function prediction and fold classification tasks show that our proposed pretraining methods outperform or are on par with the state-of-the-art sequence-based methods, while using much less pretraining data. Our implementation is available at this https URL.

Added

2026-09-26

Protein Conformation Generation via Force-Guided SE(3) Diffusion Models

Protein Conformation Generation via Force-Guided SE(3) Diffusion Models

Yan Wang, Lihao Wang, Yuning Shen, Yiqun Wang, Huizhuo Yuan, Yue Wu, Quanquan Gu

OrganizationsByteDanceTongji UniversityUniversity of California, Los Angeles

Why you should read this

Proposes ConfDiff, a force-guided SE(3) diffusion model that integrates molecular mechanics energy priors into the sampling process to generate diverse, low-energy protein conformations consistent with the Boltzmann distribution without relying on molecular dynamics training data.

The conformational landscape of proteins is crucial to understanding their functionality in complex biological processes. Traditional physics-based computational methods, such as molecular dynamics (MD) simulations, suffer from rare event sampling and long equilibration time problems, hindering their applications in general protein systems. Recently, deep generative modeling techniques, especially diffusion models, have been employed to generate novel protein conformations. However, existing score-based diffusion methods cannot properly incorporate important physical prior knowledge to guide the generation process, causing large deviations in the sampled protein conformations from the equilibrium distribution. In this paper, to overcome these limitations, we propose a force-guided SE(3) diffusion model, CONFDIFF, for protein conformation generation. By incorporating a force-guided network with a mixture of data-based score models, CONFDIFF can generate protein conformations with rich diversity while preserving high fidelity. Experiments on a variety of protein conformation prediction tasks, including 12 fast-folding proteins and the Bovine Pancreatic Trypsin Inhibitor (BPTI), demonstrate that our method surpasses the state-of-the-art method.

Added

2026-09-26

Highly accurate protein structure prediction with AlphaFold

Highly accurate protein structure prediction with AlphaFold

John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew W. Senior, Koray Kavukcuoglu, Pushmeet Kohli, Demis Hassabis

OrganizationsGoogleSeoul National University

Why you should read this

Presents a deep learning architecture that integrates physical and evolutionary constraints to provide the first computational method capable of predicting protein three-dimensional structures with atomic accuracy across diverse sequences.

Proteins are essential to life, and understanding their structure can facilitate a mechanistic understanding of their function. Through an enormous experimental effort, the structures of around 100,000 unique proteins have been determined, but this represents a small fraction of the billions of known protein sequences. Structural coverage is bottlenecked by the months to years of painstaking effort required to determine a single protein structure. Accurate computational approaches are needed to address this gap and to enable large-scale structural bioinformatics. Predicting the three-dimensional structure that a protein will adopt based solely on its amino acid sequence—the structure prediction component of the ‘protein folding problem’—has been an important open research problem for more than 50 years. Despite recent progress, existing methods fall far short of atomic accuracy, especially when no homologous structure is available. Here we provide the first computational method that can regularly predict protein structures with atomic accuracy even in cases in which no similar structure is known. We validated an entirely redesigned version of our neural network-based model, AlphaFold, in the challenging 14th Critical Assessment of protein Structure Prediction (CASP14), demonstrating accuracy competitive with experimental structures in a majority of cases and greatly outperforming other methods. Underpinning the latest version of AlphaFold is a novel machine learning approach that incorporates physical and biological knowledge about protein structure, leveraging multi-sequence alignments, into the design of the deep learning algorithm.

Added

2026-05-27

Creative Commons License