Schema mit zwei gegenläufigen Pfeilen: Links eine Kette aus Aminosäure-Bausteinen, rechts eine gefaltete dreidimensionale Proteinstruktur. Der obere Pfeil von links nach rechts ist mit Strukturvorhersage beschriftet, der untere Pfeil von rechts nach links mit Inverse Folding.

Inverse Folding

Inverse folding refers to methods that compute a matching set of building instructions made of amino acids for a desired three-dimensional shape of a protein. It is thus the reverse of the classic question of what shape a given set of building instructions will take.

Proteins are tiny tools inside living organisms. They consist of a long chain of building blocks, the amino acids. This chain folds itself into a specific spatial shape. And the shape determines what the protein can do: digesting, transporting oxygen, fighting off viruses. The usual question is: What shape emerges from a given chain? Inverse folding turns the question around. You specify a desired shape, and a computer program searches for a chain of building blocks that folds into exactly that shape.

Why researchers want to design proteins backwards

Nature has produced only a tiny fraction of all conceivable proteins. Mathematically, there are more possible chains than atoms in the universe. Anyone who can freely design a shape and then compute the matching chain is no longer dependent on existing natural products. This opens the door to medicines, vaccines, and enzymes that have never existed before.

One concrete goal is so-called binders. These are small proteins that attach to a specific site on a pathogen and block it. The shape of the target site is often known very precisely. What’s needed then is a counterpart that fits exactly into it, much like a key in a lock. Inverse folding provides the building instructions for this key.

Industry is also interested. Enzymes are proteins that speed up chemical reactions. A custom-tailored enzyme could break down plastic or make manufacturing a drug cheaper. In the past, such proteins were optimized through years of trial and error in the lab. Today, a model proposes hundreds of candidates in seconds, which are then tested.

From scaffold to building instructions

The starting point is usually a scaffold, known in technical jargon as the backbone. This is the rough spatial course of the chain, without specifying which building block sits at which position. The model goes through position after position and estimates which of the twenty natural amino acids fits best there. In doing so, it looks at the neighborhood: which sites are spatially close, and how are they oriented relative to one another?

Such models are trained on tens of thousands of known protein structures from public databases. There, both pieces of information are known: the shape and the corresponding chain. The model learns statistical patterns from this. For example, that water-repellent building blocks tend to sit inside a structure and water-friendly ones on the outside. Nobody programs these rules in by hand.

A common misconception is that there is exactly one correct solution. The opposite is true: very many different chains lead to the same shape. Models therefore output probabilities and provide entire lists of suggestions. Afterward, a folding program checks whether the suggestion actually results in the desired shape. Only then is it tested in the lab, because calculations do not replace an experiment.

Where the term shows up in the news

The field became well known through AlphaFold by Google DeepMind, which predicts shapes from chains. Inverse folding is the reverse direction of that. The best-known model of this kind is called ProteinMPNN and comes from David Baker’s lab in Seattle. In 2024, Baker received the Nobel Prize in Chemistry together with two DeepMind researchers, in part for precisely this work.

In business news, the term usually appears in the context of biotech companies. Companies such as Generate Biomedicines or EvolutionaryScale are raising capital in the hundreds of millions to design drugs on a computer. Traditional pharmaceutical companies are also buying into such methods. When people talk about AI drugs, protein design is often behind it.

For everyday life, the topic is still abstract, since so far few products have made it to market. The path from a designed protein to an approved drug takes years. Still, it’s worth knowing the term. It marks a shift: biology is increasingly becoming a discipline in which things are designed rather than merely discovered.

Related Products

Latest News

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.