AI-redesigned starting points help evolution reach proteins nature never could

Evolution is a blind climber, and the landscape it navigates is treacherous.

A protein under evolutionary pressure can only mutate one step at a time. Each change that makes it better at its job must also leave it stable enough to fold, soluble enough to function, and cooperative enough to survive in a cell. Natural selection cannot skip across valleys. It can only walk the ridge lines. So when scientists push an enzyme to evolve a new ability in the lab, they do not start from scratch. They start from a wild-type protein, one that nature already made stable and functional over millions of years. And this is where the blind climber can get stuck.

A new study from the lab of David Liu at the Broad Institute of MIT and Harvard, published in Nature on July 22, 2026, shows that this conventional wisdom has a hidden weakness: the wild-type starting point is itself a constraint. By using artificial intelligence to redesign that starting point before evolution begins, the researchers cracked open parts of the protein fitness landscape that evolution alone could never reach. The results upend a basic assumption of directed evolution and point toward a general strategy for making far better enzymes than either nature or AI can manage on their own.

The work, led by Benjamin Krasnow, Eric Xu, and Jiale Zhang, centers on a family of proteases derived from botulinum neurotoxin (BoNT). These enzymes cut other proteins at specific sequences, and researchers have spent years engineering them to target therapeutically relevant proteins instead of their natural nerve-cell substrates. The standard approach is directed evolution: introduce random mutations, screen for improved activity on the desired target, repeat. It works, but only up to a point. The authors found that the usual wild-type starting scaffolds had a serious problem: as they accumulated mutations during evolution, they lost stability.

Support independent reporting built on evidence, transparency, and scientific rigor.

Support 1ban.news

This is the climber’s dilemma. The mutations that improve an enzyme’s ability to recognize a new target tend to destabilize its structure. The protein becomes floppy, misfolds, or gets degraded before it can do its job. The wild-type sequence, for all its natural robustness, is not designed to tolerate the kinds of mutations that laboratory evolution demands. It evolved for its own ecological niche, not for therapeutic engineering.

The AI solution came from ProteinMPNN, a deep-learning model originally developed by the Baker lab for inverse protein design. Given a protein backbone structure, ProteinMPNN predicts which amino acid sequences are most likely to fold into it. The Liu team used it not to design entirely new proteins, but to redesign the existing BoNT scaffolds: keep the backbone geometry, but let the AI suggest alternative sequences that would stabilize the structure and tolerate more mutations.

The redesigned enzymes were dramatically more mutationally robust than their wild-type parents. When the researchers introduced random mutations across the board, the AI-redesigned starting points stayed folded and functional far longer. They had, in effect, been pre-optimized for the evolutionary journey ahead.

The real test came when the team pitted AI-redesigned starting points against wild-type starting points in head-to-head PACE campaigns. PACE (phage-assisted continuous evolution) is a technology developed by the Liu lab that links a protein’s activity to the replication of a bacteriophage, allowing evolution to run for hundreds of generations in days, with no human intervention. Each cycle of mutation and selection is driven by the enzyme’s performance on the target substrate, not by a researcher picking winners by hand.

In every comparison, the AI-redesigned starting points evolved more active proteases than the wild-type starting points. The gap was not marginal. Some of the best evolved variants from redesigned backgrounds had activities that the wild-type-evolved variants simply could not reach. By sequencing the evolutionary trajectories, the team discovered that the redesigned starting points led to sequences that were inaccessible from the wild-type background. It was not that the AI-redesigned enzymes evolved faster along the same path. They found entirely different paths, through entirely different sequence space, to higher peaks that the wild-type climber could not see.

This is the most striking finding of the paper. Evolution from the wild-type starting point is not a slower version of the same process. It is a fundamentally constrained process. The wild-type sequence sits at a particular spot in the fitness landscape, and the mutational paths radiating outward from that spot lead only to certain neighborhoods. The AI-redesigned sequences sit at different spots, radiating into different neighborhoods. By choosing the right starting point, the researchers could access regions of the landscape that were simply unreachable from the wild-type origin.

The team demonstrated the therapeutic potential of this approach by evolving a BoNT/E protease to selectively cleave ataxin-2, a protein whose aggregation is implicated in neurodegenerative diseases including spinocerebellar ataxia type 2 and amyotrophic lateral sclerosis. Selective proteolysis of disease-associated proteins is a major goal of protein engineering: the challenge is to make an enzyme that cuts the right target while leaving everything else alone.

Evolved from a wild-type starting point, the best BoNT/E variant for ataxin-2 cleavage had modest specificity. Evolved from an AI-redesigned starting point, the best variant achieved more than 79-fold better specificity for ataxin-2 over off-target substrates, while maintaining higher catalytic efficiency and superior thermostability. The redesigned-evolved variant was simultaneously more selective, more active, and more stable than the wild-type-evolved variant. These three properties are usually in tension; improvements in one often come at the cost of another. The AI starting point allowed evolution to break that trade-off.

What makes the work broadly significant is its generality. The pipeline has only three steps: take a protein structure, run ProteinMPNN to generate redesigned sequences, and subject the best candidates to continuous evolution against the target of interest. No custom algorithms, no protein-specific tuning, no prior knowledge of the evolutionary constraints. The AI redesign step requires no experimental work. It is a computation that runs on a laptop in minutes.

The implication is that for any protein engineering problem where evolution is part of the solution, the starting point matters as much as the evolutionary process itself. Giving evolution a better launch pad, rather than starting from the wild-type sequence that nature happened to leave us, may become the standard approach. The blind climber gets a headlamp.

Protein engineers have spent decades developing smarter ways to mutate, select, and evolve. This work suggests that the most impactful intervention may be the simplest: change where you start. The AI does not replace evolution. It just points evolution in a direction that looks unpromising from the old starting point but turns out to contain a wealth of hidden function. The wild-type sequence is not sacred. It is just the first draft. The second draft, written by a model that can see across the entire landscape at once, gives evolution the running start it needs to reach proteins that nature never could.

Scroll to Top