
For a decade, the central promise of CRISPR-based gene editing has been clear: rewrite the DNA instructions that cause disease. But there has always been a catch. The molecular machines that do the editing (proteins borrowed from bacteria and retooled for human cells) do not always hit their target. They sometimes edit the wrong gene, creating unintended mutations that could cause cancer or other harm.
Researchers have thrown a lot of computing power at this problem. They have solved the three-dimensional structures of editing proteins bound to DNA. They have simulated how those proteins flex and bend. They have cataloged thousands of off-target sites. And yet the off-target problem has remained stubbornly difficult to fix by design, because structure alone is a poor teacher of what happens when a protein meets the wrong DNA sequence.
A landmark study published in Nature this month changes that. A team led by Haopeng Meng and Zhijie Lei at the Chinese Academy of Sciences has introduced ContactSeek, an AI framework that uses DeepMind’s AlphaFold3 not simply to predict the shape of gene-editing proteins, but to measure something far subtler: how confidently each atom in the protein expects to interact with each atom in the DNA. That confidence signal, the probability of contact between amino acids and nucleotides, turns out to be exquisitely sensitive to the difference between a perfect genetic match and a dangerous near-miss.
The work, published July 23 in Nature (DOI: 10.1038/s41586-026-10794-z), represents a fundamental shift in how structural biologists think about specificity in molecular machines.
Why structure is not enough
When a protein like Cas9 binds to DNA, the complex adopts a shape. That shape can be solved by X-ray crystallography, cryo-electron microscopy, or predicted by AlphaFold. But a protein bound to its correct DNA target and one bound to a mismatched target may look nearly identical in 3D structure. The difference is not in the pose, but in the subtle loosening of atomic contacts: residues that no longer touch their intended partners.
Standard structural biology has been largely blind to these differences. “The 3D structure tells you where atoms are, but it does not tell you how much they want to be there,” Meng said in a briefing. That wanting, the energetic preference for one interaction over another, is encoded in the ensemble of possible configurations a protein-DNA complex can explore. And that ensemble is precisely what AlphaFold3 can sample.
AlphaFold3, released in 2024, is the latest iteration of the protein structure prediction system from Google DeepMind. Unlike its predecessors, which predicted a single most-likely structure, AlphaFold3 produces a distribution of possibilities. Each pair of amino acids or nucleotides receives a contact probability: a number between 0 and 1 representing how likely the model thinks those residues are to interact. These probabilities are a byproduct of AlphaFold3’s diffusion-based architecture, which generates many plausible structural models internally before settling on its best prediction.
The ContactSeek team realized that these probabilities were not noise, but a rich signal. “AlphaFold3 does not just predict one structure,” Lei explained. “It explores an entire landscape of possible structures. The contact probabilities capture how much the model has to ‘work’ to fit each interaction into a coherent whole. When the fit is bad, the confidence drops.”
Mapping the off-target landscape
To test this idea, the team focused on adenine base editors (ABEs), a widely used class of CRISPR tools that convert one DNA letter (A) to another (G) without breaking both strands of the double helix. ABEs are built from Cas9 fused to a deaminase enzyme, and while less disruptive than classic cuts, they still produce off-target edits that limit therapeutic use.
The researchers first performed genome-wide off-target screens for several ABE variants, identifying hundreds of sites where the editors produced unintended edits. They then used AlphaFold3 to predict the structure of each editor bound to each off-target DNA sequence, plus the correct target. For every complex, they extracted the contact probability matrix, roughly 350,000 pairwise interaction scores.
The result was a dataset mapping not just the shape of each complex, but the confidence landscape of every contact. And when the team compared on-target and off-target complexes, the contact probabilities revealed clear patterns that the 3D structures alone did not.
“There were complexes whose overall structure was essentially identical,” said Meng. “AlphaFold3 predicted the same fold, the same atomic positions within a few tenths of a nanometer. But the contact probabilities told a completely different story. Interactions that were highly confident in the on-target complex became ambiguous in the off-target complex, even though the atoms were in the same places.”
The key insight is that contact probability captures a property that static structure misses: the degree of frustration in the system. In the on-target complex, every residue settles into a comfortable interaction; the contact probabilities are uniformly high. In the off-target complex, some contacts are forced: the geometry may be acceptable, but the energetic fit is poor, and AlphaFold3’s ensemble sampling reflects this as reduced probability.
From insight to design
With the contact probability maps in hand, the team identified specific amino acid residues in the Cas9 domain of their ABEs that showed the largest drops in contact probability between on-target and off-target complexes. These residues, they reasoned, tolerated mismatched DNA and enabled off-target editing.
ContactSeek then guided protein engineering. By mutating the identified residues to variants that increased the gap between on-target and off-target contact probabilities, the team could systematically tighten specificity.
The results were dramatic. In human cell assays, the redesigned ABEs reduced off-target editing by 69 to 83 percent across multiple guide RNAs and genomic sites, while maintaining full on-target activity. The same approach generalized to different ABE variants and guide sequences, suggesting the method captures a fundamental principle of protein-DNA recognition.
“Crucially, the mutations we made were not in the active site or in the regions that directly contact the DNA bases,” said Lei. “They were in the backbone, in structural residues that modulate the flexibility of the complex. Contact probability let us see that these peripheral residues were the real gatekeepers of specificity, even though no structural method had flagged them before.”
Beyond base editing
ContactSeek is not limited to adenine base editors. The framework applies to any Cas9-based tool, including classic nuclease-type CRISPR, prime editors, and CRISPRi/a systems. The team has already demonstrated results with Cas12-family proteins as well.
The computational cost, while substantial, is dropping fast. Each ContactSeek analysis requires running AlphaFold3 on hundreds to thousands of protein-DNA complexes. But advances in hardware and model efficiency have brought the cost down to a few thousand dollars per target protein, a bargain compared to the trial-and-error it replaces.
“This changes the pipeline,” said David Liu, a gene-editing pioneer at the Broad Institute who was not involved in the study. “Instead of making variants and testing them blindly, you can now look at the contact probability landscape and ask: where is the protein most uncertain?”
For patients waiting for gene therapies that are both effective and safe, that shift from blind engineering to informed design may be the most important advance of all.
Reference: Meng, H., Lei, Z. et al. ContactSeek: An AlphaFold3-based framework for engineering specificity in gene-editing proteins. Nature (2026). DOI: 10.1038/s41586-026-10794-z

