Machine learning-based computational protein structure prediction tools have revolutionized the study of structural biology. After training on massive datasets of experimentally determined protein structures, tools such as AlphaFold can accurately predict the 3D structure of a protein given only its amino acid sequence. Questions that once required weeks or months of protein preparation, data collection, and complex analysis can now be answered with just the click of a button.
Proteins come in many forms, which allow them to perform a wide range of crucial biological functions. Hollow cylindrical proteins serve as pores for transport of small molecules across membranes. Y-shaped antibodies recognize specific pathogenic antigens and direct immune responses. Signaling proteins have precise key-like structures to ensure they transmit their message to the correct lock-like receptor on the right cells.
Proteins that bind DNA have devoted domains that can interact with nucleic acids. They are essential for DNA replication and repair, regulation and execution of transcription, genome organization, and more. Structural analysis can reveal the details of how a protein recognizes and interacts with DNA. However, modeling DNA structure presents several challenges. Unlike many proteins, DNA is incredibly flexible and can stretch and twist to fit into a binding groove. Its structure and flexibility are also highly dependent on the surrounding solvent and the specific protein it encounters. Even small changes in DNA sequence can have a large impact on the interaction.
Can current protein structure prediction tools accurately and reliably model protein-DNA interactions?
A recent study from Dr. Barry Stoddard’s lab in the Basic Sciences Division published in Nucleic Acids Research suggests that the answer is no, at least not yet. Researchers in the Stoddard lab experimentally determined the structures of various protein-DNA complexes by X-ray crystallography and compared them with AlphaFold3 predictions of the same complexes. While AlphaFold3 accurately determined overall protein structure, it repeatedly failed to predict the protein-DNA interaction.
The team began with testing the DNA-cutting enzyme I-Onul, whose structure in complex with DNA had already been experimentally determined and was included in the training data for AlphaFold3. In several variants of the enzyme that harbor amino acid substitutions within the DNA binding surface and thereby recognize slightly altered target sites, AlphaFold3 gave an accurate prediction of the overall protein structure but missed the mark at the protein-DNA interface. DNA base pair and backbone geometry differed between the experimentally obtained and predicted structures. The empirical structures showed that the protein-DNA binding involved interactions with solvent molecules, a detail the model failed to predict. Instead, the predicted structures shifted DNA bases and amino acid side chain into inaccurate arrangements.
The researchers moved on to a more realistic AlphaFold3 use case, predicting a protein-DNA interaction whose structure had not been previously described. They explored DNA binding to the protein I-PnoMI. They reasoned it would be a straightforward task because related proteins with solved DNA-bound structures were already known.
This time, AlphaFold3 failed spectacularly. The orientation of protein-DNA binding was flipped in the predicted structure compared to the experimental structure, and the DNA binding site was misaligned. Again, AlphaFold3 did not account for interactions with solvent molecule at the protein-DNA interface.