AlphaFold 2: predicting protein structures to near-experimental accuracy
DeepMind's neural network predicted protein structures at CASP14 with a median backbone error of under 1 angstrom, comparable to experimental methods, and the team released predicted structures for essentially every human protein within a year.
AlphaFold 2 combines multiple-sequence alignments, an attention-based Evoformer network and an equivariant structure module trained end to end on the Protein Data Bank. At the blind CASP14 assessment in 2020 it achieved a median GDT of about 92 and a median backbone RMSD95 of 0.96 angstrom on the hardest targets, versus 2.8 angstrom for the next-best method.
The model provides per-residue confidence estimates (pLDDT), enabling users to distinguish reliable regions from disordered or uncertain ones. The accompanying AlphaFold Protein Structure Database, built with EMBL-EBI, released predicted structures for the human proteome and later for more than 200 million proteins.
For cancer drug discovery, AlphaFold accelerated structure-based design for targets without crystal structures and, with AlphaFold 3 (2024) extending to protein-ligand and protein-nucleic acid complexes, is now a routine part of the target-to-lead pipeline.
- CASP14: median backbone RMSD95 of 0.96 angstrom (95% CI 0.85-1.16) vs 2.8 angstrom for the next-best method
- Median GDT score around 92 across CASP14 targets, the first time a computational method reached experimental-grade accuracy
- Per-residue confidence (pLDDT) reliably flags disordered and low-confidence regions
- Predicted structures for the entire human proteome released in 2021; over 200 million proteins by 2022
The shape of nearly every protein is now available to any researcher in seconds instead of years, which shortens the path from a cancer target to a designed molecule. It does not by itself produce drugs: binding pockets, dynamics and cellular context still need experiment.
- Predicts single static conformations; many drug targets (kinases, GPCRs, KRAS) move between states
- Accuracy is lower for proteins without evolutionary homologues, disordered regions and multi-protein complexes
- Does not predict effects of point mutations or ligand binding (AlphaFold 3 and other tools partly address this)
- The 2021 model was released with a non-commercial licence for weights, later loosened