Skip to content

AlphaFold and AI for Science

At a glance From its dominant CASP14 win to a database of 200M+ structures and a Nobel Prize in Chemistry, this article breaks down AlphaFold's MSA + Evoformer technical approach and charts the AI for Science paradigm shift, including AlphaFold3 and RFdiffusion.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

AlphaFold and AI for Science ​

AlphaFold is a protein structure prediction system developed by DeepMind: using deep learning, it turned "inferring a protein's three-dimensional structure from its amino acid sequence"—a problem that had stumped biology for half a century—from an experiment-bound scientific challenge into a computational task that can be completed automatically in minutes. In December 2020, AlphaFold2 won CASP14, the most authoritative double-blind competition in protein structure prediction, by a wide margin with a median GDT score of about 92.4; the results were published as a Nature cover paper in July 2021. In May 2024, AlphaFold3 extended prediction from single proteins to biomolecular complexes such as protein–ligand and protein–nucleic acid assemblies. The methods, open-source code, and open database built around AlphaFold, together with the Nobel-level recognition that followed, collectively define the most canonical example of the "AI for Science" research paradigm.

Why does this matter so much? Proteins are the workhorses of life—enzymes catalyze reactions, antibodies recognize pathogens, receptors transmit signals—and every one of these functions depends on the shape a protein folds into in three-dimensional space. Structure determines function, but determining structure has long been expensive, low-throughput experimental work. AlphaFold's significance is not that it "got one protein right"; it lies in driving the marginal cost of obtaining a structure down by several orders of magnitude, and in making "predict first, verify later" a workable workflow in biology. Today structural biologists, drug development teams, and synthetic biologists all use it—it has gone from being a model to being infrastructure.

This article follows a thread from "what it is" to "why it matters": first revisiting the textbook-rewriting moment of CASP14, then explaining why protein folding is hard and how AlphaFold's technical approach breaks the problem down, then placing it in the bigger picture of AI for Science and generative AI, and finally discussing its impact and what it teaches us about the research paradigm. If the underlying concepts are unfamiliar, start with What Are the Hot Concepts in AI to build the big picture, then come back.

1. CASP14: A Textbook-Rewriting Victory ​

CASP (Critical Assessment of protein Structure Prediction) is a biennial double-blind competition founded in 1994: the organizers send participants protein sequences whose structures have not yet been published; participants can only submit predicted structures based on the sequences; in the end, independent teams compare the predictions against experimentally determined structures and score them. The core metric, GDT (Global Distance Test), measures how close a predicted structure is to the real one, on a scale up to 100—the higher the better.

  • CASP13 in 2018: DeepMind's first-generation AlphaFold entered for the first time and ranked first, but its prediction accuracy still lagged noticeably behind experimentally determined structures;
  • CASP14 in December 2020: AlphaFold2 reached a median GDT of about 92.4 across all targets, turning "near-experimental accuracy" from a slogan into a report card; on the widely recognized hardest "free modeling" targets, it likewise far surpassed the historical best, opening up a cliff-edge gap from every other participating team.

What 92.4 Means

A GDT of 92.4 means that on average each residue's predicted position almost coincides with the experimental structure—on many targets, AlphaFold2's predictions are indistinguishable from experimentally determined structures to the naked eye. Structural biologists have offered an analogy: AlphaFold2 compressed "an experimental determination that might take months to years and cost tens of thousands of dollars" into "a few dozen minutes of inference on a GPU".

Nature published AlphaFold2 as a cover paper in July 2021 (Jumper et al., 2021), and the accompanying commentary called protein folding a "50-year-old grand challenge in biology". The problem traces back to the assertion of Christian Anfinsen, winner of the 1972 Nobel Prize in Chemistry—that "the amino acid sequence determines the three-dimensional structure". Over half a century, countless teams tried to compute this mapping; AlphaFold was the first method to push accuracy to near-experimental levels. Looking back on this milestone, AlphaFold was not an isolated algorithmic victory but a landmark event marking deep learning's move "from spectator to main force" in scientific computing—and a footnote to "AI going deep into research" in the brief history of AI's evolution.

2. Problem Definition: Why Protein Folding Is Hard ​

A protein is a long chain built from 20 kinds of amino acids. Inside the cell, this chain spontaneously folds into a specific three-dimensional conformation—this "sequence → structure" mapping is the protein folding problem. It is hard for three reasons:

  • Chemical complexity: a protein is made of thousands of atoms, and its potential energy surface is extremely rugged; in theory there is an astronomically large number of possible conformations (the famous Levinthal's paradox);
  • It is the carrier of function: structure determines function—an enzyme's active site, an antibody's binding surface, a membrane protein's channel are all concrete manifestations of the three-dimensional structure;
  • Sequence diversity: protein sequences in nature are vast in number, and sequenced sequences outnumber determined structures by several orders of magnitude—structure is always the bottleneck.

Before AlphaFold, structures were obtained mainly through experiments and template-based modeling, with enormous cost differences:

MethodPrincipleCostLimitations
X-ray crystallographyCrystallize the protein, reconstruct electron density from X-ray diffractionMonths to years; requires high-quality crystalsMany proteins cannot be crystallized; extremely low throughput
Cryo-electron microscopy (cryo-EM)Freeze the sample, image with an electron microscope, reconstruct in 3DExpensive equipment, demanding purificationSensitive to conformational heterogeneity
Nuclear magnetic resonance (NMR)Measure magnetic signals of atomic nuclei in solutionModerateOnly suitable for small proteins
Homology modelingBorrow the known structure of a homologous protein as a templateMinutesNearly useless without a homologous template
AlphaFoldDeep learning directly learns the "sequence → structure" mappingMinutes on a GPUDynamic conformations and complex assemblies still require experiments

Traditional computational routes come in two flavors: template-based modeling (homology modeling) and ab initio prediction. The former performs well when "a homologous structure is available" but is helpless against entirely novel folds; the latter is mathematically extremely difficult and for a long time only worked for small proteins. AlphaFold's innovation lies in this: instead of brute-forcing the physical equations, it "learns" the rules of folding from vast amounts of known structures and evolutionary information—essentially the same data-driven approach as large language models, except that its "language" is protein sequence and structure.

3. Technical Breakdown: From Sequence to Structure ​

AlphaFold2 takes an amino acid sequence as input and outputs three-dimensional coordinates for every atom. In between are three major steps: build a multiple sequence alignment (MSA) → encode with the two-track Evoformer → decode coordinates with the structure module.

1. Multiple Sequence Alignment (MSA): bringing in evolutionary information ​

Proteins tolerate a great many mutations over evolution, so homologous sequences are similar but not identical across species. Aligning the target sequence against homologous sequences in databases yields a multiple sequence alignment (MSA)—a "sequence × residue position" matrix. The co-evolutionary covariation information encoded in an MSA is enormously valuable: when two residues keep mutating in concert over evolution (when one changes, the other must follow), it often means they are in direct contact in three-dimensional space. AlphaFold takes the MSA as a core input feature—effectively "compressing" hundreds of millions of years of evolutionary experiments into the network.

2. Evoformer: an attention network, protein edition ​

Evoformer is the encoding core of AlphaFold2—think of it as a domain-specific adaptation of the Transformer and attention mechanisms. It maintains two information streams simultaneously:

  • The MSA representation: reasoning about "which pairs of residues have a co-evolutionary relationship";
  • The pair representation: a dense matrix tracking "the relative positional relationship between any two residues".

Evoformer's key design is the triangle update: information for each residue pair can only be updated along triangle-constrained paths such as "the i–j edge plus the j–k edge influencing the i–k edge". This geometry-aware inductive bias lets the network learn correct structural constraints even when training data is relatively limited, and it is AlphaFold's most important injection of domain knowledge relative to a generic transformer. In structural bioinformatics, Evoformer is often called "the BERT of protein language", but its task is not generation—it encodes MSA information into representations used for coordinate prediction.

3. Structure module: turning representations into coordinates ​

Once encoding is done, the structure module decodes the pair representation into three-dimensional coordinates of the backbone atoms. It uses SE(3)-equivariant attention layers, guaranteeing that the network's output behaves consistently under global rotations and translations—a geometric constraint that the laws of physics inherently demand. Training uses the FAPE (Frame Aligned Point Error) loss, which measures each atom's error in local coordinate frames—equivalent to decomposing "global positional error" into "local coordinate errors", so the network learns correct local geometry first and global assembly second. Finally, the network feeds its output coordinates back through recycling for several more rounds of iteration, progressively correcting itself.

Why geometric constraints matter

What AlphaFold outputs must be "legal" three-dimensional coordinates: reasonable bond lengths and bond angles, no atoms interpenetrating. Baking SE(3) equivariance and the FAPE loss into the model architecture means building these physical constraints into the network itself rather than hoping the network stumbles into them by luck—this is the fundamental difference between AlphaFold and many deep learning approaches that treat the problem as pure regression.

4. Prediction confidence: the model reports how sure it is ​

AlphaFold2 outputs two kinds of confidence scores along with the structure: pLDDT (a per-residue local confidence on a 0–100 scale; below 50 usually indicates a disordered region) and PAE (predicted aligned error—an error measure between pairs of residues, used to judge whether the relative placement of two structural domains is reliable). These two outputs mean users don't have to blindly trust the structure: low-confidence regions are precisely the "homework left for experimental verification". Reasoning with confidence is essentially "uncertainty quantification" from model evaluation and validation applied to scientific computing—a good model gives not just an answer, but a confidence interval for the answer.

python
# AlphaFold2 inference logic (conceptual sketch, not the real API)
msa = build_msa(sequence, sequence_db)           # ① homolog search + multiple sequence alignment
repr_msa, repr_pair = evoformer(msa, targets)    # ② two-track attention encoding
coords, conf = structure_module(repr_pair)       # ③ equivariant decoding → coordinates + confidence
# conf.plddt = per-residue confidence; conf.pae = pairwise error matrix

4. Is AlphaFold "Generative"? What Does It Have to Do with LLMs? ​

First, a conceptual clarification: AlphaFold does "structure prediction", not "generation". It takes a known sequence as input and outputs the three-dimensional structure corresponding to that sequence—a discriminative "mapping", not a generative model "sampling new samples from a distribution". AlphaFold cannot design brand-new proteins out of thin air, nor does it write protein sequences. This distinction maps onto the "discriminative vs generative" classification in concept boundaries.

But the two are closely related by technical descent:

  • Both benefit from deep learning, attention mechanisms, and large-scale data;
  • Evoformer is a domain-specific variant of the transformer family;
  • Both gain their generalization under the "end-to-end learned representations + large-scale pretraining" paradigm.

The differences are equally stark: an LLM's "language" is discrete tokens and its task is modeling a sequence distribution; AlphaFold's "language" is protein sequence and structure and its task is regressing three-dimensional coordinates. By the taxonomy of large language models, AlphaFold is more like a "specialized scientific model" than a general-purpose language model—which is exactly why it never became "a protein version of GPT" but rather a structure-solving machine.

The landscape of AI for Science is far bigger than AlphaFold, and a substantial part of it is genuinely generative:

DirectionRepresentative workGenerative?Link
Protein structure predictionAlphaFold2 / AlphaFold3No (prediction)This article
Protein designRFdiffusion (Baker lab, 2023)Yes (diffusion model)Diffusion models
Materials discoveryGNoME (DeepMind, 2023)No (screening + prediction)—
Weather forecastingGraphCast (DeepMind, 2023)No (GNN prediction)—
Biomolecular complexesAlphaFold3 (2024)No (prediction)This article
  • RFdiffusion: carried diffusion models from images over to proteins, "diffusing" entirely new protein backbones out of noise, then pairing with a sequence designer to generate the corresponding amino acid sequences—the same "denoising sampling" mathematics as in image generation, and a flagship of "generative AI for Science".
  • GraphCast: uses graph neural networks for medium-range weather forecasting, outperforming the European Centre for Medium-Range Weather Forecasts (ECMWF) high-resolution operational system on 10-day forecast accuracy—a landmark piece of AI scientific computing in meteorology.
  • GNoME: uses graph networks to screen crystalline materials, predicting over a million potentially stable structures and turning materials discovery from "finding a needle in a haystack" into "targeted search".

5. Impact: From a Database to a Nobel Prize ​

The AlphaFold Protein Structure Database (AlphaFold DB) is hosted by the European Bioinformatics Institute (EBI); as of this article's data cutoff (June 2025), it contains more than 200 million protein structures covering over 1 million species. For the vast number of proteins without experimentally determined structures—especially from microbiomes, deep-sea organisms, and pathogens—this is the first time humanity has had a searchable store of structural assets. This kind of open data infrastructure can be explored further in the datasets & tools archive and the models & leaderboards quick reference.

DeepMind open-sourced the complete AlphaFold2 code in 2021; the subsequent AlphaFold3 offers an online service and API, but its training code has never been fully open-sourced—sparking an ongoing debate in the AI for Science community about "open results vs open methods", structurally identical to the arguments over open-sourcing frontier models in AI safety and governance.

The 2024 Nobel Prize in Chemistry went to work in "computational protein science". In the Nobel committee's official terms: one half went to David Baker (computational protein design), and the other half was shared jointly by Demis Hassabis and John M. Jumper (for AlphaFold's protein structure prediction). This is the highest-level recognition deep learning has received since entering the sciences.

On the industry side, DeepMind founded Isomorphic Labs in 2021 to apply AlphaFold to drug development; AlphaFold3's extended ability to predict protein–ligand and protein–nucleic acid interactions is also moving "AI-assisted drug discovery" from concept to pipeline. But stay clear-eyed: a predicted structure ≠ experimental validation. AlphaFold still has clear limits on dynamic conformations, disordered regions, and the accuracy of certain complex interactions—a caveat that cannot be omitted in any serious scientific application.

The limits of AlphaFold3

AlphaFold3 greatly expanded predictive capability for "biomolecular complexes", but its quantitative predictions of binding affinity for many protein–ligand pairs still diverge from experiment; the field broadly agrees that it "raised the ceiling of structure prediction rather than replacing experiments". Before citing its conclusions, always read its confidence outputs and the applicability of its methods.

6. Takeaways: How AI Changes the Research Paradigm ​

The predict–experiment–validate loop. Traditional research is "trial and error by experiment": run the experiment, look at the result, adjust the hypothesis. AlphaFold demonstrated a new loop—AI first produces high-confidence predictions, then experiments focus on verification and correction. Structural biology can now spend its effort on "the places the AI is unsure about" (low-pLDDT regions, dynamic conformations) instead of determining every protein from scratch. This predict–experiment–validate cycle is spreading to materials, meteorology, drug discovery, and genomics.

Accelerated scientific discovery. When obtaining a structure goes from "months" to "minutes", the timelines of downstream tasks—docking, enzyme engineering, vaccine design—are compressed wholesale. Speed is not the goal in itself, but speed multiplies "how many tries a researcher can make" by one to two orders of magnitude—scientific discovery is at heart a search, and AI enlarges the search space while driving down the cost of each search.

Three ingredients, none optional. Success in AI for Science depends on combining high-quality data × domain knowledge × powerful compute. In AlphaFold's recipe, the MSA supplied the domain knowledge, the PDB (Protein Data Bank) supplied decades of accumulated structural data, and GPUs supplied the compute. For teams looking to enter this space, these three points are worth thinking through before any model architecture.

Boundaries and critique. AlphaFold proved the power of data-driven scientific computing, but it also reminds us: a model's high confidence is not the same as physical correctness, and new workflows and new evaluation protocols are needed to bridge prediction and experiment. This echoes the "metrics misaligned with objectives" lesson from evaluation and benchmarks.

From a competition champion, to a public database of 200M+ structures, to the Nobel Prize in Chemistry, AlphaFold covered in a few years a path the traditional research paradigm might have needed decades to travel. It leaves "AI for Science" with a replicable lesson: when the domain problem can be cleanly defined, the data can be collected at scale, and the mapping can be learned directly by deep learning, AI can become the primary engine of scientific discovery rather than an assistive tool.

Further Reading ​

References ​