Skip to content

RL in Scientific Research and Biomedicine

On this page AlphaFold's end-to-end structure prediction and geometric constraints, molecular design and drug discovery with RL (GFlowNets, RL-based molecule generation), experimental design as active learning, and the mathematical-reasoning RL of AlphaProof/AlphaGeometry — the big picture of RL for Science.

RL in Scientific Research and Biomedicine ​

In one sentence: this page covers the four places where RL has genuinely landed in scientific research — the "pseudo-RL" idea behind protein structure prediction, RL-based generation for molecular design and drug discovery, active learning for experimental design, and the AlphaGeometry/AlphaProof breakthroughs that bring "search × learning" to mathematical reasoning — and then takes up evaluation and ethics for "AI doing science."

1. RL for Science at a Glance: Why Scientific Problems Suit RL ​

1.1 Two Fortunate Properties of Scientific Problems ​

In most real-world business settings, RL gets stuck on "no simulator, no reward" (see RL for Financial Trading). Scientific problems, however, tend to have both:

What RL needsWhat science provides
An interactive environment / simulatorPhysics simulators (molecular dynamics, plasmas, fluids), computational experiments, proof environments
A definable objective / rewardBinding energy, stability, accuracy, verifiable theorems
Generatable dataSelf-play, randomly generated problems, sampled molecular conformations

1.2 The Four Branches of RL for Science ​

text
RL for Science
 ├─ Structure prediction: AlphaFold (mostly supervised learning, "pseudo-RL" ideas)
 ├─ Molecular design: RL/GFlowNet generates molecules with desired properties
 ├─ Experimental design: active learning / Bayesian optimization (RL-flavored decisions)
 └─ Mathematical reasoning: AlphaGeometry / AlphaProof (symbolic engine + RL search)

One theme runs through all four: the "environment" of science is a computer simulation (a model), so every branch hugs model-based RL — learn what to do inside the model first, then let experts or experiments verify.

2. Proteins: AlphaFold and the "Pseudo-RL" Lens ​

2.1 What AlphaFold Does ​

AlphaFold2 (Jumper et al., 2021, Nature) predicts a protein's 3D structure from its amino-acid sequence; at CASP14 it reached a median accuracy (GDT) of about 92 — near experimental accuracy. It is not a reinforcement learning system — the core is supervised learning (predicting each residue's position and confidence) plus geometric-constraint optimization (structure refinement).

2.2 Why Bother with a "Pseudo-RL" Lens ​

Even though AlphaFold has no explicit RL component, its design hides ideas that are isomorphic to RL:

AlphaFold componentRL counterpart
End-to-end differentiable (sequence → structure coordinates)A differentiable policy from "input state" to "output action"
Recycled structure refinementIteratively improving the output = sequential decision-making
Geometric-constraint lossReward / constraint shaping
Confidence prediction as a self-supervised signalSomething like value estimation

The more important reference point: AlphaFold's winning recipe — large-scale data + differentiable end-to-end training + integration of domain knowledge (geometry) — is "two variants of the same family" as AlphaGo's methodology: AlphaGo used RL + search, while AlphaFold used SL + geometric optimization. What they share is giving up hand-crafted rules and letting the model learn its own representations. That is why the AI4Science community routinely puts AlphaFold and AlphaGo side by side.

2.3 AlphaFold's Technical Details: Evoformer and Recycling ​

Taken apart, each of AlphaFold2's key designs holds a lesson for RL engineers:

ComponentRoleLesson for RL
Evoformer dual trackAmino-acid and co-evolution information refine each other (row/column attention)"Iteratively refined representations" beat "one pass and done"
RecyclingThe same structure is improved over multiple iterationsIsomorphic to AlphaGo's iterative search
SE(3)-constrained structure moduleOnly produces "legal" 3D geometryConstraining the action space shrinks the search dramatically
Confidence head (pLDDT)Predicts per-residue confidenceThe model carries its own "value estimate," enabling active sampling

One key point: AlphaFold regresses structure coordinates directly from the sequence, which essentially compresses "reasoning" into a forward pass; AlphaGeometry, by contrast, leaves the "reasoning" to search. The two routes illustrate a design decision that recurs throughout AI4Science: when the problem is differentiable, use end-to-end learning; when it demands exact deduction, use search × learning.

INFO

Connecting AlphaFold to RL also has a practical payoff: the "reward signals" of structure prediction (folding energy, agreement with experiment) can serve as reward functions for other scientific tasks such as molecular docking and protein design. AlphaFold's predicted structures can act directly as the simulator for downstream RL tasks — say, designing a protein that binds a specific target. This "predictive model → generative model" pipeline is the route that much of the post-2023 protein-design work has taken.

3. Molecular Design and Drug Discovery ​

3.1 RL for Molecule Generation: MolDQN and REINVENT ​

Task: generate a molecule that satisfies property constraints (e.g., "high predicted potency, low toxicity, synthesizable"). Molecules are represented as graphs or sequences, and the generation process is itself a sequential decision:

text
State s_t  : the molecular fragment generated so far (SMILES string / molecular graph)
Action a_t : add an atom / a bond / a functional group
Reward r_t : property score of the final molecule (docking score, potency, drug-likeness QED, etc.)
Terminal   : molecule complete (valence / length constraints satisfied)
  • MolDQN (Zhou et al., 2019): uses DQN to add atoms and bonds step by step on a molecular graph, directly optimizing properties (single-objective optimization on QED, logP, and the like).
  • REINVENT (Olivecrona et al., 2017): trains an RNN with policy gradients to generate SMILES, with an objective of "property reward + KL penalty to the pretrained prior." This "reward + KL-to-prior" design is isomorphic to the KL penalty in RLHF (see LLM Alignment: RLHF in Practice).

3.2 GFlowNet: From "Find the Best Molecule" to "Sample Diverse Good Ones" ​

GFlowNets (Generative Flow Networks, Bengio et al., 2021) is a framework proposed in the 2020s aimed squarely at the molecular design problem:

text
Plain RL (e.g., MolDQN): learn a policy that finds "the single highest-reward" molecule
GFlowNet             : learn a policy that samples molecules in proportion to reward
                         P(final molecule x) ∝ R(x)   (high-reward molecules are sampled more often)

Why this matters so much: in drug discovery, the "highest-reward molecule" is not necessarily the "best candidate" — property predictions carry error, synthesizability needs human judgment, and a candidate pool needs diversity. GFlowNets guarantee "sample more where the reward is high, but keep diversity," which suits scientific exploration far better than "argmax-style RL."

MethodObjectiveDiversityBest for
Bayesian optimizationFind the optimumLowSmall search spaces
RL molecule generation (MolDQN/REINVENT)Maximize rewardMediumSingle objective, scoreable properties
GFlowNetSample in proportion to rewardHighMulti-objective, diversity needed, noisy rewards

3.3 Practical Realities and Pitfalls ​

  • Molecule property scorers are themselves inaccurate (docking scores correlate only weakly with experimental binding energies) — the "wins in simulation, loses in experiment" Sim2Real problem is known in drug discovery as the sim2wet-lab gap (see the parallel discussion in Robotics Control and Sim2Real);
  • Low synthetic accessibility of RL-generated molecules is the biggest deployment blocker, so the reward must include a synthesizability term;
  • Online fine-tuning with a small amount of real experimental data (Bayesian-optimization-style active learning) works far better than pure offline generation.

3.4 Quick Reference: Evaluation Metrics for Molecular RL ​

A few industry-standard metrics you must know before doing molecular generation RL:

MetricFull name / meaningDirectionUse
QEDDrug-likeness (0–1)Higher is betterGeneral property
logPLipophilicity (partition coefficient)Moderate (~1–5)Drug-likeness
SA_scoreSynthetic accessibility (1–10)Lower is easier to synthesizeKey for deployment
Docking scoreEstimated binding affinity to the targetLower (more negative) is betterActivity proxy
Validity / noveltyFraction of legal molecules; distance from known moleculesHigher is betterGeneration quality

Pitfall: optimizing a single metric ≠ getting a good molecule (molecules with high QED often have synthesizability problems); joint multi-metric weighting or Pareto methods are more realistic — this is the molecular edition of the "multi-objective reward" section in Reward Engineering.

3.5 A Training Loop for Molecular RL: Pseudocode ​

Turning "RL molecule generation" into a runnable loop (policy-gradient family):

python
# Training loop for RL-based molecule generation (pseudocode)
policy = RNN_or_Transformer(smiles_generator)   # sequence model that generates SMILES
prior  = copy(policy)                           # pretrained prior (an ordinary molecule language model)

for step in range(train_steps):
    smiles = policy.sample(batch_size)           # generate a batch of molecules
    score  = scorer(smiles)                      # property score: weighted QED / logP / docking score
    validity = valid(smiles)                     # validity check (valence, syntax)
    r = score * validity                         # invalid molecules get a straight zero
    # key: reward + KL penalty (prevents generating gibberish that fools the scorer)
    loss = -(r - beta * kl(policy, prior)) * logprob(smiles)
    gradient_update(policy, loss)

Three engineering takeaways (matching the pitfalls above):

  1. Prior + KL penalty is the "guardrail" of molecular RL — the same idea as RLHF's reference model, preventing the model from degenerating into "churning out garbled SMILES that game the score" (see LLM Alignment: RLHF in Practice);
  2. The scorer must first pass the toxicity/synthesizability gate: optimizing QED alone yields molecules that "look like drugs but can't be made";
  3. Diversify exploration: if you need a candidate pool rather than a single optimum, go the GFlowNet route (sample in proportion to reward) instead of pure argmax.

4. Experimental Design: Active Learning and Bayesian Optimization ​

4.1 The Problem: Every Round of Experiments Is Expensive ​

Scientists run experiments — screening compounds, synthesizing materials, running simulations — under budget constraints. Experimental design answers "what should the next batch of experiments measure?" so that information or payoff is maximized. This is naturally sequential decision-making:

text
State s_t  : all measured samples so far, (input x_i, outcome y_i)
Action a_t : choose the next batch of x to measure
Reward r_t : payoff after measurement (e.g., model uncertainty reduced / a better y found)
Transition : the real or simulated experiment returns y(x)

4.2 Active Learning + RL-flavored Decision-Making ​

  • The classic approach is Bayesian optimization (BO): a surrogate model (e.g., a Gaussian process) predicts "where is most worth measuring" — exploration + exploitation, essentially the same trade-off as bandits (see the Multi-Armed Bandits page);
  • The RL-flavored idea: learn the acquisition function itself — let a policy learn "where to measure" without hand-designing UCB/EI formulas;
  • Multi-round budget allocation, batch experimental design (measure a batch at a time), and instrument-switching costs are all places where RL beats static BO.

Verdict: for small-scale experimental design, plain BO is enough (simple, sample-efficient). RL only earns its keep when the problem is high-dimensional, when the experiment has dynamics (e.g., a continuous fermentation process), or when long-horizon sequential decisions are needed. Don't use RL for the sake of using RL.

4.3 Hybrid Forms of Active Learning × RL ​

The fusion of experimental design and RL is crystallizing into a few clear routes:

RouteHow it worksMaturity
Policy-based acquisition functionUse RL to learn the policy for "what to measure next" (replacing UCB/EI)Early research
Batch design + constraintsSelect a batch of experiments at once, satisfying instrument/budget constraintsIn practical use
Closed-loop lab (self-driving lab)Robots run experiments automatically + RL decides the next roundFrontier pilots
Molecule generation + BO hybridRL generates candidates, BO ranks them, experiments verify and feed backCommon in industry

A route for teams: first get the closed loop "generate candidates → score in simulation → validate with a few experiments → retrain with feedback" running, then talk about which step to replace with RL. A data closed loop beats algorithmic sophistication — this is the "environment as product" idea from Anatomy of an RL System, replayed in the scientific setting.

5. Mathematical Reasoning: AlphaGeometry / AlphaProof ​

5.1 From Go to Mathematics: Search × Learning, Replayed ​

AlphaGeometry (Trinh et al., 2024, Nature) solves Euclidean geometry problems with a "symbolic engine + neural-network-guided search," cracking 25 of the 30 geometry problems from IMO 2015–2023 (the previous best method solved about 10). Its structure is the AlphaGo methodology, mathematics edition:

text
Symbolic engine (environment): given axioms/theorems, derives automatically (= deterministic "search rules")
Neural network (policy/value): predicts "which derivation step is most likely to lead to a proof"
Search: MCTS-style, network-guided + symbolic-engine verification
Training data: randomly generate hundreds of millions of geometric constructions → solved by the symbolic engine → supervised training

Compare the components with the AlphaGo and Monte Carlo Tree Search page: the symbolic engine ≈ the rules of Go, random data generation ≈ self-play, network guidance ≈ policy/value priors. The "search × learning" formula won again — this time in a completely different field.

5.2 AlphaProof: RL for Formal Theorem Proving ​

In July 2024, DeepMind released AlphaProof: it uses a Lean formal-proof environment + RL to train a model that solves IMO problems. Key points:

  • Translate natural-language problems into Lean formal statements;
  • The model generates proof steps, and the Lean compiler serves as "judge" (reward = does the proof check out) — rewards are fully automatically verifiable, with no human scoring needed;
  • It is the same paradigm as reasoning RL (RLVR; see the Frontier Advances page): large-scale RL driven by verifiable rewards.

AlphaProof reached silver-medal level at IMO 2024 (28 of a possible 42 points, with full marks on the algebra and number-theory problems). This shows that RL can do more than "align how models talk" — it can "create mathematical knowledge."

TIP

The biggest lesson AlphaGeometry/AlphaProof offer RL engineers is environment first: their success rests on having an automatically verifiable, infinitely sample-generable environment (a symbolic engine / the Lean compiler). To replicate this path on a scientific task, the first step is not tuning PPO — it is building the "task → verifiable environment" pipeline. This echoes "environment as product" from Anatomy of an RL System.

5.3 The Engineering Barriers of Formal Proving ​

"Doing mathematics with RL" is not a matter of calling a model API; the engineering barriers concentrate in three places:

BarrierWhat it involvesStatus
FormalizationTranslating natural-language problems into Lean statementsDepends on LLMs + human verification; a bottleneck
Scalable proof environmentSize and retrieval over the theorem library (mathlib)Large but still growing fast
Training computeHundreds of millions of samples + large-scale searchAffordable only for top institutions

Keep perspective: AlphaProof's IMO silver is an "algorithmic victory," but generalizing to real mathematical research is still far away — because the problems it solves are "problem given, answer verifiable," whereas real mathematics is "pose your own problems and organize your own proofs." This gap reminds us: AI4Science milestones should be graded by "degree of task verifiability" — don't sell a single milestone as omnipotence.

6. Physical Control Cases: Nuclear Fusion and Beyond ​

6.1 Controlling Fusion Plasma ​

DeepMind collaborated with the Swiss Plasma Center (Degrave et al., 2022, Nature) to use RL for controlling the shape and position of the plasma inside a tokamak fusion device. Key points:

  • The environment is a physics simulator of the tokamak (RAPTOR and the like), with reward = achieving the target plasma configuration + satisfying constraints (avoiding disruptions);
  • The trained policy controls the plasma from noisy observations and, crucially, can be reconfigured in 3–4 days for a new device or a new target — compared with months of manual tuning under classical control, this rapid adaptation is RL's biggest engineering value;
  • But it still lives inside the simulation → real machine framework, with very few real-machine experiments, and the risk is backstopped by safety guardrails.

6.2 The Common Recipe for "Control-Style" Scientific RL ​

ElementConcrete practice
High-fidelity simulatorPhysics simulation (fusion, fluids, materials)
Clear objective + constraintsShape, temperature, stability constraints
Domain randomization / robustnessCoping with simulation error (sim gap)
Safety guardrailsLimited real-machine trials, fail-safes
Rapid reconfigurationRL policies can be quickly retrained for new parameters

This "simulator + constraints + guardrails + rapid retraining" pipeline maps almost line by line onto Robotics Control and Sim2Real — the RL engineering template for the physical world is universal.

7. Evaluating and Governing Scientific Discovery ​

7.1 Three Layers of Evaluation ​

LayerQuestionStatus
Task layerDoes the model hit the benchmarks (CASP, IMO, molecular property prediction)?Mostly measurable, but benchmarks are easy to overfit
Science layerDoes it produce publishable, reproducible new knowledge?Most work stalls at "good predictions," without closed-loop verification
Impact layerDoes it actually accelerate discovery and change experimental paradigms?A handful of cases (the AlphaFold structure database, ESM, etc.)

The key critique: many "RL for Science" papers win only on in-silico metrics (high docking scores, small prediction error), one step short of real experimental validation. The scientific community demands wet-lab verification as the hard criterion — the Sim2Real problem yet again.

7.2 Ethics and Risks ​

  • Misuse of drug/material generation: AI-designed molecules can double as toxins (dual-use);
  • Evaluation transparency: AI-generated "discoveries" need traceability (training data, validation procedures);
  • Reproducibility: AI4Science papers are often irreproducible because training data or environments are never released — consistent with the reproducibility discipline in Evaluation and Benchmarks.

WARNING

Reward engineering in scientific settings carries a distinct ethical weight: a reward such as "maximize binding energy" may induce the model to generate molecules that cannot be synthesized in practice or are actively harmful. A reward function must define not only "what is good" but also "what must never be done" (synthesizability, toxicity constraints) — this is the most serious form of "rewards as specifications" from Reward Engineering, in the scientific domain.

8. Resource Guide ​

  1. Conceptual foundations: to understand AlphaProof's "verifiable rewards," start with Model-Based RL and Reward Engineering;
  2. Lineage: the "search × learning" evolution AlphaGo → AlphaZero → MuZero → AlphaGeometry — read AlphaGo and Monte Carlo Tree Search;
  3. Tracking the frontier: the latest work on GFlowNets, RLVR, and AI4Science — read Frontier Advances;
  4. Tools and data: a list of the AlphaFold database, molecular datasets, and simulators — see Datasets & Tools.

Quick Reference: Toolchain ​

Open-source tools commonly used in AI4Science RL projects fall into these categories:

CategoryToolsUse
Molecular representationRDKit, Open BabelSMILES parsing/generation, property calculation
Protein structureAlphaFold Database (AFDB), ESMFoldStructure retrieval/prediction
Molecule generationREINVENT, MolDQN, GFlowNet implementationsRL/generative molecular design
CheminformaticsPubChem, ChEMBLCompound and bioactivity data
Theorem provingLean 4 + mathlib, the leanprover communityFormal proofs
Physics simulationMuJoCo, Isaac, fusion simulatorsControl-style scientific tasks

Advice: get one minimal pipeline running first (e.g., "generate molecules with RDKit → score them → optimize QED with a small RL script") before thinking big. Every step of this pipeline has an off-the-shelf open-source implementation, which makes it the fastest path into AI4Science RL — and the best test of whether you actually need RL (often a Bayesian optimization suffices; see the verdict in Section 4).

Further Reading ​

References ​

  • Jumper, J., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873), 583–589.
  • Trinh, T. H., et al. (2024). Solving olympiad geometry without human demonstrations. Nature, 625(7995), 476–482. (AlphaGeometry)
  • DeepMind official blog (July 2024). AI solves IMO problems at silver medal level (AlphaProof and AlphaGeometry 2).
  • Zhou, Z., Kearnes, S., Li, L., Zare, R. N., & Riley, P. (2019). Optimization of Molecules via Deep Reinforcement Learning. Scientific Reports, 9, 10752. (MolDQN)
  • Olivecrona, M., et al. (2017). Molecular de-novo design through deep reinforcement learning. Journal of Cheminformatics, 9, 48. (REINVENT)
  • Bengio, Y., et al. (2021). GFlowNet Foundations. arXiv:2111.09266.
  • Degrave, J., et al. (2022). Magnetic control of tokamak plasmas through deep reinforcement learning. Nature, 602(7897), 414–419.
  • Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.