Appearance
RL in Scientific Research and Biomedicine
In one sentence: this page covers the four places where RL has genuinely landed in scientific research — the "pseudo-RL" idea behind protein structure prediction, RL-based generation for molecular design and drug discovery, active learning for experimental design, and the AlphaGeometry/AlphaProof breakthroughs that bring "search × learning" to mathematical reasoning — and then takes up evaluation and ethics for "AI doing science."
1. RL for Science at a Glance: Why Scientific Problems Suit RL
1.1 Two Fortunate Properties of Scientific Problems
In most real-world business settings, RL gets stuck on "no simulator, no reward" (see RL for Financial Trading). Scientific problems, however, tend to have both:
| What RL needs | What science provides |
|---|---|
| An interactive environment / simulator | Physics simulators (molecular dynamics, plasmas, fluids), computational experiments, proof environments |
| A definable objective / reward | Binding energy, stability, accuracy, verifiable theorems |
| Generatable data | Self-play, randomly generated problems, sampled molecular conformations |
1.2 The Four Branches of RL for Science
text
RL for Science
├─ Structure prediction: AlphaFold (mostly supervised learning, "pseudo-RL" ideas)
├─ Molecular design: RL/GFlowNet generates molecules with desired properties
├─ Experimental design: active learning / Bayesian optimization (RL-flavored decisions)
└─ Mathematical reasoning: AlphaGeometry / AlphaProof (symbolic engine + RL search)One theme runs through all four: the "environment" of science is a computer simulation (a model), so every branch hugs model-based RL — learn what to do inside the model first, then let experts or experiments verify.
2. Proteins: AlphaFold and the "Pseudo-RL" Lens
2.1 What AlphaFold Does
AlphaFold2 (Jumper et al., 2021, Nature) predicts a protein's 3D structure from its amino-acid sequence; at CASP14 it reached a median accuracy (GDT) of about 92 — near experimental accuracy. It is not a reinforcement learning system — the core is supervised learning (predicting each residue's position and confidence) plus geometric-constraint optimization (structure refinement).
2.2 Why Bother with a "Pseudo-RL" Lens
Even though AlphaFold has no explicit RL component, its design hides ideas that are isomorphic to RL:
| AlphaFold component | RL counterpart |
|---|---|
| End-to-end differentiable (sequence → structure coordinates) | A differentiable policy from "input state" to "output action" |
| Recycled structure refinement | Iteratively improving the output = sequential decision-making |
| Geometric-constraint loss | Reward / constraint shaping |
| Confidence prediction as a self-supervised signal | Something like value estimation |
The more important reference point: AlphaFold's winning recipe — large-scale data + differentiable end-to-end training + integration of domain knowledge (geometry) — is "two variants of the same family" as AlphaGo's methodology: AlphaGo used RL + search, while AlphaFold used SL + geometric optimization. What they share is giving up hand-crafted rules and letting the model learn its own representations. That is why the AI4Science community routinely puts AlphaFold and AlphaGo side by side.
2.3 AlphaFold's Technical Details: Evoformer and Recycling
Taken apart, each of AlphaFold2's key designs holds a lesson for RL engineers:
| Component | Role | Lesson for RL |
|---|---|---|
| Evoformer dual track | Amino-acid and co-evolution information refine each other (row/column attention) | "Iteratively refined representations" beat "one pass and done" |
| Recycling | The same structure is improved over multiple iterations | Isomorphic to AlphaGo's iterative search |
| SE(3)-constrained structure module | Only produces "legal" 3D geometry | Constraining the action space shrinks the search dramatically |
| Confidence head (pLDDT) | Predicts per-residue confidence | The model carries its own "value estimate," enabling active sampling |
One key point: AlphaFold regresses structure coordinates directly from the sequence, which essentially compresses "reasoning" into a forward pass; AlphaGeometry, by contrast, leaves the "reasoning" to search. The two routes illustrate a design decision that recurs throughout AI4Science: when the problem is differentiable, use end-to-end learning; when it demands exact deduction, use search × learning.
INFO
Connecting AlphaFold to RL also has a practical payoff: the "reward signals" of structure prediction (folding energy, agreement with experiment) can serve as reward functions for other scientific tasks such as molecular docking and protein design. AlphaFold's predicted structures can act directly as the simulator for downstream RL tasks — say, designing a protein that binds a specific target. This "predictive model → generative model" pipeline is the route that much of the post-2023 protein-design work has taken.
3. Molecular Design and Drug Discovery
3.1 RL for Molecule Generation: MolDQN and REINVENT
Task: generate a molecule that satisfies property constraints (e.g., "high predicted potency, low toxicity, synthesizable"). Molecules are represented as graphs or sequences, and the generation process is itself a sequential decision:
text
State s_t : the molecular fragment generated so far (SMILES string / molecular graph)
Action a_t : add an atom / a bond / a functional group
Reward r_t : property score of the final molecule (docking score, potency, drug-likeness QED, etc.)
Terminal : molecule complete (valence / length constraints satisfied)- MolDQN (Zhou et al., 2019): uses DQN to add atoms and bonds step by step on a molecular graph, directly optimizing properties (single-objective optimization on QED, logP, and the like).
- REINVENT (Olivecrona et al., 2017): trains an RNN with policy gradients to generate SMILES, with an objective of "property reward + KL penalty to the pretrained prior." This "reward + KL-to-prior" design is isomorphic to the KL penalty in RLHF (see LLM Alignment: RLHF in Practice).
3.2 GFlowNet: From "Find the Best Molecule" to "Sample Diverse Good Ones"
GFlowNets (Generative Flow Networks, Bengio et al., 2021) is a framework proposed in the 2020s aimed squarely at the molecular design problem:
text
Plain RL (e.g., MolDQN): learn a policy that finds "the single highest-reward" molecule
GFlowNet : learn a policy that samples molecules in proportion to reward
P(final molecule x) ∝ R(x) (high-reward molecules are sampled more often)Why this matters so much: in drug discovery, the "highest-reward molecule" is not necessarily the "best candidate" — property predictions carry error, synthesizability needs human judgment, and a candidate pool needs diversity. GFlowNets guarantee "sample more where the reward is high, but keep diversity," which suits scientific exploration far better than "argmax-style RL."
| Method | Objective | Diversity | Best for |
|---|---|---|---|
| Bayesian optimization | Find the optimum | Low | Small search spaces |
| RL molecule generation (MolDQN/REINVENT) | Maximize reward | Medium | Single objective, scoreable properties |
| GFlowNet | Sample in proportion to reward | High | Multi-objective, diversity needed, noisy rewards |
3.3 Practical Realities and Pitfalls
- Molecule property scorers are themselves inaccurate (docking scores correlate only weakly with experimental binding energies) — the "wins in simulation, loses in experiment" Sim2Real problem is known in drug discovery as the sim2wet-lab gap (see the parallel discussion in Robotics Control and Sim2Real);
- Low synthetic accessibility of RL-generated molecules is the biggest deployment blocker, so the reward must include a synthesizability term;
- Online fine-tuning with a small amount of real experimental data (Bayesian-optimization-style active learning) works far better than pure offline generation.
3.4 Quick Reference: Evaluation Metrics for Molecular RL
A few industry-standard metrics you must know before doing molecular generation RL:
| Metric | Full name / meaning | Direction | Use |
|---|---|---|---|
| QED | Drug-likeness (0–1) | Higher is better | General property |
| logP | Lipophilicity (partition coefficient) | Moderate (~1–5) | Drug-likeness |
| SA_score | Synthetic accessibility (1–10) | Lower is easier to synthesize | Key for deployment |
| Docking score | Estimated binding affinity to the target | Lower (more negative) is better | Activity proxy |
| Validity / novelty | Fraction of legal molecules; distance from known molecules | Higher is better | Generation quality |
Pitfall: optimizing a single metric ≠ getting a good molecule (molecules with high QED often have synthesizability problems); joint multi-metric weighting or Pareto methods are more realistic — this is the molecular edition of the "multi-objective reward" section in Reward Engineering.
3.5 A Training Loop for Molecular RL: Pseudocode
Turning "RL molecule generation" into a runnable loop (policy-gradient family):
python
# Training loop for RL-based molecule generation (pseudocode)
policy = RNN_or_Transformer(smiles_generator) # sequence model that generates SMILES
prior = copy(policy) # pretrained prior (an ordinary molecule language model)
for step in range(train_steps):
smiles = policy.sample(batch_size) # generate a batch of molecules
score = scorer(smiles) # property score: weighted QED / logP / docking score
validity = valid(smiles) # validity check (valence, syntax)
r = score * validity # invalid molecules get a straight zero
# key: reward + KL penalty (prevents generating gibberish that fools the scorer)
loss = -(r - beta * kl(policy, prior)) * logprob(smiles)
gradient_update(policy, loss)Three engineering takeaways (matching the pitfalls above):
- Prior + KL penalty is the "guardrail" of molecular RL — the same idea as RLHF's reference model, preventing the model from degenerating into "churning out garbled SMILES that game the score" (see LLM Alignment: RLHF in Practice);
- The scorer must first pass the toxicity/synthesizability gate: optimizing QED alone yields molecules that "look like drugs but can't be made";
- Diversify exploration: if you need a candidate pool rather than a single optimum, go the GFlowNet route (sample in proportion to reward) instead of pure argmax.
4. Experimental Design: Active Learning and Bayesian Optimization
4.1 The Problem: Every Round of Experiments Is Expensive
Scientists run experiments — screening compounds, synthesizing materials, running simulations — under budget constraints. Experimental design answers "what should the next batch of experiments measure?" so that information or payoff is maximized. This is naturally sequential decision-making:
text
State s_t : all measured samples so far, (input x_i, outcome y_i)
Action a_t : choose the next batch of x to measure
Reward r_t : payoff after measurement (e.g., model uncertainty reduced / a better y found)
Transition : the real or simulated experiment returns y(x)4.2 Active Learning + RL-flavored Decision-Making
- The classic approach is Bayesian optimization (BO): a surrogate model (e.g., a Gaussian process) predicts "where is most worth measuring" — exploration + exploitation, essentially the same trade-off as bandits (see the Multi-Armed Bandits page);
- The RL-flavored idea: learn the acquisition function itself — let a policy learn "where to measure" without hand-designing UCB/EI formulas;
- Multi-round budget allocation, batch experimental design (measure a batch at a time), and instrument-switching costs are all places where RL beats static BO.
Verdict: for small-scale experimental design, plain BO is enough (simple, sample-efficient). RL only earns its keep when the problem is high-dimensional, when the experiment has dynamics (e.g., a continuous fermentation process), or when long-horizon sequential decisions are needed. Don't use RL for the sake of using RL.
4.3 Hybrid Forms of Active Learning × RL
The fusion of experimental design and RL is crystallizing into a few clear routes:
| Route | How it works | Maturity |
|---|---|---|
| Policy-based acquisition function | Use RL to learn the policy for "what to measure next" (replacing UCB/EI) | Early research |
| Batch design + constraints | Select a batch of experiments at once, satisfying instrument/budget constraints | In practical use |
| Closed-loop lab (self-driving lab) | Robots run experiments automatically + RL decides the next round | Frontier pilots |
| Molecule generation + BO hybrid | RL generates candidates, BO ranks them, experiments verify and feed back | Common in industry |
A route for teams: first get the closed loop "generate candidates → score in simulation → validate with a few experiments → retrain with feedback" running, then talk about which step to replace with RL. A data closed loop beats algorithmic sophistication — this is the "environment as product" idea from Anatomy of an RL System, replayed in the scientific setting.
5. Mathematical Reasoning: AlphaGeometry / AlphaProof
5.1 From Go to Mathematics: Search × Learning, Replayed
AlphaGeometry (Trinh et al., 2024, Nature) solves Euclidean geometry problems with a "symbolic engine + neural-network-guided search," cracking 25 of the 30 geometry problems from IMO 2015–2023 (the previous best method solved about 10). Its structure is the AlphaGo methodology, mathematics edition:
text
Symbolic engine (environment): given axioms/theorems, derives automatically (= deterministic "search rules")
Neural network (policy/value): predicts "which derivation step is most likely to lead to a proof"
Search: MCTS-style, network-guided + symbolic-engine verification
Training data: randomly generate hundreds of millions of geometric constructions → solved by the symbolic engine → supervised trainingCompare the components with the AlphaGo and Monte Carlo Tree Search page: the symbolic engine ≈ the rules of Go, random data generation ≈ self-play, network guidance ≈ policy/value priors. The "search × learning" formula won again — this time in a completely different field.
5.2 AlphaProof: RL for Formal Theorem Proving
In July 2024, DeepMind released AlphaProof: it uses a Lean formal-proof environment + RL to train a model that solves IMO problems. Key points:
- Translate natural-language problems into Lean formal statements;
- The model generates proof steps, and the Lean compiler serves as "judge" (reward = does the proof check out) — rewards are fully automatically verifiable, with no human scoring needed;
- It is the same paradigm as reasoning RL (RLVR; see the Frontier Advances page): large-scale RL driven by verifiable rewards.
AlphaProof reached silver-medal level at IMO 2024 (28 of a possible 42 points, with full marks on the algebra and number-theory problems). This shows that RL can do more than "align how models talk" — it can "create mathematical knowledge."
TIP
The biggest lesson AlphaGeometry/AlphaProof offer RL engineers is environment first: their success rests on having an automatically verifiable, infinitely sample-generable environment (a symbolic engine / the Lean compiler). To replicate this path on a scientific task, the first step is not tuning PPO — it is building the "task → verifiable environment" pipeline. This echoes "environment as product" from Anatomy of an RL System.
5.3 The Engineering Barriers of Formal Proving
"Doing mathematics with RL" is not a matter of calling a model API; the engineering barriers concentrate in three places:
| Barrier | What it involves | Status |
|---|---|---|
| Formalization | Translating natural-language problems into Lean statements | Depends on LLMs + human verification; a bottleneck |
| Scalable proof environment | Size and retrieval over the theorem library (mathlib) | Large but still growing fast |
| Training compute | Hundreds of millions of samples + large-scale search | Affordable only for top institutions |
Keep perspective: AlphaProof's IMO silver is an "algorithmic victory," but generalizing to real mathematical research is still far away — because the problems it solves are "problem given, answer verifiable," whereas real mathematics is "pose your own problems and organize your own proofs." This gap reminds us: AI4Science milestones should be graded by "degree of task verifiability" — don't sell a single milestone as omnipotence.
6. Physical Control Cases: Nuclear Fusion and Beyond
6.1 Controlling Fusion Plasma
DeepMind collaborated with the Swiss Plasma Center (Degrave et al., 2022, Nature) to use RL for controlling the shape and position of the plasma inside a tokamak fusion device. Key points:
- The environment is a physics simulator of the tokamak (RAPTOR and the like), with reward = achieving the target plasma configuration + satisfying constraints (avoiding disruptions);
- The trained policy controls the plasma from noisy observations and, crucially, can be reconfigured in 3–4 days for a new device or a new target — compared with months of manual tuning under classical control, this rapid adaptation is RL's biggest engineering value;
- But it still lives inside the simulation → real machine framework, with very few real-machine experiments, and the risk is backstopped by safety guardrails.
6.2 The Common Recipe for "Control-Style" Scientific RL
| Element | Concrete practice |
|---|---|
| High-fidelity simulator | Physics simulation (fusion, fluids, materials) |
| Clear objective + constraints | Shape, temperature, stability constraints |
| Domain randomization / robustness | Coping with simulation error (sim gap) |
| Safety guardrails | Limited real-machine trials, fail-safes |
| Rapid reconfiguration | RL policies can be quickly retrained for new parameters |
This "simulator + constraints + guardrails + rapid retraining" pipeline maps almost line by line onto Robotics Control and Sim2Real — the RL engineering template for the physical world is universal.
7. Evaluating and Governing Scientific Discovery
7.1 Three Layers of Evaluation
| Layer | Question | Status |
|---|---|---|
| Task layer | Does the model hit the benchmarks (CASP, IMO, molecular property prediction)? | Mostly measurable, but benchmarks are easy to overfit |
| Science layer | Does it produce publishable, reproducible new knowledge? | Most work stalls at "good predictions," without closed-loop verification |
| Impact layer | Does it actually accelerate discovery and change experimental paradigms? | A handful of cases (the AlphaFold structure database, ESM, etc.) |
The key critique: many "RL for Science" papers win only on in-silico metrics (high docking scores, small prediction error), one step short of real experimental validation. The scientific community demands wet-lab verification as the hard criterion — the Sim2Real problem yet again.
7.2 Ethics and Risks
- Misuse of drug/material generation: AI-designed molecules can double as toxins (dual-use);
- Evaluation transparency: AI-generated "discoveries" need traceability (training data, validation procedures);
- Reproducibility: AI4Science papers are often irreproducible because training data or environments are never released — consistent with the reproducibility discipline in Evaluation and Benchmarks.
WARNING
Reward engineering in scientific settings carries a distinct ethical weight: a reward such as "maximize binding energy" may induce the model to generate molecules that cannot be synthesized in practice or are actively harmful. A reward function must define not only "what is good" but also "what must never be done" (synthesizability, toxicity constraints) — this is the most serious form of "rewards as specifications" from Reward Engineering, in the scientific domain.
8. Resource Guide
- Conceptual foundations: to understand AlphaProof's "verifiable rewards," start with Model-Based RL and Reward Engineering;
- Lineage: the "search × learning" evolution AlphaGo → AlphaZero → MuZero → AlphaGeometry — read AlphaGo and Monte Carlo Tree Search;
- Tracking the frontier: the latest work on GFlowNets, RLVR, and AI4Science — read Frontier Advances;
- Tools and data: a list of the AlphaFold database, molecular datasets, and simulators — see Datasets & Tools.
Quick Reference: Toolchain
Open-source tools commonly used in AI4Science RL projects fall into these categories:
| Category | Tools | Use |
|---|---|---|
| Molecular representation | RDKit, Open Babel | SMILES parsing/generation, property calculation |
| Protein structure | AlphaFold Database (AFDB), ESMFold | Structure retrieval/prediction |
| Molecule generation | REINVENT, MolDQN, GFlowNet implementations | RL/generative molecular design |
| Cheminformatics | PubChem, ChEMBL | Compound and bioactivity data |
| Theorem proving | Lean 4 + mathlib, the leanprover community | Formal proofs |
| Physics simulation | MuJoCo, Isaac, fusion simulators | Control-style scientific tasks |
Advice: get one minimal pipeline running first (e.g., "generate molecules with RDKit → score them → optimize QED with a small RL script") before thinking big. Every step of this pipeline has an off-the-shelf open-source implementation, which makes it the fastest path into AI4Science RL — and the best test of whether you actually need RL (often a Bayesian optimization suffices; see the verdict in Section 4).
Further Reading
- Model-Based RL and World Models — the unified framework of "scientific simulators as environments" and the pitfalls of model error.
- AlphaGo and Monte Carlo Tree Search — the direct ancestor of AlphaGeometry/AlphaProof's "search × learning."
- Frontier Advances — the latest work on GFlowNets, RLVR, scalable RL, and AI4Science.
- Reward Engineering — reward design for property scores, constraints, and synthesizability.
- Robotics Control and Sim2Real — the isomorphic problem of simulation → real transfer in the physical sciences.
- Datasets & Tools — data and tools for molecules, proteins, simulators, and more.
References
- Jumper, J., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873), 583–589.
- Trinh, T. H., et al. (2024). Solving olympiad geometry without human demonstrations. Nature, 625(7995), 476–482. (AlphaGeometry)
- DeepMind official blog (July 2024). AI solves IMO problems at silver medal level (AlphaProof and AlphaGeometry 2).
- Zhou, Z., Kearnes, S., Li, L., Zare, R. N., & Riley, P. (2019). Optimization of Molecules via Deep Reinforcement Learning. Scientific Reports, 9, 10752. (MolDQN)
- Olivecrona, M., et al. (2017). Molecular de-novo design through deep reinforcement learning. Journal of Cheminformatics, 9, 48. (REINVENT)
- Bengio, Y., et al. (2021). GFlowNet Foundations. arXiv:2111.09266.
- Degrave, J., et al. (2022). Magnetic control of tokamak plasmas through deep reinforcement learning. Nature, 602(7897), 414–419.
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.