Appearance
Datasets & Tools Profiles
One-line pitch: this page is the hardware-store shelf of RL engineering. Environment libraries, benchmark suites, offline datasets, and the training toolchain — all laid out in tables with the name, what it's for, what the action space looks like, how to install it, and when to pick it. No more hunting through installation docs.
Get the big picture first
RL "datasets" are not like ML datasets. Most of the time, the dataset is the environment itself — you interact with it and sample from it; only offline RL works with static datasets. So this page covers environments and benchmarks first (about 80% of daily work), then offline datasets, and finally the training toolchain.
1. Environment Libraries
| Library | Purpose / Typical Problems | Action Space | Install | Notes |
|---|---|---|---|---|
| Gymnasium | The standard interface spec plus the classic environments (CartPole, MountainCar, Pendulum, LunarLander…) | Discrete & continuous | pip install gymnasium | The de facto standard — nearly every library is compatible. Learn this one first |
| MuJoCo | Continuous-control physics simulation (robot joints, locomotion) | Continuous | pip install mujoco (then gymnasium[mujoco]) | The workhorse benchmark for deep-RL continuous control; now maintained by DeepMind |
| Brax | GPU-parallel physics simulation (up to millions of environments at once) | Continuous | pip install brax (requires JAX) | Training throughput two orders of magnitude above CPU — use it for research at scale |
| Isaac Gym / Isaac Lab | NVIDIA's GPU simulation, aimed at legged robots, dexterous hands, and embodied AI | Continuous | Requires a separate Isaac Sim download | Isaac Gym has been merged into Isaac Lab — new projects should start with Isaac Lab |
| PettingZoo | The multi-agent environment standard (cooperative and competitive) | Discrete & continuous | pip install pettingzoo | The de facto MARL standard; ships with the SuperSuit wrapper library |
| VizDoom | First-person 3D environments built on the Doom engine | Discrete | pip install vizdoom | A classic for visual navigation and partial-observability research |
| Minigrid | Minimal, highly tunable 2D grid worlds | Discrete | pip install minigrid | First choice for teaching and algorithm ablations — runs fast |
| HighwayEnv | Driving and traffic scenarios | Discrete & continuous | pip install highway-env | A lightweight research environment for autonomous-driving decision-making |
| Jumanji | InstaDeep's combinatorial-optimization environment suite | Discrete & continuous | pip install jumanji | Scheduling, routing, games, and other combinatorial problems; JAX-accelerated |
| SMAC | StarCraft II micromanagement (academic version) | Discrete | Requires the SC2 game itself | A classic MARL benchmark (no longer updated, but still widely used in papers) |
Gymnasium vs. the old Gym
OpenAI Gym was discontinued in 2022, and the community carried it on as Gymnasium (Farama Foundation). Use import gymnasium as gym in all new code. The difference between the gym.Env in legacy tutorials and the new API comes down to what reset() returns and the render modes — our step-by-step tutorial walks through a detailed comparison.
Quick Environment Install
The three most-used install commands (Python 3.9+; create an isolated environment with conda/venv first):
bash
pip install gymnasium # standard env interface + classic toy environments
pip install gymnasium[mujoco] # MuJoCo continuous-control environments on top (pulls in mujoco automatically)
pip install brax # GPU-accelerated simulation (requires JAX; works on CPU but slowly)Sanity check after installing:
python
import gymnasium as gym
env = gym.make("CartPole-v1", render_mode="rgb_array")
obs, info = env.reset(seed=42) # new API: reset returns (obs, info)
obs, reward, terminated, truncated, info = env.step(env.action_space.sample())
print(obs.shape, reward, terminated, truncated)The three most common install pitfalls
- Missing rendering dependencies:
render_modeerrors usually mean you're missingpygameand friends;pip install gymnasium[classic-control]pulls them in. - MuJoCo GL errors: on headless servers, use
render_mode="rgb_array"orMUJOCO_GL=egl. - Pin your versions: record
pip freeze > requirements.txt, or a rerun a few days later may not reproduce your results. More pitfalls in Common Pitfalls & Anti-Patterns.
2. Benchmark Suites
Benchmarks are not single "environments" — they are task sets with an agreed evaluation protocol, built for head-to-head algorithm comparison.
| Benchmark Suite | Contents | Action Space | Strengths / What to Evaluate | Source |
|---|---|---|---|---|
| Atari (ALE) | 50+ Atari games (pixel input) | Discrete (18 actions) | The standard battleground for visual RL and sample efficiency — home turf of DQN/Rainbow | Maintained by Farama: https://github.com/Farama-Foundation/Arcade-Learning-Environment |
| MuJoCo Suite (Gym versions) | HalfCheetah, Hopper, Ant, Walker2d, Humanoid | Continuous | The default setup for continuous control and the policy-gradient/actor-critic family | MuJoCo official site: https://mujoco.org/ |
| Procgen | 16 procedurally generated games | Discrete | Generalization test: the training and test distributions differ for every game | OpenAI: https://github.com/openai/procgen |
| DM Control Suite | 30+ continuous-control tasks (with visual versions) | Continuous | Finer-grained physics tasks plus visual variants — common in DeepMind papers | DeepMind: https://github.com/deepmind/dm_control |
| Meta-World | 50 tabletop robot-arm manipulation tasks | Continuous | Multi-task / meta-learning evaluation (single-task, multi-task, and fast-adaptation tiers) | Maintained by Farama: https://github.com/Farama-Foundation/Meta-World |
| RLBench | 100+ tabletop manipulation tasks (simulated arm) | Continuous | A modern benchmark for robotic manipulation plus large-scale RL training | https://github.com/stepjam/RLBench |
| Maze | Meta's multi-agent task library | Discrete & continuous | MARL and hierarchical decision-making research | Facebook Research: https://github.com/facebookresearch/maze |
Benchmark "established status" — and the criticism
Atari and MuJoCo were the golden benchmarks of 2013–2020, but the criticism today is well-founded: sample efficiency is the main bottleneck, and classic benchmarks are too "cheap" — a simple algorithm can look great on them yet fail to generalize to real problems. That's why 2020s benchmarks (Procgen, Meta-World, RLBench, Isaac Lab) emphasize "generalization" and "pre-deployment testing". For how to design a fair evaluation protocol, see Evaluation & Benchmarks and Building an RL Evaluation from Scratch.
3. Offline RL Datasets
Offline RL requires a fixed historical dataset (interaction logs). The common public ones:
| Dataset | Contents / Source | Data Format | Typical Use | Link |
|---|---|---|---|---|
| D4RL | Four domains: Gym-MuJoCo (locomotion robots), Adroit (dexterous hand), FrankaKitchen (kitchen tasks), AntMaze (ant navigation) | (s, a, r, s′) trajectories, tiered by data quality (medium/expert/replay) | The de facto offline-RL benchmark — nearly every method is evaluated here | https://github.com/Farama-Foundation/D4RL |
| RL Unplugged | Provided by DeepMind: large-scale interaction data for Atari and DM Control, tiered by the collecting agent's strength | Massive trajectory sets | Large-scale offline evaluation and data-scaling research | https://github.com/deepmind/rl_unplugged |
| NeoRL | Industry-friendly offline datasets (with data-quality tiers) | Trajectories plus offline-policy-evaluation companions | Research on offline evaluation methods | https://github.com/polixir/NeoRL |
D4RL's Four Data Domains (all inside one dataset — don't run only one domain)
D4RL is often mistaken for "a single dataset", but it's really a collection of four task domains, and algorithm performance varies enormously between them:
| Domain | Task Shape | State / Action | The Challenge Algorithms Must Overcome |
|---|---|---|---|
| Gym-MuJoCo | Locomotion (Hopper, HalfCheetah, etc.) | Continuous | Easiest to get started; nearly every offline algorithm is validated here |
| AntMaze | Ant robot navigating mazes | Continuous | Sparse reward + long horizons — the touchstone that separates strong methods from weak ones |
| Adroit | Dexterous-hand manipulation (door opening, pen spinning) | High-dimensional continuous | Scarce data (few human demos) — a test of sample utilization |
| FrankaKitchen | Multi-stage kitchen manipulation | Continuous | Multi-stage long-horizon tasks — a test of compositional generalization |
Don't extrapolate conclusions across domains
A method that shines on Gym-MuJoCo (e.g., simple BC regression) can fail completely on AntMaze, and vice versa. When reading an offline RL paper, check which domains it reports — never take one domain's results as a global conclusion. See Offline RL.
What the Data-Quality Tiers Mean
In D4RL, the same environment ships as -medium, -medium-replay, -expert, and so on — the tier directly determines task difficulty:
| Tier | How the Data Was Collected | What It Tests in an Algorithm |
|---|---|---|
-expert | Sampled from an expert policy | Almost the "easiest" — even behavior cloning does well |
-medium | Sampled from a mid-level policy | Requires the algorithm to improve on the data, not just imitate it |
-medium-replay | The replay buffer from a training run | A mixture of qualities — tests exploiting non-optimal data |
-random | Sampled from a random policy | Nearly unlearnable; serves as a lower-bound reference |
The evaluation trap of offline RL
The most insidious trap: offline RL has no "online validation" step during training. You can only evaluate on the training set, and the policy will exploit OOD overestimation to inflate its scores. Reliable checks are: (1) compare against official results on benchmarks like D4RL; (2) confirm with offline evaluation methods (e.g., fitted Q evaluation). Don't trust the training curve alone.
4. The Training Toolchain
The toolchain is organized into five categories: logging → tuning → configuration → experiment management → parallel acceleration.
| Category | Tool | What It Does | Link |
|---|---|---|---|
| Metrics logging | TensorBoard | Learning curves, scalar/histogram visualization — the default for getting started with RL | https://www.tensorflow.org/tensorboard |
| Metrics logging | Weights & Biases (W&B) | Cloud experiment tracking, comparison, and team collaboration | https://wandb.ai/ |
| Hyperparameter tuning | Optuna | Bayesian optimization / random search; lightweight and integrates with any training loop | https://optuna.org/ |
| Hyperparameter tuning | Ray Tune | Distributed hyperparameter search for large-scale parallel RL experiments | https://docs.ray.io/en/latest/tune/ |
| Configuration management | Hydra | Declarative YAML experiment configs — a reproduction lifesaver for RL | https://hydra.cc/ |
| Experiment management | MLflow | Experiment logging + model registry + deployment tracking | https://mlflow.org/ |
| Experiment management | Sacred | Academic experiment logging and config management (older but still common) | https://github.com/IDSIA/sacred |
| Distributed / acceleration | Ray | Distributed task scheduling — the foundation of RLlib | https://www.ray.io/ |
| Environment throughput | EnvPool | High-throughput parallel environment sampling (C++ core) | https://github.com/sail-sg/envpool |
| Environment throughput | Sample Factory | Asynchronous large-scale RL training framework (Atari/robotics) | https://github.com/alex-petrenko/sample-factory |
How the Pieces Fit Together: A Minimal Experiment Skeleton
More tools is not better. The skeleton below covers the full loop — get it running → log → tune → reproduce — and is enough for personal projects:
python
# 1. Fix the seed (before anything else)
seed = 42
env = gym.make("CartPole-v1")
obs, _ = env.reset(seed=seed)
# 2. Training loop (SB3 as an example)
from stable_baselines3 import PPO
model = PPO("MlpPolicy", env, seed=seed,
tensorboard_log="./logs/") # TensorBoard logging
model.learn(total_timesteps=200_000)
# 3. Tuning goes to Optuna (only key hyperparameters shown)
# search_space = {"learning_rate": (1e-5, 3e-3, "log-uniform"), "gamma": (0.9, 0.999)}
# optuna.create_study(direction="maximize").optimize(objective, n_trials=50)
# 4. Configuration & reproduction go to Hydra: put all the parameters above
# into config.yaml, pin versions, and save requirements.txtDivision of labor: TensorBoard for curves → Optuna for hyperparameters → Hydra for configs → MLflow/W&B for experiment archiving. Add Ray when a team needs scale; for a personal project, adding Ray is usually over-engineering.
The minimal engineering stack
For the vast majority of personal projects, a sufficient stack is: Gymnasium + Stable-Baselines3 (or CleanRL) + TensorBoard + Optuna + Hydra. Bring in Ray/W&B only when you need distributed training or production-grade experiment matrices. Don't start with the full stack — RL's traps live in the algorithms and environments, not in the number of tools.
5. How to Choose: A Decision Table
| Your Scenario | Environment / Benchmark | Algorithm Library | Dataset | Toolchain |
|---|---|---|---|---|
| Learning RL for the first time | CartPole / LunarLander in Gymnasium | Hand-rolled + CleanRL | Not needed | TensorBoard |
| Researching continuous control | MuJoCo / DM Control | SB3, Tianshou | D4RL (if going offline) | W&B + Optuna |
| Robotics sim → deployment | MuJoCo → Isaac Lab / RLBench | SAC/PPO (SB3 or in-house) | Self-collected data + domain randomization | Hydra + MLflow |
| Large-scale parallel experiments | Brax / EnvPool | PureJaxRL / Ray RLlib | RL Unplugged | W&B + Ray Tune |
| Multi-agent | PettingZoo / SMAC / Maze | MAPPO (in-house or MARLlib) | Self-collected | W&B + Hydra |
| Offline RL research | D4RL / NeoRL | d3rlpy (third-party offline library) or in-house | D4RL / RL Unplugged | W&B + Optuna |
d3rlpy is a well-known offline RL library (https://github.com/takuseno/d3rlpy) worth trying when running offline RL experiments.
6. How This Page Connects to the Rest of the Site
- How to evaluate environments and benchmarks fairly: see Evaluation & Benchmarks.
- How to turn an evaluation protocol into actual experiments: see Building an RL Evaluation from Scratch.
- Offline RL principles and the OOD overestimation problem: see Offline RL.
- A detailed comparison of six algorithm frameworks: see Choosing Frameworks & Tools.
- Engineering practice for building and wrapping custom environments: see Building an RL Project from Scratch.
Pin your toolchain versions
One of the most common causes of failed RL reproduction is environment/library version drift: the Gymnasium API changed, SB3 versions behave differently, MuJoCo's rendering interface changed. In practice, pin every dependency with requirements.txt (or even conda-lock) and document pip freeze in the README. Details in Common Pitfalls & Anti-Patterns.
Further Reading
- Evaluation & Benchmarks — why RL evaluation is hard and how to read the benchmark landscape.
- Choosing Frameworks & Tools — a deeper comparison of the algorithm libraries mentioned on this page.
- Building an RL Evaluation from Scratch — once you have an environment: experiment matrices, metrics, and how to write the report.
- Offline RL — the methods and traps behind D4RL and RL Unplugged.
- Building an RL Project from Scratch — the full workflow for choosing, building, and wrapping environments.
- Step-by-Step Tutorial: Three Gymnasium Versions Up and Running — run a complete project using the libraries on this page.
References
- Gymnasium official docs: https://gymnasium.farama.org/
- MuJoCo official site: https://mujoco.org/
- Brax (Google): https://github.com/google/brax
- Isaac Lab (NVIDIA): https://isaac-sim.github.io/IsaacLab/
- PettingZoo: https://pettingzoo.farama.org/
- Atari / ALE (Farama): https://github.com/Farama-Foundation/Arcade-Learning-Environment
- Procgen (OpenAI): https://github.com/openai/procgen
- DM Control (DeepMind): https://github.com/deepmind/dm_control
- Meta-World (Farama): https://github.com/Farama-Foundation/Meta-World
- RLBench: https://github.com/stepjam/RLBench
- Maze (Meta): https://github.com/facebookresearch/maze
- D4RL: https://github.com/Farama-Foundation/D4RL
- RL Unplugged (DeepMind): https://github.com/deepmind/rl_unplugged
- Optuna: https://optuna.org/ ; Hydra: https://hydra.cc/ ; Ray: https://www.ray.io/
- TensorBoard: https://www.tensorflow.org/tensorboard ; W&B: https://wandb.ai/ ; MLflow: https://mlflow.org/