Skip to content

Datasets & Tools Profiles

On this page Environment libraries (Gymnasium, MuJoCo, Brax, Isaac Gym/Lab), benchmark suites (Atari, Procgen, DM Control, Meta-World, Maze), datasets (D4RL, RL Unplugged), and the training toolchain (W&B, TensorBoard, Hydra, Optuna) — each profiled in one place.

Datasets & Tools Profiles ​

One-line pitch: this page is the hardware-store shelf of RL engineering. Environment libraries, benchmark suites, offline datasets, and the training toolchain — all laid out in tables with the name, what it's for, what the action space looks like, how to install it, and when to pick it. No more hunting through installation docs.

Get the big picture first

RL "datasets" are not like ML datasets. Most of the time, the dataset is the environment itself — you interact with it and sample from it; only offline RL works with static datasets. So this page covers environments and benchmarks first (about 80% of daily work), then offline datasets, and finally the training toolchain.

1. Environment Libraries ​

LibraryPurpose / Typical ProblemsAction SpaceInstallNotes
GymnasiumThe standard interface spec plus the classic environments (CartPole, MountainCar, Pendulum, LunarLander…)Discrete & continuouspip install gymnasiumThe de facto standard — nearly every library is compatible. Learn this one first
MuJoCoContinuous-control physics simulation (robot joints, locomotion)Continuouspip install mujoco (then gymnasium[mujoco])The workhorse benchmark for deep-RL continuous control; now maintained by DeepMind
BraxGPU-parallel physics simulation (up to millions of environments at once)Continuouspip install brax (requires JAX)Training throughput two orders of magnitude above CPU — use it for research at scale
Isaac Gym / Isaac LabNVIDIA's GPU simulation, aimed at legged robots, dexterous hands, and embodied AIContinuousRequires a separate Isaac Sim downloadIsaac Gym has been merged into Isaac Lab — new projects should start with Isaac Lab
PettingZooThe multi-agent environment standard (cooperative and competitive)Discrete & continuouspip install pettingzooThe de facto MARL standard; ships with the SuperSuit wrapper library
VizDoomFirst-person 3D environments built on the Doom engineDiscretepip install vizdoomA classic for visual navigation and partial-observability research
MinigridMinimal, highly tunable 2D grid worldsDiscretepip install minigridFirst choice for teaching and algorithm ablations — runs fast
HighwayEnvDriving and traffic scenariosDiscrete & continuouspip install highway-envA lightweight research environment for autonomous-driving decision-making
JumanjiInstaDeep's combinatorial-optimization environment suiteDiscrete & continuouspip install jumanjiScheduling, routing, games, and other combinatorial problems; JAX-accelerated
SMACStarCraft II micromanagement (academic version)DiscreteRequires the SC2 game itselfA classic MARL benchmark (no longer updated, but still widely used in papers)

Gymnasium vs. the old Gym

OpenAI Gym was discontinued in 2022, and the community carried it on as Gymnasium (Farama Foundation). Use import gymnasium as gym in all new code. The difference between the gym.Env in legacy tutorials and the new API comes down to what reset() returns and the render modes — our step-by-step tutorial walks through a detailed comparison.

Quick Environment Install ​

The three most-used install commands (Python 3.9+; create an isolated environment with conda/venv first):

bash
pip install gymnasium                 # standard env interface + classic toy environments
pip install gymnasium[mujoco]         # MuJoCo continuous-control environments on top (pulls in mujoco automatically)
pip install brax                      # GPU-accelerated simulation (requires JAX; works on CPU but slowly)

Sanity check after installing:

python
import gymnasium as gym
env = gym.make("CartPole-v1", render_mode="rgb_array")
obs, info = env.reset(seed=42)        # new API: reset returns (obs, info)
obs, reward, terminated, truncated, info = env.step(env.action_space.sample())
print(obs.shape, reward, terminated, truncated)

The three most common install pitfalls

  1. Missing rendering dependencies: render_mode errors usually mean you're missing pygame and friends; pip install gymnasium[classic-control] pulls them in.
  2. MuJoCo GL errors: on headless servers, use render_mode="rgb_array" or MUJOCO_GL=egl.
  3. Pin your versions: record pip freeze > requirements.txt, or a rerun a few days later may not reproduce your results. More pitfalls in Common Pitfalls & Anti-Patterns.

2. Benchmark Suites ​

Benchmarks are not single "environments" — they are task sets with an agreed evaluation protocol, built for head-to-head algorithm comparison.

Benchmark SuiteContentsAction SpaceStrengths / What to EvaluateSource
Atari (ALE)50+ Atari games (pixel input)Discrete (18 actions)The standard battleground for visual RL and sample efficiency — home turf of DQN/RainbowMaintained by Farama: https://github.com/Farama-Foundation/Arcade-Learning-Environment
MuJoCo Suite (Gym versions)HalfCheetah, Hopper, Ant, Walker2d, HumanoidContinuousThe default setup for continuous control and the policy-gradient/actor-critic familyMuJoCo official site: https://mujoco.org/
Procgen16 procedurally generated gamesDiscreteGeneralization test: the training and test distributions differ for every gameOpenAI: https://github.com/openai/procgen
DM Control Suite30+ continuous-control tasks (with visual versions)ContinuousFiner-grained physics tasks plus visual variants — common in DeepMind papersDeepMind: https://github.com/deepmind/dm_control
Meta-World50 tabletop robot-arm manipulation tasksContinuousMulti-task / meta-learning evaluation (single-task, multi-task, and fast-adaptation tiers)Maintained by Farama: https://github.com/Farama-Foundation/Meta-World
RLBench100+ tabletop manipulation tasks (simulated arm)ContinuousA modern benchmark for robotic manipulation plus large-scale RL traininghttps://github.com/stepjam/RLBench
MazeMeta's multi-agent task libraryDiscrete & continuousMARL and hierarchical decision-making researchFacebook Research: https://github.com/facebookresearch/maze

Benchmark "established status" — and the criticism

Atari and MuJoCo were the golden benchmarks of 2013–2020, but the criticism today is well-founded: sample efficiency is the main bottleneck, and classic benchmarks are too "cheap" — a simple algorithm can look great on them yet fail to generalize to real problems. That's why 2020s benchmarks (Procgen, Meta-World, RLBench, Isaac Lab) emphasize "generalization" and "pre-deployment testing". For how to design a fair evaluation protocol, see Evaluation & Benchmarks and Building an RL Evaluation from Scratch.

3. Offline RL Datasets ​

Offline RL requires a fixed historical dataset (interaction logs). The common public ones:

DatasetContents / SourceData FormatTypical UseLink
D4RLFour domains: Gym-MuJoCo (locomotion robots), Adroit (dexterous hand), FrankaKitchen (kitchen tasks), AntMaze (ant navigation)(s, a, r, s′) trajectories, tiered by data quality (medium/expert/replay)The de facto offline-RL benchmark — nearly every method is evaluated herehttps://github.com/Farama-Foundation/D4RL
RL UnpluggedProvided by DeepMind: large-scale interaction data for Atari and DM Control, tiered by the collecting agent's strengthMassive trajectory setsLarge-scale offline evaluation and data-scaling researchhttps://github.com/deepmind/rl_unplugged
NeoRLIndustry-friendly offline datasets (with data-quality tiers)Trajectories plus offline-policy-evaluation companionsResearch on offline evaluation methodshttps://github.com/polixir/NeoRL

D4RL's Four Data Domains (all inside one dataset — don't run only one domain) ​

D4RL is often mistaken for "a single dataset", but it's really a collection of four task domains, and algorithm performance varies enormously between them:

DomainTask ShapeState / ActionThe Challenge Algorithms Must Overcome
Gym-MuJoCoLocomotion (Hopper, HalfCheetah, etc.)ContinuousEasiest to get started; nearly every offline algorithm is validated here
AntMazeAnt robot navigating mazesContinuousSparse reward + long horizons — the touchstone that separates strong methods from weak ones
AdroitDexterous-hand manipulation (door opening, pen spinning)High-dimensional continuousScarce data (few human demos) — a test of sample utilization
FrankaKitchenMulti-stage kitchen manipulationContinuousMulti-stage long-horizon tasks — a test of compositional generalization

Don't extrapolate conclusions across domains

A method that shines on Gym-MuJoCo (e.g., simple BC regression) can fail completely on AntMaze, and vice versa. When reading an offline RL paper, check which domains it reports — never take one domain's results as a global conclusion. See Offline RL.

What the Data-Quality Tiers Mean ​

In D4RL, the same environment ships as -medium, -medium-replay, -expert, and so on — the tier directly determines task difficulty:

TierHow the Data Was CollectedWhat It Tests in an Algorithm
-expertSampled from an expert policyAlmost the "easiest" — even behavior cloning does well
-mediumSampled from a mid-level policyRequires the algorithm to improve on the data, not just imitate it
-medium-replayThe replay buffer from a training runA mixture of qualities — tests exploiting non-optimal data
-randomSampled from a random policyNearly unlearnable; serves as a lower-bound reference

The evaluation trap of offline RL

The most insidious trap: offline RL has no "online validation" step during training. You can only evaluate on the training set, and the policy will exploit OOD overestimation to inflate its scores. Reliable checks are: (1) compare against official results on benchmarks like D4RL; (2) confirm with offline evaluation methods (e.g., fitted Q evaluation). Don't trust the training curve alone.

4. The Training Toolchain ​

The toolchain is organized into five categories: logging → tuning → configuration → experiment management → parallel acceleration.

CategoryToolWhat It DoesLink
Metrics loggingTensorBoardLearning curves, scalar/histogram visualization — the default for getting started with RLhttps://www.tensorflow.org/tensorboard
Metrics loggingWeights & Biases (W&B)Cloud experiment tracking, comparison, and team collaborationhttps://wandb.ai/
Hyperparameter tuningOptunaBayesian optimization / random search; lightweight and integrates with any training loophttps://optuna.org/
Hyperparameter tuningRay TuneDistributed hyperparameter search for large-scale parallel RL experimentshttps://docs.ray.io/en/latest/tune/
Configuration managementHydraDeclarative YAML experiment configs — a reproduction lifesaver for RLhttps://hydra.cc/
Experiment managementMLflowExperiment logging + model registry + deployment trackinghttps://mlflow.org/
Experiment managementSacredAcademic experiment logging and config management (older but still common)https://github.com/IDSIA/sacred
Distributed / accelerationRayDistributed task scheduling — the foundation of RLlibhttps://www.ray.io/
Environment throughputEnvPoolHigh-throughput parallel environment sampling (C++ core)https://github.com/sail-sg/envpool
Environment throughputSample FactoryAsynchronous large-scale RL training framework (Atari/robotics)https://github.com/alex-petrenko/sample-factory

How the Pieces Fit Together: A Minimal Experiment Skeleton ​

More tools is not better. The skeleton below covers the full loop — get it running → log → tune → reproduce — and is enough for personal projects:

python
# 1. Fix the seed (before anything else)
seed = 42
env = gym.make("CartPole-v1")
obs, _ = env.reset(seed=seed)

# 2. Training loop (SB3 as an example)
from stable_baselines3 import PPO
model = PPO("MlpPolicy", env, seed=seed,
            tensorboard_log="./logs/")   # TensorBoard logging
model.learn(total_timesteps=200_000)

# 3. Tuning goes to Optuna (only key hyperparameters shown)
#   search_space = {"learning_rate": (1e-5, 3e-3, "log-uniform"), "gamma": (0.9, 0.999)}
#   optuna.create_study(direction="maximize").optimize(objective, n_trials=50)

# 4. Configuration & reproduction go to Hydra: put all the parameters above
#    into config.yaml, pin versions, and save requirements.txt

Division of labor: TensorBoard for curves → Optuna for hyperparameters → Hydra for configs → MLflow/W&B for experiment archiving. Add Ray when a team needs scale; for a personal project, adding Ray is usually over-engineering.

The minimal engineering stack

For the vast majority of personal projects, a sufficient stack is: Gymnasium + Stable-Baselines3 (or CleanRL) + TensorBoard + Optuna + Hydra. Bring in Ray/W&B only when you need distributed training or production-grade experiment matrices. Don't start with the full stack — RL's traps live in the algorithms and environments, not in the number of tools.

5. How to Choose: A Decision Table ​

Your ScenarioEnvironment / BenchmarkAlgorithm LibraryDatasetToolchain
Learning RL for the first timeCartPole / LunarLander in GymnasiumHand-rolled + CleanRLNot neededTensorBoard
Researching continuous controlMuJoCo / DM ControlSB3, TianshouD4RL (if going offline)W&B + Optuna
Robotics sim → deploymentMuJoCo → Isaac Lab / RLBenchSAC/PPO (SB3 or in-house)Self-collected data + domain randomizationHydra + MLflow
Large-scale parallel experimentsBrax / EnvPoolPureJaxRL / Ray RLlibRL UnpluggedW&B + Ray Tune
Multi-agentPettingZoo / SMAC / MazeMAPPO (in-house or MARLlib)Self-collectedW&B + Hydra
Offline RL researchD4RL / NeoRLd3rlpy (third-party offline library) or in-houseD4RL / RL UnpluggedW&B + Optuna

d3rlpy is a well-known offline RL library (https://github.com/takuseno/d3rlpy) worth trying when running offline RL experiments.

6. How This Page Connects to the Rest of the Site ​

Pin your toolchain versions

One of the most common causes of failed RL reproduction is environment/library version drift: the Gymnasium API changed, SB3 versions behave differently, MuJoCo's rendering interface changed. In practice, pin every dependency with requirements.txt (or even conda-lock) and document pip freeze in the README. Details in Common Pitfalls & Anti-Patterns.

Further Reading ​

References ​