Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Results

These results are seeded reference runs for r2l’s PPO implementation. They exercise both supported tensor backends across classic control, Box2D, MuJoCo, Bullet, and partially observable environments, providing broad integration coverage and a baseline for identifying behavioral regressions as the implementation evolves. Each configuration uses seed 0 to reduce variation between benchmark passes.

Stable Baselines3 (SB3) curves are included as a familiar external reference. They help make unexpected learning behavior easier to spot and are interesting to compare, but the runs are not designed as a controlled framework ranking.

Setup

All three implementations use the per-environment PPO configurations in the benchmark configuration, which are based on RL Baselines3 Zoo settings. The r2l runs use the same configuration for the Candle and Burn backends. Evaluation follows each implementation’s existing evaluation path.

The runs use r2l 0.0.3 at commit 6030dd8 and execute as Google Cloud Batch spot jobs in europe-west1 on CPU-only c4d-highcpu-2 workers. Each task is allocated one vCPU and 1,250 MiB of memory, with Rayon and OpenBLAS limited to one thread. Up to 36 tasks run concurrently.

Summary

The current snapshot contains evaluation data for 28 environments on each r2l backend and Stable Baselines3.

Environment familyEnvironments
Classic control5
Box2D4
MuJoCo10
Bullet8
PopGym1
Total28

Both r2l backends produce learning curves across the full suite through the same public API and PPO configuration. Many Candle and Burn curves follow similar broad trends, while some environments show substantial differences between the backends and SB3.

All evaluation curves

The horizontal axis is the number of environment interactions sampled during training. Faint lines show the recorded evaluation values and solid lines show moving averages. Higher reward is better, but reward scales should only be compared within an environment. The plots are generated by book/scripts/plot_results.py.

Classic control

Acrobot-v1

Acrobot-v1 evaluation curves

CartPole-v1

CartPole-v1 evaluation curves

MountainCar-v0

MountainCar-v0 evaluation curves

MountainCarContinuous-v0

MountainCarContinuous-v0 evaluation curves

Pendulum-v1

Pendulum-v1 evaluation curves

Box2D

BipedalWalker-v3

BipedalWalker-v3 evaluation curves

BipedalWalkerHardcore-v3

BipedalWalkerHardcore-v3 evaluation curves

LunarLander-v3

LunarLander-v3 evaluation curves

LunarLanderContinuous-v3

LunarLanderContinuous-v3 evaluation curves

MuJoCo

Ant-v4

Ant-v4 evaluation curves

HalfCheetah-v4

HalfCheetah-v4 evaluation curves

Hopper-v4

Hopper-v4 evaluation curves

Humanoid-v4

Humanoid-v4 evaluation curves

HumanoidStandup-v2

HumanoidStandup-v2 evaluation curves

InvertedDoublePendulum-v2

InvertedDoublePendulum-v2 evaluation curves

InvertedPendulum-v2

InvertedPendulum-v2 evaluation curves

Reacher-v2

Reacher-v2 evaluation curves

Swimmer-v4

Swimmer-v4 evaluation curves

Walker2d-v4

Walker2d-v4 evaluation curves

Bullet

AntBulletEnv-v0

AntBulletEnv-v0 evaluation curves

HalfCheetahBulletEnv-v0

HalfCheetahBulletEnv-v0 evaluation curves

HopperBulletEnv-v0

HopperBulletEnv-v0 evaluation curves

HumanoidBulletEnv-v0

HumanoidBulletEnv-v0 evaluation curves

InvertedDoublePendulumBulletEnv-v0

InvertedDoublePendulumBulletEnv-v0 evaluation curves

InvertedPendulumSwingupBulletEnv-v0

InvertedPendulumSwingupBulletEnv-v0 evaluation curves

ReacherBulletEnv-v0

ReacherBulletEnv-v0 evaluation curves

Walker2DBulletEnv-v0

Walker2DBulletEnv-v0 evaluation curves

PopGym

popgym-BattleshipEasy-v0

popgym-BattleshipEasy-v0 evaluation curves