Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Overview

Warning

Pre-alpha: This library is under active development. APIs may change between releases, and some planned features are not implemented yet. Basic familiarity with reinforcement learning concepts and rust knowledge is assumed.

Why r2l

The goal of r2l is to be a customizable, ergonomic and easily embeddable library. To be more exact:

  • Customizable: users have control over how agents are trained. r2l defines how components interact and exposes lifecycle hooks for custom behavior.
  • Ergonomic: most users are not necessarily concerned with implementation details. High-level builders provide common configurations.
  • Embeddable: r2l describes its backend requirements with traits. Candle and Burn implementations are currently available.

The near-term scope of r2l is a dependable, well-tested on-policy stack for PPO and A2C. Stable Baselines3 is used as a benchmark reference, not as a feature-parity target. The next planned extension is recurrent-policy support, starting with an end-to-end recurrent PPO path across both tensor backends. Broader capabilities will be added as independently tested vertical slices rather than by trying to reproduce another library’s entire algorithm catalog.

Potential longer-term work includes hyperparameter tuning, monitoring integrations, additional persistence formats, and multi-agent support. These directions do not currently have release commitments.

About this book

This book will help you get up to speed with using and hacking r2l. In particular:

  • User Guide: Introduces how environments are to be implemented and how to work with the higher level APIs. Most users should start here. Some basic examples are also shown.
  • On policy algorithms: A detailed architectural overview of the components of on-policy algorithms, how the pieces fit together, and how to create your own custom hook system.
  • Off-policy algorithms: Current support status.

Getting started

For most applications, r2l is the main dependency. It provides complete PPO and A2C builders while the lower-level workspace crates define environments, samplers, agents, and backend integrations. This getting started guide primarily uses the r2l crate, which itself builds on lower-level crates. Implementing a custom environment also requires the error type from r2l-core. If the current setup does not satisfy you, the lower level hooks allow for a lot of hackability.

Shortest setup

The shortest Gymnasium-based PPO setup is:

r2l = { version = "0.0.3", features = ["gym"] }
extern crate r2l;
use r2l::PPOBuilder;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let builder = PPOBuilder::gym("Pendulum-v1", 4)?;
    let mut algorithm = builder.build()?;
    algorithm.train()?;
    Ok(())
}

The gym feature requires Python 3.11 or newer and the gymnasium Python package. A Gymnasium environment id is passed to GymEnvBuilder, which maps supported Gymnasium spaces to r2l space descriptions. Discrete spaces currently require start = 0; non-zero start values are not supported. The builders support other environment types through a different construction, which will be introduced later on.

Saving training artifacts

We rarely want to train algorithms for the sake of it. Once an algorithm is learned, we are usually curious about

  • how the chosen hyperparameters affect learning performance
  • how to run inference using the model
  • less often, maybe we are curious about the running performance of the algorithm

r2l allows saving these artifacts using the with_training_artifacts builder method.

extern crate r2l;
use r2l::{PPOBuilder, TrainingArtifactsConfig};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let builder = PPOBuilder::gym("Pendulum-v1", 4)?;
    let artifacts_config = TrainingArtifactsConfig::new("runs/pendulum")
        .with_evaluation_results(true)
        .with_training_timings(true)
        .with_inference_artifacts(true);
    let mut algorithm = builder
        .with_training_artifacts(artifacts_config)
        .build()?;
    algorithm.train()?;
    Ok(())
}

Under the hood, r2l evaluates the policy with dedicated environments that are reset and reused between evaluation passes, and remembers the best-performing policy. Evaluation runs after every rollout by default. Its frequency and other settings can be customized with EvaluationSettings. Once training is done, the following new files are created in the artifacts folder:

$ tree runs/pendulum
runs/pendulum
├── actor.safetensors
├── evaluations.csv
├── inference.yaml
└── training_timings.csv

1 directory, 4 files

Running inference

Using the inference artifacts and a new environment, you can create an InferenceRunner.

extern crate r2l;
use r2l::{GymEnv, InferenceRunner};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let env = GymEnv::new("Pendulum-v1", Some("human".to_owned()))?;
    let mut inference = InferenceRunner::load_from_env("runs/pendulum", env)?;
    for _ in 0..4 {
        inference.run_episode()?;
    }
    Ok(())
}

The inference configuration records the backend, policy architecture settings, and observation normalization mode needed to rebuild the policy. It does not contain training hyperparameters such as the learning rate or discount factor.

Applications that receive observations outside an Env can load an InferencePolicy directly and request actions from raw observations. For a stateful control loop, implement InferenceEnv and initialize an InferenceRunner with the first raw observation. Full Env implementations also support resetting and running complete episodes.

Environments

Environments implement the Env trait.

/// Environment interface used by samplers.
pub trait Env {
    /// Tensor type used for observations and actions.
    type Tensor: R2lTensor;

    /// Resets the environment and returns the initial observation.
    ///
    /// # Errors
    ///
    /// Returns an error if the environment cannot be reset.
    fn reset(&mut self, seed: u64) -> Result<Self::Tensor, Error>;

    /// Applies one action and returns the resulting transition snapshot.
    ///
    /// # Errors
    ///
    /// Returns an error if the environment cannot apply the action.
    fn step(&mut self, action: Self::Tensor) -> Result<Snapshot<Self::Tensor>, Error>;

    /// Returns static observation/action space metadata.
    fn env_description(&self) -> EnvDescription<Self::Tensor>;
}

Algorithm builders receive an EnvBuilder rather than a concrete environment so that each sampler worker can construct its environment in the place where it runs.

/// Factory for constructing environments of one compatible type.
pub trait EnvBuilder: Sync + Send + 'static {
    /// Environment type produced by this builder.
    type Env: Env;

    /// Builds a fresh environment instance.
    ///
    /// # Errors
    ///
    /// Returns an error if the environment cannot be constructed.
    fn build_env(&self) -> Result<Self::Env, Error>;

    /// Returns the environment description for produced environments.
    ///
    /// # Errors
    ///
    /// Returns an error if a representative environment cannot be constructed.
    fn env_description(&self) -> Result<EnvDescription<<Self::Env as Env>::Tensor>, Error> {
        let env = self.build_env()?;
        Ok(env.env_description())
    }
}

A closure or function returning r2l_core::error::Result<E> automatically implements EnvBuilder.

let env_builder = || -> r2l_core::error::Result<MyEnv> { Ok(MyEnv) };
let ppo_builder = PPOBuilder::new(env_builder, 10)?;
let ppo = ppo_builder.build()?;

For a more detailed example of how to implement the Env and EnvBuilder traits, see the environment building example.

Hyperparameters

Both the PPO and A2C builders expose a great deal of hyperparameters that can be tuned.

Backends

The algorithm builders default to Candle. Use with_candle(device) to choose a Candle device explicitly, or with_burn() to use the default Burn autodifferentiation backend.

Backend selection changes the concrete builder type, so call it before methods that are specific to a chosen backend when following compiler suggestions.

Rollout collection

with_rollout_steps(n) collects n steps per environment for each rollout. with_rollout_episodes(n) switches to episode-bounded sampling and collects n completed episodes per environment.

The builders default to SamplerExecutionMode::MultiThreaded, which runs workers on dedicated threads. Use SamplerExecutionMode::SingleThreaded to step workers sequentially on the calling thread:

extern crate r2l;
use r2l::{PPOBuilder, SamplerExecutionMode};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let builder = PPOBuilder::gym("Pendulum-v1", 4)?
        .with_execution_mode(SamplerExecutionMode::SingleThreaded);
    Ok(())
}

Gymnasium calls still execute under Python’s interpreter lock, so threaded execution should not be assumed to improve Python-environment throughput.

with_observation_normalizer(Some(clip)) switches to staged sampling and enables observation normalization. Passing None switches to staged sampling without a normalizer. Step-bounded rollouts can also normalize discounted rewards with with_reward_normalizer(gamma, clip_reward).

Training schedules

TrainingLimit::rollouts(n) stops after n rollout collections. TrainingLimit::steps(n) stops after at least n sampled environment steps across all workers.

LearningRateSchedule::Constant(rate) keeps the configured rate fixed. LearningRateSchedule::Linear(rate) decays it from rate to zero over the configured training schedule.

For the underlying traits and hook points, continue with On-policy algorithms. For exact builder methods, use the r2l reference.

Examples

The workspace includes runnable examples in crates/r2l-examples. They are intended to show complete setups that can be inspected, copied, and changed without reading every lower-level crate first.

Run examples from the workspace root:

cargo run -p r2l-examples --example <name>

The examples in this section focus on two common starting points:

  • Environment building: implementing Env, using EnvBuilder, and plugging Gymnasium environments into builders.
  • Algorithms: configuring PPO and A2C runs with backends, rollout bounds, training schedules, reporting, and artifacts.

Gymnasium-backed examples require Python 3.11 or newer with gymnasium installed in the Python environment used by the process.

Environment building

Algorithms are built from an EnvBuilder, not from a single environment instance. This lets samplers create one environment per worker and keeps environment state local to the worker that steps it.

The env_building example shows the supported construction styles:

  • implementing Env for a custom environment;
  • implementing EnvBuilder for a custom builder type;
  • passing a closure or function that returns r2l_core::error::Result<E>;
  • using GymEnvBuilder directly;
  • using PPOBuilder::gym for Gymnasium environment ids.

Run it from the workspace root:

cargo run -p r2l-examples --example env_building

The full example is:

use r2l::{Env, EnvBuilder, EnvDescription, PPOBuilder, Snapshot, Space, VecTensor};
use r2l_core::error::Error;
use r2l_gym::GymEnvBuilder;

// Not a working implementation an actual env
pub struct MyEnv;

impl Env for MyEnv {
    type Tensor = VecTensor;

    fn reset(&mut self, _seed: u64) -> Result<Self::Tensor, Error> {
        Ok(VecTensor::new(vec![0., 0.], vec![2])?)
    }

    fn step(&mut self, _action: Self::Tensor) -> Result<Snapshot<Self::Tensor>, Error> {
        let state = VecTensor::new(vec![0., 0.], vec![2])?;
        let reward = 0.;
        let terminated = false;
        let truncated = false;
        let snapshot = Snapshot::new(state, reward, terminated, truncated);
        Ok(snapshot)
    }

    fn env_description(&self) -> EnvDescription<Self::Tensor> {
        let observation_space = Space::Box {
            min: None,
            max: None,
            shape: vec![2],
        };
        let action_space = Space::Discrete(2);
        EnvDescription::new(observation_space, action_space)
    }
}

struct MyEnvBuilder;

impl EnvBuilder for MyEnvBuilder {
    type Env = MyEnv;

    fn build_env(&self) -> Result<Self::Env, Error> {
        Ok(MyEnv)
    }
}

#[allow(clippy::unnecessary_wraps)]
fn build_env() -> Result<MyEnv, Error> {
    Ok(MyEnv)
}

fn main() -> anyhow::Result<()> {
    // Anything that implements Into<GymEnvBuilder> can be used with PPOBuilder::gym.
    // method. This includes &str, String and GymEnvBuilder itself (or your own implementation)
    let ppo_builder0 = PPOBuilder::gym("Pendulum-v1", 10)?;
    let _ppo0 = ppo_builder0.build()?;

    // Since GymEnvBuilder is an EnvBuilder, it can be used with PPOBuilder::new.
    let gym_env_builder = GymEnvBuilder::new("Pendulum-v1");
    let ppo_builder1 = PPOBuilder::new(gym_env_builder, 10)?;
    let _ppo1 = ppo_builder1.build()?;

    // This closure that returns an environment can be used as an environment builder
    let env_builder = || Ok(MyEnv);
    let ppo_builder = PPOBuilder::new(env_builder, 10)?;
    let _ppo = ppo_builder.build()?;

    // This function that returns an environment can also be used as an environment builder
    let ppo_builder3 = PPOBuilder::new(build_env, 10)?;
    let _ppo3 = ppo_builder3.build()?;

    // We can implement our own environment builder to be used with PPOBuilder::new.
    let ppo_builder4 = PPOBuilder::new(MyEnvBuilder, 10)?;
    let _ppo4 = ppo_builder4.build()?;
    Ok(())
}

For real environments, make sure env_description accurately describes the flattened observation and action spaces. The policy builder uses those spaces to choose the policy distribution and network dimensions.

Algorithms

The algorithm examples show complete PPO and A2C configurations through r2l. They are useful when you want a runnable training loop before customizing lower-level hooks or samplers.

PPO

The PPO example trains on Pendulum-v1, writes training artifacts, reloads the best policy through InferenceRunner, and runs rendered inference episodes.

Run it from the workspace root:

cargo run -p r2l-examples --example ppo
use r2l::{Error, GymEnv, InferenceRunner, PPOBuilder, TrainingArtifactsConfig, TrainingLimit};

fn main() -> Result<(), Error> {
    const ENV_NAME: &str = "Pendulum-v1";

    // Path where the training artifacts are going to be stored. Training artifacts could include:
    // - Parameters of the model that was trained + the weights as a safetensor and the optional obs normalizer serialized
    // - Measurements on how the trained model perfoms after each training round
    // - Measurements on how long parts of the trainig run took
    const ARTIFACT_DIR: &str = "runs/pendulum";
    let artifacts_config = TrainingArtifactsConfig::new(ARTIFACT_DIR);

    // An environmnet builder how environments are to be constructed. Environment construction can
    // be elaborate (especially when working with external dependencies), so r2l opts to not pass
    // the environment directly (the environment would have to be Send for multi sampling), but
    // instead accepts anything that implements the `EnvBuilder` trait. Simplest example is just a
    // function/closure that returns the Env.
    let env_builder = || GymEnv::new(ENV_NAME, None);

    // The algorightm is constructed through a PPOBuilder. For A2C, the A2Cbuilder would be
    // equivalent. Builders expose a lot of common parameter setters. To check all the options, you
    // check https://docs.rs/r2l/latest/r2l/type.PPOBuilder.html.
    let mut ppo = PPOBuilder::new(env_builder, 10)?
        .with_training_artifacts(artifacts_config)
        .with_policy_hidden_layers(vec![64, 64])
        .with_lambda(0.95)
        .with_gamma(0.9)
        .with_learning_rate(0.001)
        .with_training_limit(TrainingLimit::rollouts(30))
        .build()?;

    // This kicks off and finishes training.
    ppo.train()?;

    // Once training in done, the training artifacts as serialized. You can reuse the trained model
    // by constructing an InferenceRunner. InferenceRunner can single step or run episodes on the
    // environment it recieves.
    let env = GymEnv::new(ENV_NAME, Some("human".to_owned()))?;
    let mut inference = InferenceRunner::load_from_env(ARTIFACT_DIR, env)?;
    for _ in 0..10 {
        inference.run_episode()?;
    }

    Ok(())
}

The important pieces are with_training_artifacts, which writes actor.safetensors, inference.yaml, and metrics files, and InferenceRunner::load_from_env, which rebuilds the inference runner from those saved files.

A2C

The A2C example selects the Candle backend, configures rollout collection, and uses a reporter channel to observe training statistics.

Run it from the workspace root:

cargo run -p r2l-examples --example a2c
use std::{
    sync::mpsc::{self, Receiver, Sender},
    thread,
};

use candle_core::Device;
use r2l::{A2CBuilder, A2CRolloutStats, SamplerExecutionMode, TrainingLimit};

fn main() -> anyhow::Result<()> {
    let (update_tx, update_rx): (Sender<A2CRolloutStats>, Receiver<A2CRolloutStats>) =
        mpsc::channel();

    let a2c_builder = A2CBuilder::gym("Pendulum-v1", 10)?
        .with_candle(Device::Cpu)
        .with_seed(0)
        .with_entropy_coefficient(0.2)
        .with_gradient_clipping(Some(0.5))
        .with_rollout_steps(2048)
        .with_execution_mode(SamplerExecutionMode::SingleThreaded)
        .with_training_limit(TrainingLimit::rollouts(300))
        .with_rollout_reporter(Some(update_tx));
    let mut a2c = a2c_builder.build()?;
    let t = thread::spawn(move || {
        while let Ok(stats) = update_rx.recv() {
            println!("avg reward: {}", stats.average_reward);
        }
    });
    a2c.train()?;
    drop(a2c);
    t.join()
        .map_err(|_| anyhow::anyhow!("A2C reporter thread panicked"))?;
    Ok(())
}

Both PPO and A2C builders expose the same broad setup concepts: choose an environment builder, select a backend, set rollout bounds, configure a learning schedule, then call build() and train().

Results

These results are seeded reference runs for r2l’s PPO implementation. They exercise both supported tensor backends across classic control, Box2D, MuJoCo, Bullet, and partially observable environments, providing broad integration coverage and a baseline for identifying behavioral regressions as the implementation evolves. Each configuration uses seed 0 to reduce variation between benchmark passes.

Stable Baselines3 (SB3) curves are included as a familiar external reference. They help make unexpected learning behavior easier to spot and are interesting to compare, but the runs are not designed as a controlled framework ranking.

Setup

All three implementations use the per-environment PPO configurations in the benchmark configuration, which are based on RL Baselines3 Zoo settings. The r2l runs use the same configuration for the Candle and Burn backends. Evaluation follows each implementation’s existing evaluation path.

The runs use r2l 0.0.3 at commit 6030dd8 and execute as Google Cloud Batch spot jobs in europe-west1 on CPU-only c4d-highcpu-2 workers. Each task is allocated one vCPU and 1,250 MiB of memory, with Rayon and OpenBLAS limited to one thread. Up to 36 tasks run concurrently.

Summary

The current snapshot contains evaluation data for 28 environments on each r2l backend and Stable Baselines3.

Environment familyEnvironments
Classic control5
Box2D4
MuJoCo10
Bullet8
PopGym1
Total28

Both r2l backends produce learning curves across the full suite through the same public API and PPO configuration. Many Candle and Burn curves follow similar broad trends, while some environments show substantial differences between the backends and SB3.

All evaluation curves

The horizontal axis is the number of environment interactions sampled during training. Faint lines show the recorded evaluation values and solid lines show moving averages. Higher reward is better, but reward scales should only be compared within an environment. The plots are generated by book/scripts/plot_results.py.

Classic control

Acrobot-v1

Acrobot-v1 evaluation curves

CartPole-v1

CartPole-v1 evaluation curves

MountainCar-v0

MountainCar-v0 evaluation curves

MountainCarContinuous-v0

MountainCarContinuous-v0 evaluation curves

Pendulum-v1

Pendulum-v1 evaluation curves

Box2D

BipedalWalker-v3

BipedalWalker-v3 evaluation curves

BipedalWalkerHardcore-v3

BipedalWalkerHardcore-v3 evaluation curves

LunarLander-v3

LunarLander-v3 evaluation curves

LunarLanderContinuous-v3

LunarLanderContinuous-v3 evaluation curves

MuJoCo

Ant-v4

Ant-v4 evaluation curves

HalfCheetah-v4

HalfCheetah-v4 evaluation curves

Hopper-v4

Hopper-v4 evaluation curves

Humanoid-v4

Humanoid-v4 evaluation curves

HumanoidStandup-v2

HumanoidStandup-v2 evaluation curves

InvertedDoublePendulum-v2

InvertedDoublePendulum-v2 evaluation curves

InvertedPendulum-v2

InvertedPendulum-v2 evaluation curves

Reacher-v2

Reacher-v2 evaluation curves

Swimmer-v4

Swimmer-v4 evaluation curves

Walker2d-v4

Walker2d-v4 evaluation curves

Bullet

AntBulletEnv-v0

AntBulletEnv-v0 evaluation curves

HalfCheetahBulletEnv-v0

HalfCheetahBulletEnv-v0 evaluation curves

HopperBulletEnv-v0

HopperBulletEnv-v0 evaluation curves

HumanoidBulletEnv-v0

HumanoidBulletEnv-v0 evaluation curves

InvertedDoublePendulumBulletEnv-v0

InvertedDoublePendulumBulletEnv-v0 evaluation curves

InvertedPendulumSwingupBulletEnv-v0

InvertedPendulumSwingupBulletEnv-v0 evaluation curves

ReacherBulletEnv-v0

ReacherBulletEnv-v0 evaluation curves

Walker2DBulletEnv-v0

Walker2DBulletEnv-v0 evaluation curves

PopGym

popgym-BattleshipEasy-v0

popgym-BattleshipEasy-v0 evaluation curves

On policy algorithms

On-policy algorithms

r2l separates an on-policy training run into three components:

  • an Agent owns the trainable policy and learns from trajectory batches;
  • a Sampler collects trajectories with a snapshot of the current actor;
  • OnPolicyAlgorithmHooks controls initialization, stopping, evaluation, and shutdown.

OnPolicyAlgorithm holds a runtime and a hook implementation. Its training loop repeatedly collects rollouts, invokes the post-rollout hook, updates the agent, and invokes the post-training hook. A hook can stop the loop by returning HookResult::Break. The runtime converts actor and trajectory tensors through the common R2lTensor representation when the agent and sampler use different tensor types.

On-policy algorithm overview

The public traits and their method contracts are documented in the r2l-core API.

Samplers

r2l-sampler provides two sampler implementations:

  • DirectSampler lets workers write transitions directly to output buffers;
  • StagedSampler receives transitions from workers and can update and apply clipped observation normalization before committing trajectories.

Both samplers support SamplerExecutionMode::SingleThreaded, which steps environments on the current thread, and SamplerExecutionMode::MultiThreaded, which assigns each environment to a worker thread. Gymnasium environments still execute Python code under Python’s interpreter lock, so threaded sampling should not be assumed to improve Gymnasium throughput.

Rollout collection is hook-driven. The high-level algorithm builders expose the standard policies through with_rollout_steps and with_rollout_episodes.

Sampler overview

Agents

r2l-agents contains the lower-level PPO, A2C, and VPG learning logic. r2l composes those agents with Candle or Burn learning modules and provides defaults for loss configuration, reporting, evaluation, and learning schedules.

Applications construct complete runs with PPOBuilder or A2CBuilder. Backend and sampler choices change the concrete builder type while preserving the shared configuration. Applications that need custom agents, samplers, or hook compositions can use the lower-level traits from r2l-core, r2l-agents, and r2l-sampler directly.

PPO hooks

The PPO agent exposes hooks at three points:

  1. after advantages and return targets are computed;
  2. after each PPO epoch, where the hook decides whether another epoch runs;
  3. after a minibatch loss is computed and before the optimizer update.

The default r2l hook uses these points for advantage normalization, entropy and value-loss coefficients, target-KL stopping, progress reporting, and statistics.

A2C hooks

The A2C agent exposes hooks before minibatching, before each optimizer update, and after all minibatches have been processed. The default hook provides advantage normalization, entropy and value-loss coefficients, reporting, and statistics.

Off policy algorithms

Off-policy algorithms

Off-policy algorithms are not implemented in r2l v0.0.3. The current training stack supports the on-policy PPO, A2C, and lower-level VPG implementations.

Off-policy replay buffers, agents, and high-level builders remain roadmap items. Applications targeting v0.0.3 should use the on-policy interfaces described in the user guide.