Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Getting started

For most applications, r2l is the main dependency. It provides complete PPO and A2C builders while the lower-level workspace crates define environments, samplers, agents, and backend integrations. This getting started guide primarily uses the r2l crate, which itself builds on lower-level crates. Implementing a custom environment also requires the error type from r2l-core. If the current setup does not satisfy you, the lower level hooks allow for a lot of hackability.

Shortest setup

The shortest Gymnasium-based PPO setup is:

r2l = { version = "0.0.3", features = ["gym"] }
extern crate r2l;
use r2l::PPOBuilder;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let builder = PPOBuilder::gym("Pendulum-v1", 4)?;
    let mut algorithm = builder.build()?;
    algorithm.train()?;
    Ok(())
}

The gym feature requires Python 3.11 or newer and the gymnasium Python package. A Gymnasium environment id is passed to GymEnvBuilder, which maps supported Gymnasium spaces to r2l space descriptions. Discrete spaces currently require start = 0; non-zero start values are not supported. The builders support other environment types through a different construction, which will be introduced later on.

Saving training artifacts

We rarely want to train algorithms for the sake of it. Once an algorithm is learned, we are usually curious about

  • how the chosen hyperparameters affect learning performance
  • how to run inference using the model
  • less often, maybe we are curious about the running performance of the algorithm

r2l allows saving these artifacts using the with_training_artifacts builder method.

extern crate r2l;
use r2l::{PPOBuilder, TrainingArtifactsConfig};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let builder = PPOBuilder::gym("Pendulum-v1", 4)?;
    let artifacts_config = TrainingArtifactsConfig::new("runs/pendulum")
        .with_evaluation_results(true)
        .with_training_timings(true)
        .with_inference_artifacts(true);
    let mut algorithm = builder
        .with_training_artifacts(artifacts_config)
        .build()?;
    algorithm.train()?;
    Ok(())
}

Under the hood, r2l evaluates the policy with dedicated environments that are reset and reused between evaluation passes, and remembers the best-performing policy. Evaluation runs after every rollout by default. Its frequency and other settings can be customized with EvaluationSettings. Once training is done, the following new files are created in the artifacts folder:

$ tree runs/pendulum
runs/pendulum
├── actor.safetensors
├── evaluations.csv
├── inference.yaml
└── training_timings.csv

1 directory, 4 files

Running inference

Using the inference artifacts and a new environment, you can create an InferenceRunner.

extern crate r2l;
use r2l::{GymEnv, InferenceRunner};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let env = GymEnv::new("Pendulum-v1", Some("human".to_owned()))?;
    let mut inference = InferenceRunner::load_from_env("runs/pendulum", env)?;
    for _ in 0..4 {
        inference.run_episode()?;
    }
    Ok(())
}

The inference configuration records the backend, policy architecture settings, and observation normalization mode needed to rebuild the policy. It does not contain training hyperparameters such as the learning rate or discount factor.

Applications that receive observations outside an Env can load an InferencePolicy directly and request actions from raw observations. For a stateful control loop, implement InferenceEnv and initialize an InferenceRunner with the first raw observation. Full Env implementations also support resetting and running complete episodes.

Environments

Environments implement the Env trait.

/// Environment interface used by samplers.
pub trait Env {
    /// Tensor type used for observations and actions.
    type Tensor: R2lTensor;

    /// Resets the environment and returns the initial observation.
    ///
    /// # Errors
    ///
    /// Returns an error if the environment cannot be reset.
    fn reset(&mut self, seed: u64) -> Result<Self::Tensor, Error>;

    /// Applies one action and returns the resulting transition snapshot.
    ///
    /// # Errors
    ///
    /// Returns an error if the environment cannot apply the action.
    fn step(&mut self, action: Self::Tensor) -> Result<Snapshot<Self::Tensor>, Error>;

    /// Returns static observation/action space metadata.
    fn env_description(&self) -> EnvDescription<Self::Tensor>;
}

Algorithm builders receive an EnvBuilder rather than a concrete environment so that each sampler worker can construct its environment in the place where it runs.

/// Factory for constructing environments of one compatible type.
pub trait EnvBuilder: Sync + Send + 'static {
    /// Environment type produced by this builder.
    type Env: Env;

    /// Builds a fresh environment instance.
    ///
    /// # Errors
    ///
    /// Returns an error if the environment cannot be constructed.
    fn build_env(&self) -> Result<Self::Env, Error>;

    /// Returns the environment description for produced environments.
    ///
    /// # Errors
    ///
    /// Returns an error if a representative environment cannot be constructed.
    fn env_description(&self) -> Result<EnvDescription<<Self::Env as Env>::Tensor>, Error> {
        let env = self.build_env()?;
        Ok(env.env_description())
    }
}

A closure or function returning r2l_core::error::Result<E> automatically implements EnvBuilder.

let env_builder = || -> r2l_core::error::Result<MyEnv> { Ok(MyEnv) };
let ppo_builder = PPOBuilder::new(env_builder, 10)?;
let ppo = ppo_builder.build()?;

For a more detailed example of how to implement the Env and EnvBuilder traits, see the environment building example.

Hyperparameters

Both the PPO and A2C builders expose a great deal of hyperparameters that can be tuned.

Backends

The algorithm builders default to Candle. Use with_candle(device) to choose a Candle device explicitly, or with_burn() to use the default Burn autodifferentiation backend.

Backend selection changes the concrete builder type, so call it before methods that are specific to a chosen backend when following compiler suggestions.

Rollout collection

with_rollout_steps(n) collects n steps per environment for each rollout. with_rollout_episodes(n) switches to episode-bounded sampling and collects n completed episodes per environment.

The builders default to SamplerExecutionMode::MultiThreaded, which runs workers on dedicated threads. Use SamplerExecutionMode::SingleThreaded to step workers sequentially on the calling thread:

extern crate r2l;
use r2l::{PPOBuilder, SamplerExecutionMode};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let builder = PPOBuilder::gym("Pendulum-v1", 4)?
        .with_execution_mode(SamplerExecutionMode::SingleThreaded);
    Ok(())
}

Gymnasium calls still execute under Python’s interpreter lock, so threaded execution should not be assumed to improve Python-environment throughput.

with_observation_normalizer(Some(clip)) switches to staged sampling and enables observation normalization. Passing None switches to staged sampling without a normalizer. Step-bounded rollouts can also normalize discounted rewards with with_reward_normalizer(gamma, clip_reward).

Training schedules

TrainingLimit::rollouts(n) stops after n rollout collections. TrainingLimit::steps(n) stops after at least n sampled environment steps across all workers.

LearningRateSchedule::Constant(rate) keeps the configured rate fixed. LearningRateSchedule::Linear(rate) decays it from rate to zero over the configured training schedule.

For the underlying traits and hook points, continue with On-policy algorithms. For exact builder methods, use the r2l reference.