Off-policy algorithms
Off-policy algorithms are not implemented in r2l v0.0.3. The current
training stack supports the on-policy PPO, A2C, and lower-level VPG
implementations.
Off-policy replay buffers, agents, and high-level builders remain roadmap items. Applications targeting v0.0.3 should use the on-policy interfaces described in the user guide.