finds.dev← search

// the find

pytorch/rl

★ 3,594 · Python · MIT · updated Oct 2026

A modular, primitive-first, python-first PyTorch library for Reinforcement Learning.

TorchRL is PyTorch's own RL library, built around TensorDict (a dict-like tensor container) as the common data format connecting environments, collectors, replay buffers, and loss modules. It's aimed at people who want composable RL building blocks inside the PyTorch ecosystem rather than a monolithic trainer or a Gym-wrapper toy repo — covers everything from PPO/SAC to MARL, model-based RL, and LLM post-training (GRPO).

TensorDict is a real design decision, not just plumbing — named, batched, nested keys mean a collector output can go straight into a replay buffer and then a loss module without anyone writing glue code to reshape tuples or dicts. The collector/replay-buffer infrastructure is genuinely substantial: async multiprocess and distributed collectors, memmap-backed storage, prioritized replay with optional CUDA kernels — the kind of plumbing most RL repos never bother building correctly. Algorithm and environment coverage is wide and current: PPO/SAC/DQN/TD3/Dreamer/Decision Transformer/MAPPO/IPPO plus wrappers for Gymnasium, DMControl, Brax, Jumanji, PettingZoo, VMAS, Isaac Lab. CI discipline stands out — benchmark tracking over time, flaky-test tracking, and separate test matrices for dozens of optional dependencies, which is unusual rigor for a research-adjacent library.

It's still labeled a PyTorch 'beta feature' in their own README — breaking changes ship across releases, so pin your version or expect to rebase onto their deprecation cycle. TensorDict is a mandatory abstraction, not an optional one: every policy, env, and loss has to speak in named keys, and debugging a shape mismatch buried in a nested TensorDict is less obvious than a plain tensor stack trace. The prioritized replay buffer's C++ extension has known undefined-symbol issues tied to PyTorch version mismatches (they link their own versioning troubleshooting doc for it), which suggests that ABI boundary is still fragile. The surface area is enormous — recurrent RL, MARL, model-based, robotics, LLM post-training — but the trusted-collaborator list is only three people, so any given corner of this gets thinner review attention than the breadth implies.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →