// the find
mees/calvin
CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
CALVIN is a simulated benchmark (PyBullet-based) for training and evaluating robots that follow long, chained natural-language instructions like 'open the drawer... now push the block in.' It's aimed at robotics/VLA researchers who want a standardized long-horizon language-conditioned manipulation task, not at people building production robot software.
Won the RA-L 2022 Best Paper Award and is still the reference benchmark for a large chunk of current VLA research — the README lists dozens of SOTA papers (2023-2025) built on it, so it's not an abandoned academic drop. The action/observation space is genuinely flexible: swap between absolute/relative/joint control and RGB/depth/tactile sensor combos via Hydra config overrides without touching code. The changelog is honest about real bugs (wrong scene_info.npy, bad language annotations in ABC/ABCD) and gives exact fix commands instead of silently patching and moving on.
Setup is heavy and dated: pinned to Python 3.8 via conda, a git submodule (calvin_env) you must remember to pull, and a README-acknowledged pyhash/setuptools install footgun. It's built around PyBullet + EGL GPU rendering with cluster/SLURM assumptions baked in (see the multi-GPU EGL device-mismatch FAQ entry), so getting it running well on a single dev machine takes real fiddling. There's no visible test suite in the repo, and evaluation of a custom model requires hand-implementing an interface class in evaluate_policy.py rather than a documented plugin API. Datasets are large (the smallest 'debug' split is 1.3GB, full splits much bigger) and hosted on an external university server, not versioned alongside the code.