// the find
huangwl18/VoxPoser
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Official demo implementation of VoxPoser, a Stanford paper that has an LLM write code to compose 3D voxel-space value maps (affordance, avoidance, rotation, velocity) from a language instruction, which a greedy planner then turns into a robot trajectory. This is for robotics/embodied-AI researchers who want to reproduce or build on the zero-shot manipulation approach, not for anyone wanting a deployable robot stack.
The LMP (Language Model Program) decomposition is genuinely clever and the code structure mirrors the paper closely — interfaces.py, planners.py, and the prompts/rlbench directory make it easy to see exactly which prompt produces which map. LLM_cache.py disk-caches model outputs, which matters a lot for iterating on a project where every run fires off several LLM calls. The README is upfront that this is a demo on RLBench's ground-truth object masks, not the real-world perception stack, and it names concrete modern replacements (SAM 2, OWL-ViT) instead of just gesturing at 'future work.'
The actual hard part of the real system — open-vocabulary detection plus segmentation and tracking to turn object names into masks — is explicitly not in this repo; you get RLBench's oracle masks instead, so going from this to a real robot is a substantial rebuild, not a port. Setup requires PyRep, RLBench, and CoppeliaSim, a notoriously fragile install chain that typically needs a display and pins you to old dependency versions. It's hardwired to the OpenAI API with no provider abstraction, there's no eval harness to reproduce the paper's quantitative results, and the only usage path is a single Jupyter notebook with no test suite.