Data & Benchmarks
CALVIN
CALVIN (Composing Actions from Language and Vision) is a simulated benchmark for long-horizon, language-conditioned manipulation, introduced by Mees et al. in 2022. A Franka arm in four tabletop environments must complete chains of up to five sequential instructions such as opening drawers, pushing blocks, and toggling lights, using onboard RGB observations. Its evaluation protocol measures how many consecutive subtasks a policy completes, testing both instruction grounding and long-horizon consistency.
Why it matters for physical AI
CALVIN remains a standard yardstick for language-conditioned policies and hierarchical vision-language-action systems, exposing failure modes in instruction chaining that single-task benchmarks miss.
Related terms
Build physical AI
Put these concepts to work on real hardware
Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.