Foundation Models

VoxPoser

VoxPoser is a zero-shot manipulation method from Huang et al. (2023) in which a large language model writes code that composes vision-language-model queries into 3D voxel value maps encoding task affordances and constraints, and a motion planner then synthesizes end-effector trajectories by optimizing over those maps. It executes free-form language instructions without any task-specific policy training or demonstrations.

Why it matters for physical AI

VoxPoser demonstrated that foundation models can ground language into executable spatial plans without robot training data, an influential template for zero-shot manipulation and for hybrids of learned reasoning with classical planning.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.