Foundation Models

AutoRT

AutoRT is a system from Google DeepMind, presented in 2024, that uses foundation models to orchestrate autonomous data collection across a fleet of mobile manipulators. A vision-language model describes each robot's scene, a large language model proposes candidate tasks and filters them through a written constitution of safety and feasibility rules, and tasks are then executed by teleoperators, scripted policies, or learned policies such as RT-2. It coordinated over twenty robots collecting tens of thousands of episodes.

Why it matters for physical AI

Scaling robot learning requires data engines, not just models, and AutoRT prototyped how language models can supervise fleets to gather diverse real-world experience with minimal human oversight.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.