Foundation Models

Gemini Robotics

Gemini Robotics is a family of vision-language-action models from Google DeepMind, introduced in 2025, that builds robot control on top of the Gemini multimodal foundation model. The initial release paired a VLA that outputs robot actions with Gemini Robotics-ER, an embodied-reasoning variant for spatial understanding, and demonstrated dexterous bimanual tasks and cross-embodiment operation; subsequent versions added on-device execution and agentic long-horizon behavior.

Why it matters for physical AI

Frontier-lab entries like Gemini Robotics test whether internet-scale multimodal pretraining transfers decisive advantages to control, and their embodied-reasoning splits illustrate one architecture for coupling semantic planning with low-level action.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.