Foundation Models

Video Prediction Model

A video prediction model is a generative model that forecasts future video frames conditioned on past frames and, optionally, on actions or language. Action-conditioned video prediction enabled early visual model-predictive control in the Visual Foresight line of work, and modern large-scale video generators are increasingly treated as learned world simulators from which robot policies can plan or be distilled.

Why it matters for physical AI

A model that predicts how scenes evolve under actions is effectively a physics simulator learned from data, offering a path to planning and policy training that scales with web video rather than hand-built simulation.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.