Robot Learning
Visual Imitation
Visual imitation is imitation learning in which the policy consumes raw camera images rather than privileged state, and, in a broader sense, learning behaviors from visual observation of another agent, including humans, without access to expert actions. The first form dominates modern manipulation through methods like ACT and Diffusion Policy; the second, learning from human video, must additionally bridge the embodiment gap between demonstrator and robot.
Why it matters for physical AI
Cameras are the most scalable sensor for collecting demonstrations, and progress in visual imitation determines how directly the vast supply of human manipulation video can be converted into robot skills.
Related terms
Build physical AI
Put these concepts to work on real hardware
Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.