Foundation Models

DINOv2

DINOv2 is a family of self-supervised Vision Transformers released by Meta AI in 2023, trained with a discriminative self-distillation objective on a curated dataset of 142 million images. Its frozen features transfer strongly to classification, segmentation, depth estimation, and dense correspondence without fine-tuning. The 2025 successor DINOv3 scaled the approach further.

Why it matters for physical AI

Frozen self-supervised backbones like DINOv2 are widely used as visual encoders for manipulation policies and correspondence-based grasping, providing robust, semantically rich features from limited robot data.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.