Data & Benchmarks

COLOSSEUM

COLOSSEUM is a simulation benchmark introduced in 2024 for evaluating the generalization of manipulation policies under systematic environmental perturbations. Built on RLBench tasks, it applies controlled variations across factors including object color, texture, size, lighting, distractor objects, camera pose, and physical properties, measuring how policy success degrades relative to the training conditions. It quantifies robustness gaps that standard same-distribution evaluations conceal.

Why it matters for physical AI

Generalization under visual and physical variation is the central obstacle to deploying learned manipulation, and perturbation benchmarks like COLOSSEUM provide the controlled measurements needed to compare robustness across policy architectures.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.