Figure AI Tackles the Physical Data Gap with New Program

Humanoid robotics firms are racing to build the physical world's largest dataset, a strategy mirroring how Tesla leveraged its fleet to dominate autonomous driving.
Figure AI has launched a program called Index, designed to construct what it describes as the world's largest and most diverse physical dataset. The move highlights a critical bottleneck in robotics: unlike software companies that can generate synthetic data, robots need to learn from real-world interactions. This approach mirrors the strategy that gave Tesla a significant edge in autonomous driving, where every vehicle on the road became a data collection unit.
The core challenge for humanoid robots is that their intelligence depends on understanding physical environments, not just digital ones. By systematically gathering data on how objects move, how surfaces feel, and how spaces are navigated, Figure aims to train its models to operate reliably in unstructured settings. This is a departure from the traditional method of building robots for specific tasks and hoping they generalize well.
Tesla’s fleet became a data engine
Tesla’s advantage in self-driving technology was not just its software, but its ability to collect vast amounts of real-world video. Every electric vehicle sold included hardware for autonomous driving, regardless of whether the customer paid for the feature. This allowed the company to amass billions of miles of driving data, which became the foundation for its full self-driving software.
This data moat is difficult to replicate. While competitors like Waymo have collected millions of miles of autonomous driving data, Tesla’s fleet has generated over 14 billion miles. This sheer volume of real-world experience allowed Tesla to train its vision-based systems on a scale that dedicated robotaxi fleets cannot match. The data also served as the basis for developing its humanoid robot, Optimus.
Physical data is harder to obtain
Replicating this model in robotics is more complex than in driving. A car drives on roads, which are relatively predictable. A humanoid robot must navigate warehouses, homes, and factories, where every object and interaction is different. Figure’s Index program seeks to solve this by creating a centralized repository of physical interactions, essentially turning every robot deployment into a data source.
The trade-off for this strategy is cost and infrastructure. Building a dataset requires not just robots, but the ability to process, store, and analyze the resulting data. For Figure, this means investing heavily in backend systems and ensuring that the data collected is high-quality and diverse enough to be useful for training.
The catch behind the data rush
The primary catch is that data volume does not automatically translate to capability. Tesla’s data advantage took years to build and still requires constant refinement. Figure faces the same reality: collecting data is the easy part; making sense of it is the hard part. The company will need to ensure that its dataset is not just large, but actually improves robot performance in real-world scenarios.
As reported by GN auto tech/robotics, this competition for physical data is becoming the new battleground in robotics. The firm that can best leverage real-world experience to train its models will likely have the upper hand. For now, the race is on to see who can build the most useful dataset, not just the biggest one.






