Figure's Helix 2.5 Tests Humanoid Robots in Unseen Homes

Figure AI has released a new neural network that allows humanoid robots to perform complex household tasks in homes they have never encountered, without requiring new training data for each location.
Figure AI has introduced Helix 2.5, a neural network designed to enable humanoid robots to perform whole-body household tasks in unfamiliar environments. The system was tested across 30 different homes in the Bay Area, where the robot completed chores such as tidying toys and making beds without collecting any new data or adapting its model to those specific layouts. This approach aims to solve a major hurdle in robotics: the ability to generalize skills learned from human behavior to new, unseen places.
The company states that this new model significantly reduces the cost and data requirements for deploying robotic behaviors. By leveraging a foundation model pretrained on a large dataset of human actions, Helix 2.5 requires half as much task-specific data as previous versions while expanding its operational scope from a single environment to dozens of unique homes. The primary goal is to move away from location-specific training and toward a more flexible, general-purpose capability.
Testing Across Thirty Unseen Environments
The evaluation process involved placing the robot in 30 homes that were not part of its training dataset. The tasks required the robot to use its whole body to pick up 13 to 15 scattered toys and place them in a basket, fold towels, and make a bed. Success was defined strictly: every single toy or towel had to be handled, and the bed had to be fully made with pillows and comforter corners in place. The same model checkpoint was used for all 30 homes, ensuring that no local adjustments were made.
Figure reported that the robot demonstrated improved self-correction during these longer tasks. When it made a mistake, such as misplacing an object or losing balance, it could step back, reposition its body, and continue the task. This ability to recover from errors is crucial for real-world deployment, as it allows the robot to handle the unpredictability of messy household environments without human intervention.
Pretraining Drives Significant Performance Gains
The core of Helix 2.5’s capability comes from pretraining on Figure’s Index dataset, which contains recorded human behavior. In controlled comparisons, a policy trained with this pretraining data succeeded on 56% of trials, compared to only 9% for a model trained from scratch using the same task-specific data. This suggests that learning from broad human examples provides a much stronger foundation for robotic action than narrow, task-specific instruction alone.
The company notes that performance improves predictably as more human-behavior data is added. This scaling relationship allows Figure to forecast the results of larger training runs with high accuracy. By investing heavily in compute resources, the company aims to continue expanding this dataset, which currently generates a substantial amount of human experience data per second, further enhancing the robot's ability to generalize across diverse settings.
Trade-offs in Data and Deployment
While the zero-shot capability is a significant step forward, it relies heavily on the quality and volume of the pretraining data. The system does not adapt to each home individually, meaning it must rely entirely on the patterns learned from human examples. If a home contains objects or layouts that differ significantly from the training data, the robot may struggle. This creates a trade-off between flexibility and reliability, as the robot must balance its need to generalize with the risk of encountering unfamiliar scenarios.
Furthermore, the high computational cost of generating and processing this data remains a barrier. Figure has committed billions of dollars in compute to refine these models, but the economic viability of deploying such robots at scale depends on reducing these costs. The current success rate of 56% indicates that while the technology is advancing, it is not yet at a level where it can be relied upon for critical tasks without human oversight. The path to widespread adoption requires further improvements in consistency and error handling.






