Figure's 56% Success Rate Redefines Humanoid Robot Expectations

A new benchmark suggests humanoid robots can now handle household chores in unfamiliar homes without prior training, though a significant failure rate remains.
In a significant shift for embodied AI, a leading robotics firm has demonstrated that its latest humanoid model can perform complex household tasks in completely unfamiliar environments. The robot successfully completed chores such as making beds and folding towels in thirty different homes where it had never been seen or trained. This marks a notable improvement in what the industry calls generalization, allowing the machine to apply skills learned from other scenarios to new, unseen settings.
The achievement, reported by GN auto tech/robotics, sees the task success rate jump from a previous 9% to 56%. While this is a substantial leap, it means that nearly half of the attempts still end in failure. The evaluation was strict: there were no partial points, and any human intervention counted as a failed trial. This rigorous approach provides a clearer, albeit harsh, picture of the current capabilities of household robots.
Strict Testing in Unfamiliar Homes
The trial involved thirty families in the San Francisco Bay Area who allowed the robot into their private residences. The robot was assigned random tasks, such as collecting scattered toys or arranging pillows, without any specific adaptation for the unique layout of each house. The objects it handled, from specific towels to particular beds, were all new to the machine. This setup tested the robot's ability to generalize its physical movements across diverse and unpredictable domestic spaces.
The results varied by task, with bed-making showing a 67% success rate and towel folding at 62%, while toy collection lagged behind at 40%. This variance highlights that while the robot has improved in some areas, it still struggles with others. The inability to guarantee success in every home means that for practical consumer use, the robot cannot yet be relied upon to handle chores independently without a high risk of needing human help.
Pre-Training Drives Performance Gains
The core of this improvement lies in a new pre-training method that uses human behavioral data. By training on a massive dataset of human movements, the robot learns to understand physical interactions before it ever performs a specific task. This approach allowed the system to achieve a 56% success rate compared to just 9% for a model trained from scratch on the same data. The technique effectively scales the robot's understanding of the physical world without requiring it to learn every single chore from zero.
However, this method comes with its own trade-offs. The system requires a significant volume of human data to generate meaningful improvements, and the scaling laws observed suggest that further gains will depend on exponentially larger datasets. For other companies in the field, this sets a high bar for data collection and processing. The ability to predict performance based on data volume is a strong signal, but it also underscores the heavy reliance on external human input to drive progress in autonomy.
Remaining Limitations for Daily Use
Despite the headline number, a 56% success rate is far from the reliability required for a seamless domestic assistant. In a real-world scenario, failing nearly half the time means the robot would frequently need to be corrected by a human, negating much of the time-saving promise of automation. The trade-off between improved generalization and consistent reliability remains the central challenge for the industry.
As the technology advances, the focus will likely shift from broad generalization to consistent execution. Until the failure rate drops significantly, these robots will remain experimental tools rather than practical household helpers. The next steps for the industry involve closing this reliability gap, ensuring that the impressive generalization capabilities translate into dependable, everyday utility for consumers.






